EDBT 2026 Demo / reviewers in the wild / expert
Duc Thien Nguyen
dblp:118/3928
· DBLP profile ↗
14ranked-venue papers
8as first author
4since 2021 · last 2025
0009-0009-7485-7852ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 64% Multi-agent systems · 23% Planning, search and constraint satisfaction · 13% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 61% Distributed computing theory · 30% Algorithms and data structures · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Smart cities and intelligent transportation · 100% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.8 | 3 | 2018 | Credit Assignment For Collective Multiagent RL With Global Rewards · NeurIPS 2018 Policy Gradient With Value Function Approximation For Collective Multiagent Planning · NIPS 2017 Decentralized Multi-Agent Reinforcement Learning in Average-Reward Dynamic DCOPs · AAAI 2014 |
Machine learning › Reinforcement learning
actor-critic methods |
0.6 | 2 | 2018 | Credit Assignment For Collective Multiagent RL With Global Rewards · NeurIPS 2018 Policy Gradient With Value Function Approximation For Collective Multiagent Planning · NIPS 2017 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent decision making |
0.4 | 1 | 2019 | Multiagent Decision Making For Maritime Traffic Management · AAAI 2019 |
Mathematical optimization › discrete optimization
mixed integer linear programming |
0.4 | 1 | 2019 | Improving Law Enforcement Daily Deployment Through Machine Learning-Informed Optimization under Uncertainty · IJCAI 2019 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment |
0.3 | 1 | 2018 | Credit Assignment For Collective Multiagent RL With Global Rewards · NeurIPS 2018 |
Knowledge, reasoning and agents › Multi-agent systems › multi-agent decision making
decentralized markov decision process |
0.3 | 1 | 2017 | Collective Multiagent Sequential Decision Making Under Uncertainty · AAAI 2017 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
multi-agent planning |
0.3 | 1 | 2017 | Collective Multiagent Sequential Decision Making Under Uncertainty · AAAI 2017 |
Machine learning › Reinforcement learning
value function approximation |
0.3 | 1 | 2017 | Policy Gradient With Value Function Approximation For Collective Multiagent Planning · NIPS 2017 |
Machine learning › Reinforcement learning › markov decision process
average-reward reinforcement learning |
0.2 | 1 | 2014 | Decentralized Multi-Agent Reinforcement Learning in Average-Reward Dynamic DCOPs · AAAI 2014 |
Knowledge, reasoning and agents › Multi-agent systems
distributed constraint optimization |
0.2 | 1 | 2014 | Decentralized Multi-Agent Reinforcement Learning in Average-Reward Dynamic DCOPs · AAAI 2014 |
Machine learning › Reinforcement learning › large-scale reinforcement learning
distributed reinforcement learning |
0.2 | 1 | 2014 | Decentralized Multi-Agent Reinforcement Learning in Average-Reward Dynamic DCOPs · AAAI 2014 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › constraint programming
tractable class |
0.2 | 1 | 2014 | A Simple Polynomial-Time Randomized Distributed Algorithm for Connected Row Convex Constraints · AAAI 2014 |
Distributed computing theory › distributed algorithms › distributed coordination
distributed constraint satisfaction |
0.2 | 1 | 2014 | A Simple Polynomial-Time Randomized Distributed Algorithm for Connected Row Convex Constraints · AAAI 2014 |
Smart cities and intelligent transportation › traffic management
maritime traffic management |
0.1 | 1 | 2019 | Multiagent Decision Making For Maritime Traffic Management · AAAI 2019 |
Algorithms and data structures
randomized algorithms |
0.1 | 1 | 2014 | A Simple Polynomial-Time Randomized Distributed Algorithm for Connected Row Convex Constraints · AAAI 2014 |
Methods — techniques the papers use, named apart from their topics
policy gradient · 1.4traffic simulation · 0.8sample average approximation · 0.8multi-agent reinforcement learning · 0.8machine learning · 0.8iterated local search · 0.8actor-critic · 0.6polynomial-time analysis · 0.4difference rewards · 0.3value function decomposition · 0.3lifted inference · 0.3finite partial exchangeability · 0.3collective graphical model · 0.3randomized algorithm · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mastering the Craft of Data Synthesis for CodeLLMsabstractMeng Chen, Philip Arthur, Qianyu Feng, Cong Duy Vu Hoang, Yu-Heng Hong, Mahdi Kazemi Moghaddam, Omid Nezami, Duc Thien Nguyen, Gioacchino Tangari, Duy Vu, Thanh Vu, Mark Johnson, Krishnaram Kenthapadi, Don Dharmasiri, Long Duong, Yuan-Fang Li. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Philip Arthur, Qianyu Feng, Cong Duy Vu Hoang, Yu-Heng Hong, Mahdi Kazemi Moghaddam, Omid Nezami, Duc Thien Nguyen, Gioacchino Tangari, Duy Vu, Mark Johnson 0001, Krishnaram Kenthapadi, Don Dharmasiri, Long Duong, Yuan-Fang Li |
NAACL (Long Papers) | 8 |
| 2025 | Imputation of time-varying edge flows in graphs by multilinear kernel regression and manifold learning
Duc Thien Nguyen, Konstantinos Slavakis, Dimitris A. Pados |
Signal Process. | 1 |
| 2024 | Multi-Linear Kernel Regression and Imputation VIA Manifold Learning: the Dynamic MRI CaseabstractThis paper introduces an efficient multi-linear nonparametric (kernel-based) approximation framework for data regression and imputation. Data features are assumed to reside in or close to a smooth and userunknown manifold embedded in a reproducing kernel Hilbert space. Landmark points are identified to describe concisely the point cloud of features by linear approximating patches which mimic the concept of tangent spaces to smooth manifolds. The multi-linear model effects dimensionality reduction, enables efficient computations, and extracts data patterns and their geometry without any training data or additional information. Numerical tests on highly accelerated dynamic magnetic-resonance imaging (dMRI) data demonstrate remarkable improvements in efficiency and accuracy of the proposed approach over its predecessors and popular "shallow" data modeling methods, while offering substantial computational savings with regards to a deep-image-prior scheme. Duc Thien Nguyen, Konstantinos Slavakis |
ICASSP | 1 |
| 2022 | Neural-progressive hedging: Enforcing constraints in reinforcement learning with stochastic programmingabstractWe propose a framework, called neural-progressive hedging (NP), that leverages stochastic programming during the online phase of executing a reinforcement learning (RL) policy. The goal is to ensure feasibility with respect to constraints and risk-based objectives such as conditional value-at-risk (CVaR) during the execution of the policy, using probabilistic models of the state transitions to guide policy adjustments. The framework is particularly amenable to the class of sequential resource allocation problems since feasibility with respect to typical resource constraints cannot be enforced in a scalable manner. The NP framework provides an alternative that adds modest overhead during the online phase. Experimental results demonstrate the efficacy of the NP framework on two continuous real-world tasks: (i) the portfolio optimization problem with liquidity constraints for financial planning, characterized by non-stationary state distributions; and (ii) the dynamic repositioning problem in bike sharing systems, that embodies the class of supply-demand matching problems. We show that the NP framework produces policies that are better than deep RL and other baseline approaches, adapting to non-stationarity, whilst satisfying structural constraints and accommodating risk measures in the resulting policies. Additional benefits of the NP framework are ease of implementation and better explainability of the policies. Supriyo Ghosh, Laura Wynter, Shiau Hong Lim, Duc Thien Nguyen |
UAI | 4 |
| 2019 | Multiagent Decision Making For Maritime Traffic ManagementabstractWe address the problem of maritime traffic management in busy waterways to increase the safety of navigation by reducing congestion. We model maritime traffic as a large multiagent systems with individual vessels as agents, and VTS authority as the regulatory agent. We develop a maritime traffic simulator based on historical traffic data that incorporates realistic domain constraints such as uncertain and asynchronous movement of vessels. We also develop a traffic coordination approach that provides speed recommendation to vessels in different zones. We exploit the nature of collective interactions among agents to develop a scalable policy gradient approach that can scale up to real world problems. Empirical results on synthetic and real world problems show that our approach can significantly reduce congestion while keeping the traffic throughput high. Arambam James Singh, Duc Thien Nguyen, Akshat Kumar, Hoong Chuin Lau |
AAAI | 2 |
| 2019 | Improving Law Enforcement Daily Deployment Through Machine Learning-Informed Optimization under UncertaintyabstractUrban law enforcement agencies are under great pressure to respond to emergency incidents effectively while operating within restricted budgets. Minutes saved on emergency response times can save lives and catch criminals, and a responsive police force can deter crime and bring peace of mind to citizens. To efficiently minimize the response times of a law enforcement agency operating in a dense urban environment with limited manpower, we consider in this paper the problem of optimizing the spatial and temporal deployment of law enforcement agents to predefined patrol regions in a real-world scenario informed by machine learning. To this end, we develop a mixed integer linear optimization formulation (MIP) to minimize the risk of failing response time targets. Given the stochasticity of the environment in terms of incident numbers, location, timing, and duration, we use Sample Average Approximation (SAA) to find a robust deployment plan. To overcome the sparsity of real data, samples are provided by an incident generator that learns the spatio-temporal distribution and demand parameters of incidents from a real world historical dataset and generates sets of training incidents accordingly. To improve runtime performance across multiple samples, we implement a heuristic based on Iterated Local Search (ILS), as the solution is intended to create deployment plans quickly on a daily basis. Experimental results demonstrate that ILS performs well against the integer model while offering substantial gains in execution time. Jonathan Chase, Duc Thien Nguyen, Hoong Chuin Lau |
IJCAI | 2 |
| 2019 | Distributed Gibbs: A Linear-Space Sampling-Based DCOP AlgorithmabstractResearchers have used distributed constraint optimization problems (DCOPs) to model various multi-agent coordination and resource allocation problems. Very recently, Ottens et al. proposed a promising new approach to solve DCOPs that is based on confidence bounds via their Distributed UCT (DUCT) sampling-based algorithm. Unfortunately, its memory requirement per agent is exponential in the number of agents in the problem, which prohibits it from scaling up to large problems. Thus, in this article, we introduce two new sampling-based DCOP algorithms called Sequential Distributed Gibbs (SD-Gibbs) and Parallel Distributed Gibbs (PD-Gibbs). Both algorithms have memory requirements per agent that is linear in the number of agents in the problem. Our empirical results show that our algorithms can find solutions that are better than DUCT, run faster than DUCT, and solve some large problems that DUCT failed to solve due to memory limitations. Duc Thien Nguyen, William Yeoh 0001, Hoong Chuin Lau, Roie Zivan |
J. Artif. Intell. Res. | 1 |
| 2018 | Credit Assignment For Collective Multiagent RL With Global RewardsabstractScaling decision theoretic planning to large multiagent systems is challenging due to uncertainty and partial observability in the environment. We focus on a multiagent planning model subclass, relevant to urban settings, where agent interactions are dependent on their ``collective influence'' on each other, rather than their identities. Unlike previous work, we address a general setting where system reward is not decomposable among agents. We develop collective actor-critic RL approaches for this setting, and address the problem of multiagent credit assignment, and computing low variance policy gradient estimates that result in faster convergence to high quality solutions. We also develop difference rewards based credit assignment methods for the collective setting. Empirically our new approaches provide significantly better solutions than previous methods in the presence of global rewards on two real world problems modeling taxi fleet optimization and multiagent patrolling, and a synthetic grid navigation domain. Duc Thien Nguyen, Akshat Kumar, Hoong Chuin Lau |
NeurIPS | 1 |
| 2017 | Collective Multiagent Sequential Decision Making Under UncertaintyabstractMultiagent sequential decision making has seen rapid progress with formal models such as decentralized MDPs and POMDPs. However, scalability to large multiagent systems and applicability to real world problems remain limited. To address these challenges, we study multiagent planning problems where the collective behavior of a population of agents affects the joint-reward and environment dynamics. Our work exploits recent advances in graphical models for modeling and inference with a population of individuals such as collective graphical models and the notion of finite partial exchangeability in lifted inference. We develop a collective decentralized MDP model where policies can be computed based on counts of agents in different states. As the policy search space over counts is combinatorial, we develop a sampling based framework that can compute open and closed loop policies. Comparisons with previous best approaches on synthetic instances and a real world taxi dataset modeling supply-demand matching show that our approach significantly outperforms them w.r.t. solution quality. Duc Thien Nguyen, Akshat Kumar, Hoong Chuin Lau |
AAAI | 1 |
| 2017 | Policy Gradient With Value Function Approximation For Collective Multiagent PlanningabstractDecentralized (PO)MDPs provide an expressive framework for sequential decision making in a multiagent system. Given their computational complexity, recent research has focused on tractable yet practical subclasses of Dec-POMDPs. We address such a subclass called CDec-POMDP where the collective behavior of a population of agents affects the joint-reward and environment dynamics. Our main contribution is an actor-critic (AC) reinforcement learning method for optimizing CDec-POMDP policies. Vanilla AC has slow convergence for larger problems. To address this, we show how a particular decomposition of the approximate action-value function over agents leads to effective updates, and also derive a new way to train the critic based on local reward signals. Comparisons on a synthetic benchmark and a real world taxi fleet optimization problem show that our new AC approach provides better quality solutions than previous best approaches. Duc Thien Nguyen, Akshat Kumar, Hoong Chuin Lau |
NIPS | 1 |
| 2016 | Approximate Inference Using DC Programming For Collective Graphical ModelsabstractCollective graphical models (CGMs) provide a framework for reasoning about a population of independent and identically distributed individuals when only noisy and aggregate observations are given. Previous approaches for inference in CGMs work on a junction-tree representation, thereby highly limiting their scalability. To remedy this, we show how the Bethe entropy approximation naturally arises for the inference problem in CGMs. We reformulate the resulting optimization problem as a difference-of-convex functions program that can capture different types of CGM noise models. Using the concave-convex procedure, we then develop a scalable message-passing algorithm. Empirically, our approach is highly scalable and accurate for large graphs, more than an order-of-magnitude faster than a generic optimization solver, and is guaranteed to converge unlike the previous message-passing approach NLBP that fails in several loopy graphs. Duc Thien Nguyen, Akshat Kumar, Hoong Chuin Lau, Daniel Sheldon |
AISTATS | 1 |
| 2014 | A Simple Polynomial-Time Randomized Distributed Algorithm for Connected Row Convex ConstraintsabstractIn this paper, we describe a simple randomized algorithm that runs in polynomial time and solves connected row convex (CRC) constraints in distributed settings. CRC constraints generalize many known tractable classes of constraints like 2-SAT and implicational constraints. They can model problems in many domains including temporal reasoning and geometric reasoning, and generally speaking, play the role of ``Gaussians'' in the logical world. Our simple randomized algorithm for solving them in distributed settings, therefore, has a number of important applications. We support our claims through a theoretical analysis and empirical results. T. K. Satish Kumar, Duc Thien Nguyen, William Yeoh 0001, Sven Koenig |
AAAI | 2 |
| 2014 | Decentralized Multi-Agent Reinforcement Learning in Average-Reward Dynamic DCOPsabstractResearchers have introduced the Dynamic Distributed Constraint Optimization Problem (Dynamic DCOP) formulation to model dynamically changing multi-agent coordination problems, where a dynamic DCOP is a sequence of (static canonical) DCOPs, each partially different from the DCOP preceding it. Existing work typically assumes that the problem in each time step is decoupled from the problems in other time steps, which might not hold in some applications. Therefore, in this paper, we make the following contributions: (i) We introduce a new model, called Markovian Dynamic DCOPs (MD-DCOPs), where the DCOP in the next time step is a function of the value assignments in the current time step; (ii) We introduce two distributed reinforcement learning algorithms, the Distributed RVI Q-learning algorithm and the Distributed R-learning algorithm, that balance exploration and exploitation to solve MD-DCOPs in an online manner; and (iii) We empirically evaluate them against an existing multi-arm bandit DCOP algorithm on dynamic DCOPs. Duc Thien Nguyen, William Yeoh 0001, Hoong Chuin Lau, Shlomo Zilberstein, Chongjie Zhang |
AAAI | 1 |
| 2012 | Dynamic Stochastic Orienteering Problems for Risk-Aware Applications
Hoong Chuin Lau, William Yeoh 0001, Pradeep Varakantham, Duc Thien Nguyen, HuaXing Chen |
UAI | 4 |