Thanh Vinh Vo

dblp:222/7878 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning
abstract
Reward shaping is a reinforcement learning technique that addresses the sparse-reward problem by providing frequent, informative feedback. We propose an efficient self-adaptive reward-shaping mechanism that uses success rates derived from historical experiences as shaped rewards. The success rates are sampled from Beta distributions, which evolve from uncertainty to reliability as data accumulates. Initially, shaped rewards are stochastic to encourage exploration, gradually becoming more certain to promote exploitation and maintain a natural balance between exploration and exploitation. We apply Kernel Density Estimation (KDE) with Random Fourier Features (RFF) to derive Beta distributions, providing a computationally efficient solution for continuous and high-dimensional state spaces. Our method, validated on tasks with extremely sparse rewards, improves sample efficiency and convergence stability over relevant baselines.
Haozhe Ma, Zhengding Luo, Thanh Vinh Vo, Kuankuan Sima, Tze-Yun Leong
ICLR3
2025 Catching Two Birds with One Stone: Reward Shaping with Dual Random Networks for Balancing Exploration and Exploitation
abstract
Existing reward shaping techniques for sparse-reward reinforcement learning generally fall into two categories: novelty-based exploration bonuses and significance-based hidden state values. The former promotes exploration but can lead to distraction from task objectives, while the latter facilitates stable convergence but often lacks sufficient early exploration. To address these limitations, we propose Dual Random Networks Distillation (DuRND), a novel reward shaping framework that efficiently balances exploration and exploitation in a unified mechanism. DuRND leverages two lightweight random network modules to simultaneously compute two complementary rewards: a novelty reward to encourage directed exploration and a contribution reward to assess progress toward task completion. With low computational overhead, DuRND excels in high-dimensional environments with challenging sparse rewards, such as Atari, VizDoom, and MiniWorld, outperforming several benchmarks.
Haozhe Ma, Fangling Li, Jing Yu Lim, Zhengding Luo, Thanh Vinh Vo, Tze-Yun Leong
ICML5
2025 Centralized Reward Agent for Knowledge Sharing and Transfer in Multi-Task Reinforcement Learning
abstract
Reward shaping is effective in addressing the sparse-reward challenge in reinforcement learning (RL) by providing immediate feedback through auxiliary, informative rewards. Based on the reward shaping strategy, we propose a novel multi-task reinforcement learning framework that integrates a centralized reward agent (CRA) and multiple distributed policy agents. The CRA functions as a knowledge pool, aimed at distilling knowledge from various tasks and distributing it to individual policy agents to improve learning efficiency. Specifically, the shaped rewards serve as a straightforward metric for encoding knowledge. This framework not only enhances knowledge sharing across established tasks but also adapts to new tasks by transferring meaningful reward signals. We validate the proposed method on both discrete and continuous domains, including the representative Meta-World benchmark, demonstrating its robustness in multi-task sparse-reward settings and its effective transferability to unseen tasks.
Haozhe Ma, Zhengding Luo, Thanh Vinh Vo, Kuankuan Sima, Tze-Yun Leong
NeurIPS3
2025 Federated causal inference from observational data
Thanh Vinh Vo, Tze-Yun Leong
Mach. Learn.1
2024 Reward Shaping for Reinforcement Learning with An Assistant Reward Agent
abstract
Reward shaping is a promising approach to tackle the sparse-reward challenge of reinforcement learning by reconstructing more informative and dense rewards. This paper introduces a novel dual-agent reward shaping framework, composed of two synergistic agents: a policy agent to learn the optimal behavior and a reward agent to generate auxiliary reward signals. The proposed method operates as a self-learning approach, without reliance on expert knowledge or hand-crafted functions. By restructuring the rewards to capture future-oriented information, our framework effectively enhances the sample efficiency and convergence stability. Furthermore, the auxiliary reward signals facilitate the exploration of the environment in the early stage and the exploitation of the policy agent in the late stage, achieving a self-adaptive balance. We evaluate our framework on continuous control tasks with sparse and delayed rewards, demonstrating its robustness and superiority over existing methods.
Haozhe Ma, Kuankuan Sima, Thanh Vinh Vo, Di Fu, Tze-Yun Leong
ICML3
2023 Discovering Low-Dimensional Causal Pathways between Multiple Interacting Neuronal Populations
Evangelos Sigalas, Thanh Vinh Vo, Tze-Yun Leong, Camilo Libedinsky
CogSci2
2023 Transfer Kernel Learning for Multi-Source Transfer Gaussian Process Regression
abstract
Multi-source transfer regression is a practical and challenging problem where capturing the diverse relatedness of different domains is the key of adaptive knowledge transfer. In this paper, we propose an effective way of explicitly modeling the domain relatedness of each domain pair through transfer kernel learning. Specifically, we first discuss the advantages and disadvantages of existing transfer kernels in handling the multi-source transfer regression problem. To cope with the limitations of the existing transfer kernels, we further propose a novel multi-source transfer kernel$k_{ms}$. The proposed$k_{ms}$assigns a learnable parametric coefficient to model the relatedness of each inter-domain pair, and simultaneously regulates the relatedness of the intra-domain pair to be 1. Moreover, to capture the heterogeneous data characteristics of multiple domains,$k_{ms}$exploits different standard kernels for different domain pairs. We further provide a theorem that not only guarantees the positive semi-definiteness of$k_{ms}$but also conveys a semantic interpretation to the learned domain relatedness. Moreover, the theorem can be easily used in the learning of the corresponding transfer Gaussian process model with$k_{ms}$. Extensive empirical studies show the effectiveness of our proposed method on domain relatedness modelling and transfer performance.
Pengfei Wei 0001, Thanh Vinh Vo, Xinghua Qu, Yew-Soon Ong, Zejun Ma 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Adaptive Multi-Source Causal Inference from Observational Data
abstract
We propose a new approach to estimate causal effects from observational data. We leverage multiple data sources which share similar causal mechanisms with the scarce target observations to help infer causal effects in the target domain. The data sources may be available in sequence or some unplanned order. Causal inference can be carried out without prior knowledge of the data discrepancy between the source and target observations. We introduce three levels of knowledge transfer through modelling the outcomes, treatments, and confounders to achieve consistent positive transfer. We incorporate parametric transfer factors to adaptively control the transfer strength, thus achieving a fair and balanced knowledge transfer between the sources and the target. We also empirically show the effectiveness of the proposed method as compared with recent baselines.
Thanh Vinh Vo, Pengfei Wei 0001, Trong Nghia Hoang, Tze-Yun Leong
CIKM1
2022 An Adaptive Kernel Approach to Federated Learning of Heterogeneous Causal Effects
abstract
We propose a new causal inference framework to learn causal effects from multiple, decentralized data sources in a federated setting. We introduce an adaptive transfer algorithm that learns the similarities among the data sources by utilizing Random Fourier Features to disentangle the loss function into multiple components, each of which is associated with a data source. The data sources may have different distributions; the causal effects are independently and systematically incorporated. The proposed method estimates the similarities among the sources through transfer coefficients, and hence requiring no prior information about the similarity measures. The heterogeneous causal effects can be estimated with no sharing of the raw training data among the sources, thus minimizing the risk of privacy leak. We also provide minimax lower bounds to assess the quality of the parameters learned from the disparate sources. The proposed method is empirically shown to outperform the baselines on decentralized data sources with dissimilar distributions.
Thanh Vinh Vo, Arnab Bhattacharyya 0001, Tze-Yun Leong
NeurIPS1
2022 Bayesian federated estimation of causal effects from observational data
abstract
We propose a Bayesian framework for estimating causal effects from federated observational data sources. Bayesian causal inference is an important approach to learning the distribution of the causal estimands and understanding the uncertainty of causal effects. Our framework estimates the posterior distributions of the causal effects to compute the higher-order statistics that capture the uncertainty. We integrate local causal effects from different data sources without centralizing them. We then estimate the treatment effects from observational data using a non-parametric reformulation of the classical potential outcomes framework. We model the potential outcomes as a random function distributed by Gaussian processes, with defining parameters that can be efficiently learned from multiple data sources. Our method avoids exchanging raw data among the sources, thus contributing towards privacy-preserving causal learning. The promise of our approach is demonstrated through a set of simulated and real-world examples.
Thanh Vinh Vo, Trong Nghia Hoang, Tze-Yun Leong
UAI1
2021 Causal Modeling with Stochastic Confounders
abstract
This work extends causal inference in temporal models with stochastic confounders. We propose a new approach to variational estimation of causal inference based on a representer theorem with a random input space. We estimate causal effects involving latent confounders that may be interdependent and time-varying from sequential, repeated measurements in an observational study. Our approach extends current work that assumes independent, non-temporal latent confounders with potentially biased estimators. We introduce a simple yet elegant algorithm without parametric specification on model components. Our method avoids the need for expensive and careful parameterization in deploying complex models, such as deep neural networks in existing approaches, for causal inference and analysis. We demonstrate the effectiveness of our approach on various benchmark temporal datasets.
Thanh Vinh Vo, Pengfei Wei 0001, Wicher Bergsma, Tze-Yun Leong
AISTATS1
2018 Z-Transforms and its Inference on Partially Observable Point Processes
abstract
This paper proposes an inference framework based on the Z-transform for a specific class of non-homogeneous point processes. This framework gives an alternative method to maximum likelihood estimation which is omnipresent in the field of point processes. The inference strategy is to couple or match the theoretical Z-transform with its empirical counterpart from the observed samples. This procedure fully characterizes the distribution of the point process since there exists a one-to-one mapping with the Z-transform. We illustrate how to use the methodology to estimate a point process whose intensity is driven by a general neural network.
Thanh Vinh Vo, Kar Wai Lim, Harold Soh
IJCAI2
2018 Generation meets recommendation: proposing novel items for groups of users
abstract
Consider a movie studio aiming to produce a set of new movies for summer release: What types of movies it should produce? Who would the movies appeal to? How many movies should it make? Similar issues are encountered by a variety of organizations, e.g., mobile-phone manufacturers and online magazines, who have to create new (non-existent) items to satisfy groups of users with different preferences. In this paper, we present a joint problem formalization of these interrelated issues, and propose generative methods that address these questions simultaneously. Specifically, we leverage on the latent space obtained by training a deep generative model---the Variational Autoencoder (VAE)---via a loss function that incorporates both rating performance and item reconstruction terms. We use a greedy search algorithm that utilize this learned latent space to jointly obtain K plausible new items, and user groups that would find the items appealing. An evaluation of our methods on a synthetic dataset indicates that our approach is able to generate novel items similar to highly-desirable unobserved items. As case studies on real-world data, we applied our method on the MART abstract art and Movielens Tag Genome datasets, which resulted in promising results: small and diverse sets of novel items.
Thanh Vinh Vo, Harold Soh
RecSys1