Ngo Anh Vien

dblp:87/439 · DBLP profile ↗
← Back
43ranked-venue papers
16as first author
15since 2021 · last 2025
0000-0001-9646-267XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 15 first-author · 14 since 2021Systems, architecture and hardware · 10 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorComputer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Geometry-aware RL for Manipulation of Varying Shapes and Deformable Objects
abstract
Manipulating objects with varying geometries and deformable objects is a major challenge in robotics. Tasks such as insertion with different objects or cloth hanging require precise control and effective modelling of complex dynamics. In this work, we frame this problem through the lens of a heterogeneous graph that comprises smaller sub-graphs, such as actuators and objects, accompanied by different edge types describing their interactions. This graph representation serves as a unified structure for both rigid and deformable objects tasks, and can be extended further to tasks comprising multiple actuators. To evaluate this setup, we present a novel and challenging reinforcement learning benchmark, including rigid insertion of diverse objects, as well as rope and cloth manipulation with multiple end-effectors. These tasks present a large search space, as both the initial and target configurations are uniformly sampled in 3D space. To address this issue, we propose a novel graph-based policy model, dubbed Heterogeneous Equivariant Policy (HEPi), utilizing $SE(3)$ equivariant message passing networks as the main backbone to exploit the geometric symmetry. In addition, by modeling explicit heterogeneity, HEPi can outperform Transformer-based and non-heterogeneous equivariant policies in terms of average returns, sample efficiency, and generalization to unseen objects. Our project page is available at https://thobotics.github.io/hepi.
Tai Hoang, Philipp Becker, Ngo Anh Vien, Gerhard Neumann
ICLR4
2025 Efficient Off-Policy Learning for High-Dimensional Action Spaces
abstract
Existing off-policy reinforcement learning algorithms often rely on an explicit state-action-value function representation, which can be problematic in high-dimensional action spaces due to the curse of dimensionality. This reliance results in data inefficiency as maintaining a state-action-value function in such spaces is challenging. We present an efficient approach that utilizes only a state-value function as the critic for off-policy deep reinforcement learning. This approach, which we refer to as Vlearn, effectively circumvents the limitations of existing methods by eliminating the necessity for an explicit state-action-value function. To this end, we leverage a weighted importance sampling loss for learning deep value functions from off-policy data. While this is common for linear methods, it has not been combined with deep value function networks. This transfer to deep methods is not straightforward and requires novel design choices such as robust policy updates, twin value function networks to avoid an optimization bias, and importance weight clipping. We also present a novel analysis of the variance of our estimate compared to commonly used importance sampling estimators such as V-trace. Our approach improves sample complexity as well as final performance and ensures consistent and robust performance across various benchmark tasks. Eliminating the state-action-value function in Vlearn facilitates a streamlined learning process, yielding high-return agents.
Fabian Otto, Philipp Becker, Ngo Anh Vien, Gerhard Neumann
ICLR3
2025 How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?
abstract
Recent advances in 3D point cloud transformers have led to state-of-the-art results in tasks such as semantic segmentation and reconstruction. However, these models typically rely on dense token representations, incurring high computational and memory costs during training and inference. In this work, we present the finding that tokens are remarkably redundant, leading to substantial inefficiency. We introduce \textbf{GitMerge3D}, a \textbf{g}lobally \textbf{i}nformed graph \textbf{t}oken \textbf{merging} method that can reduce the token count by up to 90–95\% while maintaining competitive performance. This finding challenges the prevailing assumption that more tokens inherently yield better performance and highlights that many current models are over-tokenized and under-optimized for scalability. We validate our method across multiple 3D vision tasks and show consistent improvements in computational efficiency. This work is the first to assess redundancy in large-scale 3D transformer models, providing insights into the development of more efficient 3D foundation architectures. Our code and checkpoints are publicly available at \href{https://gitmerge3d.github.io/}{https://gitmerge3d.github.io}.
Duy M. H. Nguyen, Hoai-Chau Tran, Michael Barz, Khoa D. Doan, Roger Wattenhofer, Ngo Anh Vien, Mathias Niepert, Daniel Sonntag, Paul Swoboda
NeurIPS7
2024 Pseudo Labeling and Contextual Curriculum Learning for Online Grasp Learning in Robotic Bin Picking
abstract
The prevailing grasp prediction methods predominantly rely on offline learning, overlooking the dynamic grasp learning that occurs during real-time adaptation to novel picking scenarios. These scenarios may involve previously unseen objects, variations in camera perspectives, and bin configurations, among other factors. In this paper, we introduce a novel approach, SSL-ConvSAC, that combines semi-supervised learning and reinforcement learning for online grasp learning. By treating pixels with reward feedback as labeled data and others as unlabeled, it efficiently exploits unlabeled data to enhance learning. In addition, we address the imbalance between labeled and unlabeled data by proposing a contextual curriculum-based method. We ablate the proposed approach on real-world evaluation data and demonstrate promise for improving online grasp learning on bin picking tasks using a physical 7-DoF Franka Emika robot arm with a suction gripper. Video: https://youtu.be/OAro5pg8I9U
Philipp Schillinger, Miroslav Gabriel, Alexander Qualmann, Ngo Anh Vien
ICRA5
2024 Efficient End-to-End Detection of 6-DoF Grasps for Robotic Bin Picking
abstract
Bin picking is an important building block for many robotic systems, in logistics, production or in household use-cases. In recent years, machine learning methods for the prediction of 6-DoF grasps on diverse and unknown objects have shown promising progress. However, existing approaches only consider a single ground truth grasp orientation at a grasp location during training and therefore can only predict limited grasp orientations which leads to a reduced number of feasible grasps in bin picking with restricted reachability. In this paper, we propose a novel approach for learning dense and diverse 6-DoF grasps for parallel-jaw grippers in robotic bin picking. We introduce a parameterized grasp distribution model based on Power-Spherical distributions that enables a training based on all possible ground truth samples. Thereby, we also consider the grasp uncertainty enhancing the model’s robustness to noisy inputs. As a result, given a single top-down view depth image, our model can generate diverse grasps with multiple collision-free grasp orientations. Experimental evaluations in simulation and on a real robotic bin picking setup demonstrate the model’s ability to generalize across various object categories achieving an object clearing rate of around 90% in simulation and real-world experiments. We also outperform state of the art approaches. Moreover, the proposed approach exhibits its usability in real robot experiments without any refinement steps, even when only trained on a synthetic dataset, due to the probabilistic grasp distribution modeling.
Alexander Qualmann, Zehao Yu 0002, Miroslav Gabriel, Philipp Schillinger, Markus Spies, Ngo Anh Vien, Andreas Geiger 0001
ICRA7
2024 Uncertainty-driven Exploration Strategies for Online Grasp Learning
abstract
Existing grasp prediction approaches are mostly based on offline learning, while, ignoring the exploratory grasp learning during online adaptation to new picking scenarios, i.e., objects that are unseen or out-of-domain (OOD), camera and bin settings, etc. In this paper, we present an uncertainty-based approach for online learning of grasp predictions for robotic bin picking. Specifically, the online learning algorithm with an effective exploration strategy can significantly improve its adaptation performance to unseen environment settings. To this end, we first propose to formulate online grasp learning as an RL problem that will allow us to adapt both grasp reward prediction and grasp poses. We propose various uncertainty estimation schemes based on Bayesian uncertainty quantification and distributional ensembles. We carry out evaluations on real-world bin picking scenes of varying difficulty. The objects in the bin have various challenging physical and perceptual characteristics that can be characterized by semi- or total transparency, and irregular or curved surfaces. The results of our experiments demonstrate a notable improvement of grasp performance in comparison to conventional online learning methods which incorporate only naive exploration strategies. Video: https://youtu.be/fPKOrjC2QrU
Yitian Shi, Philipp Schillinger, Miroslav Gabriel, Alexander Qualmann, Zohar Feldman, Hanna Carolin Ziesche, Ngo Anh Vien
ICRA7
2023 Model-Free Grasping with Multi-Suction Cup Grippers for Robotic Bin Picking
abstract
This paper presents a novel method for model-free prediction of grasp poses for suction grippers with multiple suction cups. Our approach is agnostic to the design of the gripper and does not require gripper-specific training data. In particular, we propose a two-step approach, where first, a neural network predicts pixel-wise grasp quality for an input image to indicate areas that are generally graspable. Second, an optimization step determines the optimal gripper selection and corresponding grasp poses based on configured gripper layouts and activation schemes. In addition, we introduce a method for automated labeling for supervised training of the grasp quality network. Experimental evaluations on a real-world industrial application with bin picking scenes of varying difficulty demonstrate the effectiveness of our method.
Philipp Schillinger, Miroslav Gabriel, Alexander Kuss, Hanna Carolin Ziesche, Ngo Anh Vien
IROS5
2022 What Matters For Meta-Learning Vision Regression Tasks?
abstract
Meta-learning is widely used in few-shot classification and function regression due to its ability to quickly adapt to unseen tasks. However, it has not yet been well explored on regression tasks with high dimensional inputs such as images. This paper makes two main contributions that help understand this barely explored area. First, we design two new types of cross-category level vision regression tasks, namely object discovery and pose estimation of unprecedented complexity in the meta-learning domain for computer vision. To this end, we (i) exhaustively evaluate common meta-learning techniques on these tasks, and (ii) quantitatively analyze the effect of various deep learning techniques commonly used in recent meta-learning algorithms in order to strengthen the generalization capability: data augmentation, domain randomization, task augmentation and meta-regularization. Finally, we (iii) provide some insights and practical recommendations for training meta-learning algorithms on vision regression tasks. Second, we propose the addition of functional contrastive learning (FCL) over the task representations in Conditional Neural Processes (CNPs) and train in an end-to-end fashion. The experimental results show that the results of prior work are misleading as a consequence of a poor choice of the loss function as well as too small meta-training sets. Specifically, we find that CNPs outperform MAML on most tasks without fine-tuning. Furthermore, we observe that naive task augmentation without a tailored design results in underfitting.
Ning Gao 0005, Hanna Carolin Ziesche, Ngo Anh Vien, Michael Volpp, Gerhard Neumann
CVPR3
2022 FusionVAE: A Deep Hierarchical Variational Autoencoder for RGB Image Fusion
Fabian Duffhauss, Ngo Anh Vien, Hanna Carolin Ziesche, Gerhard Neumann
ECCV (39)2
2022 A Hybrid Approach for Learning to Shift and Grasp with Elaborate Motion Primitives
abstract
Many possible fields of application of robots in real world settings hinge on the ability of robots to grasp objects. As a result, robot grasping has been an active field of research for many years. With our publication we contribute to the endeavor of enabling robots to grasp, with a particular focus on bin picking applications. Bin picking is especially challenging due to the often cluttered and unstructured arrangement of objects and the often limited graspability of objects by simple top down grasps. To tackle these challenges, we propose a fully self-supervised reinforcement learning approach based on a hybrid discrete-continuous adaptation of soft actor-critic (SAC). We employ parametrized motion primitives for pushing and grasping movements in order to enable a flexibly adaptable behavior to the difficult setups we consider. Furthermore, we use data augmentation to increase sample efficiency. We demonstrate our proposed method on challenging picking scenarios in which planar grasp learning or action discretization methods would face a lot of difficulties.
Zohar Feldman, Hanna Carolin Ziesche, Ngo Anh Vien, Dotan Di Castro
ICRA3
2022 Hierarchical Policy Learning for Mechanical Search
abstract
Retrieving objects from clutters is a complex task, which requires multiple interactions with the environment until the target object can be extracted. These interactions involve executing action primitives like grasping or pushing as well as setting priorities for the objects to manipulate and the actions to execute. Mechanical Search (MS) [1] is a framework for object retrieval, which uses a heuristic algorithm for pushing and rule-based algorithms for high-level planning. While rule-based policies profit from human intuition in how they work, they usually perform sub-optimally in many cases. Deep reinforcement learning (RL) has shown great performance in complex tasks such as taking decisions through evaluating pixels, which makes it suitable for training policies in the context of object-retrieval. In this work, we first formulate the MS problem in a principled formulation as a hierarchical POMDP. Based on this formulation, we propose a hierarchical policy learning approach for the MS problem. For demonstration, we present two main parameterized sub-policies: a push policy and an action selection policy. When integrated into the hierarchical POMDP's policy, our proposed sub-policies increase the success rate of retrieving the target object from less than 32% to nearly 80%, while reducing the computation time for push actions from multiple seconds to less than 10 milliseconds.
Oussama Zenkri, Ngo Anh Vien, Gerhard Neumann
ICRA2
2021 Differentiable Trust Region Layers for Deep Reinforcement Learning
Fabian Otto, Philipp Becker, Ngo Anh Vien, Hanna Carolin Ziesche, Gerhard Neumann
ICLR3
2021 Residual Feedback Learning for Contact-Rich Manipulation Tasks with Uncertainty
abstract
While classic control theory offers state of the art solutions in many problem scenarios, it is often desired to improve beyond the structure of such solutions and surpass their limitations. To this end, residual policy learning (RPL) offers a formulation to improve existing controllers with reinforcement learning (RL) by learning an additive "residual" to the output of a given controller. However, the applicability of such an approach highly depends on the structure of the controller. Often, internal feedback signals of the controller limit an RL algorithm to adequately change the policy and, hence, learn the task. We propose a new formulation that addresses these limitations by also modifying the feedback signals to the controller with an RL policy and show superior performance of our approach on a contact-rich peg-insertion task under position and orientation uncertainty. In addition, we use a recent Cartesian impedance control architecture as the control framework which can be available to us as a black-box while assuming no knowledge about its input/output structure, and show the difficulties of standard RPL. Furthermore, we introduce an adaptive curriculum for the given task to gradually increase the task difficulty in terms of position and orientation uncertainty. A video showing the results can be found at https://youtu.be/SAZm_Krze7U.
Alireza Ranjbar, Ngo Anh Vien, Hanna Carolin Ziesche, Joschka Boedecker, Gerhard Neumann
IROS2
2021 Non-local Graph Convolutional Network for joint Activity Recognition and Motion Prediction
abstract
3D skeleton-based motion prediction and activity recognition are two interwoven tasks in human behaviour analysis. In this work, we propose a motion context modeling methodology that provides a new way to combine the advantages of both graph convolutional neural networks and recurrent neural networks for joint human motion prediction and activity recognition. Our approach is based on using an LSTM encoder-decoder and a non-local feature extraction attention mechanism to model the spatial correlation of human skeleton data and temporal correlation among motion frames. The proposed network can easily include two output branches, one for Activity Recognition and one for Future Motion Prediction, which can be jointly trained for enhanced performance. Experimental results on Human 3.6M, CMU Mocap and NTU RGB-D datasets show that our proposed approach provides the best prediction capability among baseline LSTM-based methods, while achieving comparable performance to other state-of-the-art methods.
Dianhao Zhang, Ngo Anh Vien, Mien Van, Seán F. McLoone
IROS2
2021 Deep Learning-Aided Multicarrier Systems
abstract
This paper proposes a deep learning (DL)-aided multicarrier (MC) system operating on fading channels, where both modulation and demodulation blocks are modeled by deep neural networks (DNNs), regarded as the encoder and decoder of an autoencoder (AE) architecture, respectively. Unlike existing AE-based systems, which incorporate domain knowledge of a channel equalizer to suppress the effects of wireless channels, the proposed scheme, termed as MC-AE, directly feeds the decoder with the channel state information and received signal, which are then processed in a fully data-driven manner. This new approach enables MC-AE to jointly learn the encoder and decoder to optimize the diversity and coding gains over fading channels. In particular, the block error rate of MC-AE is analyzed to show its higher performance gains than existing hand-crafted baselines, such as various recent index modulation-based MC schemes. We then extend MC-AE to multiuser scenarios, wherein the resultant system is termed as MU-MC-AE. Accordingly, two novel DNN structures for uplink and downlink MU-MC-AE transmissions are proposed, along with a novel cost function that ensures a fast training convergence and fairness among users. Finally, simulation results are provided to show the superiority of the proposed DL-based schemes over current baselines, in terms of both the error performance and receiver complexity.
Thien Van Luong, Youngwook Ko, Michail Matthaiou, Ngo Anh Vien, Minh-Tuan Le, Vu-Duc Ngo
IEEE Trans. Wirel. Commun.4
2020 Fast Analysis and Prediction in Large Scale Virtual Machines Resource Utilisation
abstract
Most Cloud providers running Virtual Machines (VMs) have a constant goal of preventing downtime, increas- ing performance and power management among others. The most effective way to achieve these goals is to be proactive by predicting the behaviours of the VMs. Analysing VMs is important, as it can help cloud providers gain insights to understand the needs of their customers, predict their demands, and optimise the use of resources. To manage the resources in the cloud efficiently, and to ensure the performance of cloud ser- vices, it is crucial to predict the behaviour of VMs accurately. This will also help the cloud provider improve VM placement, scheduling, consolidation, power management, etc. In this paper, we propose a framework for fast analysis and prediction in large scale VM CPU utilisation. We use a novel approach both in terms of the algorithms employed for prediction and in terms of the tools used to run these algorithms with a large dataset to deliver a solid VM CPU utilisation predictor. We processed over two million VMs from Microsoft Azure VM traces and filter out the VMs with complete one month of data which amount to 28,858VMs. The filtered VMs were subsequently used for prediction. Our Statistical analysis reveals that 94% of these VMs are predictable. Furthermore, we investigate the patterns and behaviours of those VMs and realised that most VMs have one or several spikes of which the majority are not seasonal. For all the 28,858VMs analysed and forecasted, we accurately predicted 17,523 (61%) VMs based on their CPU. We use Apache Spark for parallel and distributed processing to achieve fast processing. In terms of fast processing (execution time), on average, each VM is analysed and predicted within three seconds.
Sakil Barbhuiya, Peter Kilpatrick, Ngo Anh Vien, Dimitrios S. Nikolopoulos
CLOSER4
2020 Graph-Based Motion Planning Networks
Tai Hoang, Ngo Anh Vien
ECML/PKDD (2)2
2020 Asynchronous framework with Reptile+ algorithm to meta learn partially observable Markov decision process
Dang Quang Nguyen, Ngo Anh Vien, Viet-Hung Dang, TaeChoong Chung
Appl. Intell.2
2020 Deep Energy Autoencoder for Noncoherent Multicarrier MU-SIMO Systems
abstract
We propose a novel deep energy autoencoder (EA) for noncoherent multicarrier multiuser single-input multipleoutput (MU-SIMO) systems under fading channels.In particular, a single-user noncoherent EA-based (NC-EA) system, based on the multicarrier SIMO framework, is first proposed, where both the transmitter and receiver are represented by deep neural networks (DNNs), known as the encoder and decoder of an EA.Unlike existing systems, the decoder of the NC-EA is fed only with the energy combined from all receive antennas, while its encoder outputs a real-valued vector whose elements stand for the subcarrier power levels.Using the NC-EA, we then develop two novel DNN structures for both uplink and downlink NC-EA multiple access (NC-EAMA) schemes, based on the multicarrier MU-SIMO framework.Note that NC-EAMA allows multiple users to share the same sub-carriers, thus enables to achieve higher performance gains than noncoherent orthogonal counterparts.By properly training, the proposed NC-EA and NC-EAMA can efficiently recover the transmitted data without any channel state information estimation.Simulation results clearly show the superiority of our schemes in terms of reliability, flexibility and complexity over baseline schemes.
Thien Van Luong, Youngwook Ko, Ngo Anh Vien, Michail Matthaiou, Hien Quoc Ngo
IEEE Trans. Wirel. Commun.3
2018 Bayesian Functional Optimization
abstract
Bayesian optimization (BayesOpt) is a derivative-free approach for sequentially optimizing stochastic black-box functions. Standard BayesOpt, which has shown many successes in machine learning applications, assumes a finite dimensional domain which often is a parametric space. The parameter space is defined by the features used in the function approximations which are often selected manually. Therefore, the performance of BayesOpt inevitably depends on the quality of chosen features. This paper proposes a new Bayesian optimization framework that is able to optimize directly on the domain of function spaces. The resulting framework, Bayesian Functional Optimization (BFO), not only extends the application domains of BayesOpt to functional optimization problems but also relaxes the performance dependency on the chosen parameter space. We model the domain of functions as a reproducing kernel Hilbert space (RKHS), and use the notion of Gaussian processes on a real separable Hilbert space. As a result, we are able to define traditional improvement-based (PI and EI) and optimistic acquisition functions (UCB) as functionals. We propose to optimize the acquisition functionals using analytic functional gradients that are also proved to be functions in a RKHS. We evaluate BFO in three typical functional optimization tasks: i) a synthetic functional optimization problem, ii) optimizing activation functions for a multi-layer perceptron neural network, and iii) a reinforcement learning task whose policies are modeled in RKHS.
Ngo Anh Vien, Heiko Zimmermann, Marc Toussaint
AAAI1
2018 Scalable and Interpretable One-Class SVMs with Deep Learning and Random Fourier Features
Minh-Nghia Nguyen, Ngo Anh Vien
ECML/PKDD (1)2
2017 A Covariance Matrix Adaptation Evolution Strategy for Direct Policy Search in Reproducing Kernel Hilbert Space
abstract
The covariance matrix adaptation evolution strategy (CMA-ES) is an efficient derivative-free optimization algorithm. It optimizes a black-box objective function over a well defined parameter space. In some problems, such parameter spaces are defined using function approximation in which feature functions are manually defined. Therefore, the performance of those techniques strongly depends on the quality of chosen features. Hence, enabling CMA-ES to optimize on a more complex and general function class of the objective has long been desired. Specifically, we consider modeling the input space for black-box optimization in reproducing kernel Hilbert spaces (RKHS). This modeling leads to a functional optimization problem whose domain is a function space that enables us to optimize in a very rich function class. In addition, we propose CMA-ES-RKHS, a generalized CMA-ES framework, that performs black-box functional optimization in RKHS. A search distribution, represented as a Gaussian process, is adapted by updating both its mean function and covariance operator. Adaptive representation of the mean function and the covariance operator is achieved by resorting to sparsification. CMA-ES-RKHS is evaluated on two simple functional optimization problems and two bench-mark reinforcement learning (RL) domains. For an application in RL, we model policies for MDPs in RKHS and transform a cumulative return objective as a functional of RKHS policies, which can be optimized via CMA-ES-RKHS. This formulation results in a black-box functional policy search framework.
Ngo Anh Vien, Viet-Hung Dang, TaeChoong Chung
ACML1
2016 Relational activity processes for modeling concurrent cooperation
abstract
In human-robot collaboration, multi-agent domains, or single-robot manipulation with multiple end-effectors, the activities of the involved parties are naturally concurrent. Such domains are also naturally relational as they involve objects, multiple agents, and models should generalize over objects and agents. We propose a novel formalization of relational concurrent activity processes that allows us to transfer methods from standard relational MDPs, such as Monte-Carlo planning and learning from demonstration, to concurrent cooperation domains. We formally compare the formulation to previous propositional models of concurrent decision making and demonstrate planning and learning from demonstration methods on a real-world human-robot assembly task.
Marc Toussaint, Thibaut Munzer, Yoan Mollard, Li Yang Wu, Ngo Anh Vien, Manuel Lopes 0001
ICRA5
2016 Policy Search in Reproducing Kernel Hilbert Space
Ngo Anh Vien, Peter Englert, Marc Toussaint
IJCAI1
2016 Bayes-adaptive hierarchical MDPs
Ngo Anh Vien, SeungGwan Lee, TaeChoong Chung
Appl. Intell.1
2015 Hierarchical Monte-Carlo Planning
abstract
Monte-Carlo Tree Search, especially UCT and its POMDP version POMCP, have demonstrated excellent performanceon many problems. However, to efficiently scale to large domains one should also exploit hierarchical structure if present. In such hierarchical domains, finding rewarded states typically requires to search deeply; covering enough such informative states very far from the root becomes computationally expensive in flat non-hierarchical search approaches. We propose novel, scalable MCTS methods which integrate atask hierarchy into the MCTS framework, specifically lead-ing to hierarchical versions of both, UCT and POMCP. The new method does not need to estimate probabilistic models of each subtask, it instead computes subtask policies purely sample-based. We evaluate the hierarchical MCTS methods on various settings such as a hierarchical MDP, a Bayesian model-based hierarchical RL problem, and a large hierarchical POMDP.
Ngo Anh Vien, Marc Toussaint
AAAI1
2015 POMDP manipulation via trajectory optimization
abstract
Efficient object manipulation based only on force feedback typically requires a plan of actively contact-seeking actions to reduce uncertainty over the true environmental model. In principle, that problem could be formulated as a full partially observable Markov decision process (POMDP) whose observations are sensed forces indicating the presence/absence of contacts with objects. Such a naive application leads to a very large POMDP with high-dimensional continuous state, action and observation spaces. Solving such large POMDPs is practically prohibitive. In other words, we are facing three challenging problems: 1) uncertainty over discontinuous contacts with objects; 2) high-dimensional continuous spaces; 3) optimization for not only trajectory cost but also execution time. As trajectory optimization is a powerful model-based method for motion generation, it can handle the last two issues effectively by computing locally optimal trajectories. This paper aims to integrate advantages of trajectory optimization into existing POMDP solvers. The full POMDP formulation is solved using sample-based approaches, where each sampled model is quickly evaluated via trajectory optimization instead of simulating a large number of rollouts. To further accelerate the solver, we propose to integrate temporal abstraction, i.e. macro actions or temporal actions, into the POMDP model. We demonstrate the proposed method on a simulated 7 DoF KUKA arm and a physical Willow Garage PR2 platform. The results show that our proposed method could effectively seek contacts in complex scenarios, and achieve near-optimal performance of path planing.
Ngo Anh Vien, Marc Toussaint
IROS1
2014 Model-Based Relational RL When Object Existence is Partially Observable
abstract
We consider learning and planning in relational MDPs when object existence is uncertain and new objects may appear or disappear depending on previous actions or properties of other objects. Optimal policies actively need to discover objects to achieve a goal; planning in such domains in general amounts to a POMDP problem, where the belief is about the existence and properties of potential not-yet-discovered objects. We propose a computationally efficient extension of model-based relational RL methods that approximates these beliefs using discrete uncertainty predicates. In this formulation the belief update is learned using probabilistic rules and planning in the approximated belief space can be achieved using an extension of existing planners. We prove that the learned belief update rules encode an approximation of the exact belief updates of a POMDP formulation and demonstrate experimentally that the proposed approach successfully learns a set of relational rules appropriate to solve such problems.
Ngo Anh Vien, Marc Toussaint
ICML1
2014 Approximate planning for bayesian hierarchical reinforcement learning
Ngo Anh Vien, Hung Quoc Ngo 0001, Sungyoung Lee 0001, TaeChoong Chung
Appl. Intell.1
2014 Efficient Interactive Multiclass Learning from Binary Feedback
abstract
We introduce a novel algorithm called upper confidence - weighted learning (UCWL) for online multiclass learning from binary feedback (e.g., feedback that indicates whether the prediction was right or wrong). UCWL combines the upper confidence bound (UCB) framework with the soft confidence-weighted (SCW) online learning scheme. In UCB, each instance is classified using both score and uncertainty. For a given instance in the sequence, the algorithm might guess its class label primarily to reduce the class uncertainty. This is a form of informed exploration, which enables the performance to improve with lower sample complexity compared to the case without exploration. Combining UCB with SCW leads to the ability to deal well with noisy and nonseparable data, and state-of-the-art performance is achieved without increasing the computational cost. A potential application setting is human-robot interaction (HRI), where the robot is learning to classify some set of inputs while the human teaches it by providing only binary feedback—or sometimes even the wrong answer entirely. Experimental results in the HRI setting and with two benchmark datasets from other settings show that UCWL outperforms other state-of-the-art algorithms in the online binary feedback setting—and surprisingly even sometimes outperforms state-of-the-art algorithms that get full feedback (e.g., the true class label), whereas UCWL gets only binary feedback on the same data sequence.
Hung Quoc Ngo 0001, Matthew D. Luciw, Jawad Nagi, Alexander Förster, Jürgen Schmidhuber, Ngo Anh Vien
ACM Trans. Interact. Intell. Syst.6
2013 Upper Confidence Weighted Learning for Efficient Exploration in Multiclass Prediction with Binary Feedback
Hung Quoc Ngo 0001, Matthew D. Luciw, Ngo Anh Vien, Jürgen Schmidhuber
IJCAI3
2013 Learning via human feedback in continuous state and action spaces
Ngo Anh Vien, Wolfgang Ertel, TaeChoong Chung
Appl. Intell.1
2013 Monte-Carlo tree search for Bayesian reinforcement learning
Ngo Anh Vien, Wolfgang Ertel, Viet-Hung Dang, TaeChoong Chung
Appl. Intell.1
2012 Monte Carlo Tree Search for Bayesian Reinforcement Learning
abstract
Bayesian model-based reinforcement learning can be formulated as a partially observable Markov decision process (POMDP) to provide a principled framework for optimally balancing exploitation and exploration. Then, a POMDP solver can be used to solve the problem. If the prior distribution over the environment's dynamics is a product of Dirichlet distributions, the POMDP's optimal value function can be represented using a set of multivariate polynomials. Unfortunately, the size of the polynomials grows exponentially with the problem horizon. In this paper, we examine the use of an online Monte-Carlo tree search (MCTS) algorithm for large POMDPs, to solve the Bayesian reinforcement learning problem online. We will show that such an algorithm successfully searches for a near-optimal policy. In addition, we examine the use of a parameter tying method to keep the model search space small, and propose the use of nested mixture of tied models to increase robustness of the method when our prior information does not allow us to specify the structure of tied models exactly. Experiments show that the proposed methods substantially improve scalability of current Bayesian reinforcement learning methods.
Ngo Anh Vien, Wolfgang Ertel
ICMLA (1)1
2011 Nomogram Visualization for Ranking Support Vector Machine
Nguyen Thi Thanh Thuy, Nguyen Thi Ngoc Vinh, Ngo Anh Vien
ISNN (2)3
2011 Hessian matrix distribution for Bayesian policy gradient reinforcement learning
Ngo Anh Vien, Hwanjo Yu, TaeChoong Chung
Inf. Sci.1
2010 Monte Carlo Value Iteration for Continuous-State POMDPs
Haoyu Bai, David Hsu, Wee Sun Lee, Ngo Anh Vien
WAFR4
2009 VRIFA: a nonlinear SVM visualization tool using nomogram and localized radial basis function (LRBF) kernels
abstract
Prediction problems are prevalent in medical domains. For example, computer-aided diagnosis or prognosis is a key component in a CDSS (Clinical Decision Support System). SVMs, especially SVMs with nonlinear kernels such RBF kernels, have shown superior accuracy in prediction problems. However, they are not favorably used by physicians for medical prediction problems because nonlinear SVMs are difficult to visualize, thus it is hard to provide intuitive interpretation of prediction results to physicians. Nomogram was proposed to visualize SVM classification models. However, it cannot visualize nonlinear SVM models. Localized RBF (LRBF) kernel was proposed which shows comparable accuracy as the RBF kernel while the LRBF kernel is easier to interpret since it can be linearly decomposed. This paper presents a new tool named VRIFA, which integrates the nomogram and LRBF kernel to provide users with an interactive visualization of nonlinear SVM models. VRIFA graphically exposes the internal structure of nonlinear SVM models showing the effect of each feature, the magnitude of the effect, and the change at the prediction output. VRIFA also performs nomogram-based feature selection while training a model in order to remove noise or redundant features and improve the prediction accuracy. The tool has been used by biomedical researchers for computer-aided diagnosis and risk factor analysis for diseases. VRIFA is accessible at http://dm.postech.ac.kr/vrifa .
Ngo Anh Vien, Nguyen Hoang Viet, TaeChoong Chung, Hwanjo Yu, Sungchul Kim, Baek Hwan Cho
CIKM1
2009 Probabilistic Ranking Support Vector Machine
Nguyen Thi Thanh Thuy, Ngo Anh Vien, Nguyen Hoang Viet, TaeChoong Chung
ISNN (2)2
2008 Obstacle Avoidance Path Planning for Mobile Robot Based on Multi Colony Ant Algorithm
abstract
The task of planning trajectories for a mobile robot has received considerable attention in the research literature. The problem involves computing a collision-free path between a start point and a target point in environment of known obstacles. In this paper, we study an obstacle avoidance path planning problem using multi ant colony system, in which several colonies of ants cooperate in finding good solution by exchanging good information. In the simulation, we experimentally investigate the behaviour of multi colony ant algorithm with different kinds of information among the colonies. At last we will compare the behaviour of different number of colonies with a multi start single colony ant algorithm to show the good improvement.
Nguyen Hoang Viet, Ngo Anh Vien, SeungGwan Lee, TaeChoong Chung
ACHI2
2008 Policy Gradient Semi-markov Decision Process
abstract
This paper proposes a simulation-based algorithm for optimizing the average reward in a parameterized continuous-time, finite-state semi-Markov decision process (SMDP). Our contributions are twofold: First, we compute the approximate gradient of the average reward with respect to the parameters in SMDP controlled by parameterized stochastic policies. Then stochastic gradient ascent method is used to adjust the parameters in order to optimize the average reward. Second, we present a simulation-based algorithm to estimate the approximate average gradient of the average reward (GSMDP), using only single sample path of the underlying Markov chain. We prove the almost sure convergence of this estimate to the true gradient of the average reward when the number of iterations goes to infinity.
Ngo Anh Vien, TaeChoong Chung
ICTAI (2)1
2007 Natural Gradient Policy for Average Cost SMDP Problem
abstract
Semi-markov decision processes (SMDP) are continuous time generalizations of discrete time Markov Decision Process. A number of value and policy iteration algorithms have been developed for the solution of SMDP problem. But solving SMDP problem requires prior knowledge of the deterministic kernels, and suffers from the curse of dimensionality. In this paper, we present the steepest descent direction based on a family of parameterized policies to overcome those limitations. The update rule is based on stochastic policy gradients employing Amari's natural gradient approach that is moving toward choosing a greedy optimal action. We then show considerable performance improvements of this method in the simple two-state SMDP problem and in the more complex SMDP of call admission control problem.
Ngo Anh Vien, TaeChoong Chung
ICTAI (1)1
2007 Obstacle Avoidance Path Planning for Mobile Robot Based on Ant-Q Reinforcement Learning Algorithm
Ngo Anh Vien, Nguyen Hoang Viet, SeungGwan Lee, TaeChoong Chung
ISNN (1)1