Dinh Tuan Tran

dblp:05/10583 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0001-7443-9102ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Systems, architecture and hardware · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SafePCA: Enhancing Autonomous Robot Navigation in Dynamic Crowds Using Proximal Policy Optimization and Cellular Automata
abstract
Navigating robots in dynamic environments, such as human crowds, is a major challenge due to the trade-off between performance and robustness. Traditional reinforcement learning methods, such as Proximal Policy Optimization (PPO), have shown strong adaptation capabilities but require extensive training and lack explicit mechanisms for collision avoidance. On the other hand, rule-based approaches, such as the Dynamic Window Approach (DWA), offer computational efficiency but struggle with generalization to unseen crowd behaviors. The proposed SafePCA framework aims to address this trade-off by integrating Cellular Automata (CA) into PPO-based navigation. CA enhances robustness by predicting high-risk areas based on pedestrian movement patterns, reducing unnecessary collisions. However, this approach may lead to conservative behavior, potentially affecting navigation performance in reaching the goal efficiently. The core research question addressed in this work is whether SafePCA can balance these trade-offs to ensure safe yet efficient robot navigation in dynamic crowds. Experiments demonstrate that SafePCA outperforms traditional PPO by providing superior risk assessment and avoidance strategies, achieving optimal performance with fewer training episodes. SafePCA's real-time adaptability ensures robust navigation in dynamic environments. By leveraging PPO's adaptive learning and CA's risk analysis, SafePCA offers an efficient solution for autonomous robot navigation in crowded environments, advancing the field and broadening application possibilities.
Ardiansyah Al Farouq, Dinh Tuan Tran, Joo-Ho Lee 0001
ICRA2
2025 Improved 3D Point-Line Mapping Regression for Camera Relocalization
abstract
In this paper, we present a new approach for improving 3D point and line mapping regression for camera re-localization. Previous methods typically rely on feature matching (FM) with stored descriptors or use a single network to encode both points and lines. While FM-based methods perform well in large-scale environments, they become computationally expensive with a growing number of mapping points and lines. Conversely, approaches that learn to encode mapping features within a single network reduce memory footprint but are prone to overfitting, as they may capture unnecessary correlations between points and lines. We propose that these features should be learned independently, each with a distinct focus, to achieve optimal accuracy. To this end, we introduce a new architecture that learns to prioritize each feature independently before combining them for localization. Experimental results demonstrate that our approach significantly enhances the 3D map point and line regression performance for camera re-localization. The implementation of our method will be publicly available at: https://github.com/ais-lab/pl2map/.
Bach-Thuan Bui, Huy-Hoang Bui, Yasuyuki Fujii, Dinh Tuan Tran, Joo-Ho Lee 0001
IROS4
2025 PainDiffusion: Learning to Express Pain
abstract
Accurate pain expression synthesis is essential for improving clinical training and human-robot interaction. Current Robotic Patient Simulators (RPSs) lack realistic pain facial expressions, limiting their effectiveness in medical training. In this work, we introduce PainDiffusion, a generative model that synthesizes naturalistic facial pain expressions. Unlike traditional heuristic or autoregressive methods, PainDiffusion operates in a continuous latent space, ensuring smoother and more natural facial motion while supporting indefinite-length generation via diffusion forcing. Our approach incorporates intrinsic characteristics such as pain expressiveness and emotion, allowing for personalized and controllable pain expression synthesis. We train and evaluate our model using the BioVid HeatPain Database. Additionally, we integrate PainDiffusion into a robotic system to assess its applicability in real-time rehabilitation exercises. Qualitative studies with clinicians reveal that PainDiffusion produces realistic pain expressions, with a 31.2% ± 4.8% preference rate against ground-truth recordings. Our results suggest that PainDiffusion can serve as a viable alternative to real patients in clinical training and simulation, bridging the gap between synthetic and naturalistic pain expression. Code and videos are available at: https://damtien444.github.io/paindf/.
Quang Tien Dam, Tri Tung Nguyen Nguyen, Yuuki Endo, Dinh Tuan Tran, Joo-Ho Lee 0001
IROS4
2025 When Less is More: A Sparse Facial Motion Structure for Listening Motion Learning
abstract
Effective human behavior modeling is critical for successful human–robot interaction. Current state-of-the-art approaches for predicting listening head behavior during dyadic conversations employ continuous-to-discrete representations, where continuous facial motion sequence is converted into discrete latent tokens. However, nonverbal facial motion presents unique challenges owing to its temporal variance and multimodal nature. State-of-the-art discrete motion token representation struggles to capture underlying nonverbal facial patterns making training the listening head inefficient with low-fidelity generated motion. This study proposes a novel method for representing and predicting nonverbal facial motion by encoding long sequences into a sparse sequence of keyframes and transition frames. By identifying crucial motion steps and interpolating intermediate frames, our method preserves the temporal structure of motion while enhancing instance-wise diversity during the learning process. Additionally, we apply this novel sparse representation to the task of listening head prediction, demonstrating its contribution to improving the explanation of facial motion patterns.
Tri Tung Nguyen Nguyen, Tien Quang Dam, Dinh Tuan Tran, Joo-Ho Lee 0001
IEEE Trans. Comput. Soc. Syst.3
2024 Finite Scalar Quantization as Facial Tokenizer for Dyadic Reaction Generation
abstract
Creating a human-like interface in human-robot interaction is a formidable challenge. Many efforts have been made to mimic the human ability of attentive listening and synchronous participation in conversations, especially in terms of facial expressions and head movements. By taking advantage of transformer-based sequence generation models and quantization techniques, this advantage is further enhanced in the areas of text, video, and audio generation. Using Finite Scalar Quantization, we develop a facial expression tokenization module that is able to encode facial expressions in a finite, semantically meaningful vocabulary. Using this module, we establish a more powerful cross-modality transformer-based, non-deterministic model that is able to learn multiple appropriate facial responses in a dyadic conversational context. 1
Quang Tien Dam, Tri Tung Nguyen Nguyen, Dinh Tuan Tran, Joo-Ho Lee 0001
FG3
2024 Leveraging Neural Radiance Field in Descriptor Synthesis for Keypoints Scene Coordinate Regression
abstract
Classical structural-based visual localization methods offer high accuracy but face trade-offs in terms of storage, speed, and privacy. A recent innovation, keypoint scene coordinate regression (KSCR) named D2S addresses these issues by leveraging graph attention networks to enhance keypoint relationships and predict their 3D coordinates using a simple multilayer perceptron (MLP). Camera pose is then determined via PnP+RANSAC, using established 2D-3D correspondences. While KSCR achieves competitive results, rivaling state-of-the-art image-retrieval methods like HLoc across multiple benchmarks, its performance is hindered when data samples are limited due to the deep learning model’s reliance on extensive data. This paper proposes a solution to this challenge by introducing a pipeline for keypoint descriptor synthesis using Neural Radiance Field (NeRF). By generating novel poses and feeding them into a trained NeRF model to create new views, our approach enhances the KSCR’s generalization capabilities in data-scarce environments. The proposed system could significantly improve localization accuracy by up to 50% and cost only a fraction of time for data synthesis. Furthermore, its modular design allows for the integration of multiple NeRFs, offering a versatile and efficient solution for visual localization. The implementation is publicly available at: https://github.com/ais-lab/DescriptorSynthesis4Feat2Map.
Huy-Hoang Bui, Bach-Thuan Bui, Dinh Tuan Tran, Joo-Ho Lee 0001
IROS3
2024 Representing 3D sparse map points and lines for camera relocalization
abstract
Recent advancements in visual localization and mapping have demonstrated considerable success in integrating point and line features. However, expanding the localization framework to include additional mapping components frequently results in increased demand for memory and computational resources dedicated to matching tasks. In this study, we show how a lightweight neural network can learn to represent both 3D point and line features, and exhibit leading pose accuracy by harnessing the power of multiple learned mappings. Specifically, we utilize a single transformer block to encode line features, effectively transforming them into distinctive point-like descriptors. Subsequently, we treat these point and line descriptor sets as distinct yet interconnected feature sets. Through the integration of self- and cross-attention within several graph layers, our method effectively refines each feature before regressing 3D maps using two simple MLPs. In comprehensive experiments, our indoor localization findings surpass those of Hloc and Limap across both point-based and line-assisted configurations. Moreover, in outdoor scenarios, our method secures a significant lead, marking the most considerable enhancement over state-of-the-art learning-based methodologies. The source code and demo videos of this work are publicly available at: https://thpjp.github.io/pl2map/.
Bach-Thuan Bui, Huy-Hoang Bui, Dinh Tuan Tran, Joo-Ho Lee 0001
IROS3
2022 Evaluation of position-keeping strategies for symmetrically-shaped autonomous water-surface robots under disturbances
abstract
Extensive research has been conducted on autonomous surface robots and underwater robots for various tasks in aquatic environments. The duration of the operation of autonomous field robots depends on the capacity of the mounted battery, as they are not typically connected to an external power supply. Therefore, smart strategies which are optimized for each task are required to extend the working time of autonomous field robots. We have developed a symmetrically-shaped au-tonomous surface robot for the long-term monitoring of water quality. In this study, we propose position-keeping strategies to prolong the duration of the symmetrically-shaped surface robot for in-situ monitoring. The proposed position-keeping strategies are evaluated in terms of the power consumption and mean error distance in both practical and simulation environments. The experimental results demonstrate that a robot placed on a water surface with disturbance determines the best course of action to maintain its position based on the environmental conditions and application.
Yasuyuki Fujii, Dinh Tuan Tran, Joo-Ho Lee 0001
IROS2
2021 Pain Expression-based Visual Feedback Method for Care Training Assistant Robot with Musculoskeletal Symptoms
abstract
A human patient simulator (HPS) can achieve effective visual-, auditory-, text-, and alarm-based feedback methods in care or nursing education. Among these, the method of visual feedback is important to design an HPS that can express emotions or feelings of pain like an actual human does because this method allows an immediate reaction between robots and humans. This study aims to develop an avatar-based visual feedback method for a care training assistant robot that can express pain states in joint care education. First, this study introduces its own pain facial expression database from Ritsumeikan University (RU-PITENS) for an avatar with pain expression. The RU-PITENS database contains pain images of 41 Japanese people in their 20s, 30s, 40s, and 60s, and an experiment of pain stimulus is conducted based on transcutaneous electrical nerve stimulation, which is low-cost and easy to use in daily life. Based on the pain images in the RU-PITENS database, we generated an avatar with pain expression to achieve the goal of our study. Since the RUPITENS database does not contain the quantitative pain level, the Siamese network was used to calculate the pain intensity. In addition, the care training assistant robot (CaTARo) developed in our previous study reproduces symptoms of musculoskeletal diseases, and the pain of CaTARo was measured using fuzzy logic theory. As a result, a visual feedback system was constructed to express five types of pain (no pain at all, very faint, weak, moderate, and strong pain) with avatars according to the intensity of the pain output of CaTARo in care training environments.
Miran Lee, Dinh Tuan Tran, Joo-Ho Lee 0001
IROS2
2020 Multi-scale affined-HOF and dimension selection for view-unconstrained action recognition
Dinh Tuan Tran, Hirotake Yamazoe, Joo-Ho Lee 0001
Appl. Intell.1
2015 Integration of a topic probability distribution into surgical phase estimation with a hidden Markov model
abstract
In this paper, we present two new methods to integrate latent Dirichlet allocation (LDA) which is a topic model for surgical workflow phase estimation with a hidden Markov model (HMM). The proposed methods are able to detect surgical phases automatically based on codebook which is built by quantizing the extracted optical flow vectors from the recorded videos of surgical processes. To detect the current phase at a given time point of an operation, some sets of training data with correct phase labels need to be learned by LDA. All documents which are actually short clips divided from the recorded videos are presented as mixtures over learned latent topics. These presentations are then quantized as observed values of a HMM. The major difference between two proposed methods is that while the first method quantizes all topic-based presentations based on k-means, the second method does this based on multivariate Gaussian mixture model. A Left to Right HMM is appropriate for this work because there is no switching the order between surgical phases.
Dinh Tuan Tran, Ryuhei Sakurai, Joo-Ho Lee 0001
IECON1