Phat Nguyen

dblp:225/8229 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Autonomous driving · 42% Trustworthy machine learning · 18% 3D vision · 16%
Software engineering, system software, and programming languages
1 paper
Program analysis · 67% Program verification · 33%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%
Human-computer interaction and pervasive computing
1 paper
Immersive interaction · 77% Learning and educational technologies · 23%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › inverse problem
inverse design
0.912025
ReGen: Generative Robot Simulation via Inverse Design · ICLR 2025
Robotics › Robot manipulation
robot simulation
0.912025
ReGen: Generative Robot Simulation via Inverse Design · ICLR 2025
Robotics › Autonomous driving
safety validation
0.912025
Generating Out-of-Distribution Scenarios Using Language Models · ICRA 2025
Robotics › Autonomous driving
scenario generation
0.912025
Generating Out-of-Distribution Scenarios Using Language Models · ICRA 2025
Robotics › Autonomous driving › autonomous vehicle testing
simulation-based testing
0.912025
Generating Out-of-Distribution Scenarios Using Language Models · ICRA 2025
Program analysis
constraint solving
0.912025
Large Language Models for Safe Minimization · ICSE 2025
Program analysis › control flow analysis
infeasible path detection
0.912025
Large Language Models for Safe Minimization · ICSE 2025
Program verification › decision procedure
satisfiability modulo theories
0.912025
Large Language Models for Safe Minimization · ICSE 2025
Machine learning › Trustworthy machine learning › interpretability
explainable AI
0.812024
XGA-Osteo: Towards XAI-Enabled Knee Osteoarthritis Diagnosis with Adversarial Learning · IJCAI 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
XGA-Osteo: Towards XAI-Enabled Knee Osteoarthritis Diagnosis with Adversarial Learning · IJCAI 2024
Computer vision › 3D vision
camera pose estimation
0.712023
Robust Frame-to-Frame Camera Rotation Estimation in Crowded Scenes · ICCV 2023
Computer vision › 3D vision › camera pose estimation
camera rotation estimation
0.712023
Robust Frame-to-Frame Camera Rotation Estimation in Crowded Scenes · ICCV 2023
Immersive interaction › virtual reality locomotion
redirected walking
0.312018
Experiencing an Invisible World War I Battlefield Through Narrative-Driven Redirected Walking in Virtual Reality · VR 2018
Automated reasoning and model checking › constraint solving
string constraint solving
0.312025
Large Language Models for Safe Minimization · ICSE 2025
Learning and educational technologies › immersive learning
virtual reality learning
0.112018
Experiencing an Invisible World War I Battlefield Through Narrative-Driven Redirected Walking in Virtual Reality · VR 2018

Methods — techniques the papers use, named apart from their topics

large language model · 3.5sample-and-enumerate decoding · 1.7adversarial learning · 1.5vision-language model · 0.9inverse design · 0.9cause-and-effect graph · 0.9SMT solvers · 0.9SMT solver · 0.9optical flow · 0.7RANSAC · 0.7Hough transform on SO(3) · 0.7redirected walking · 0.3narrative-driven resets · 0.3
YearPublicationVenuePosition
2025 ReGen: Generative Robot Simulation via Inverse Design
abstract
Simulation plays a key role in scaling robot learning and validating policies, but constructing simulations remains labor-intensive. In this paper, we introduce ReGen, a generative simulation framework that automates this process using inverse design. Given an agent's behavior (such as a motion trajectory or objective function) and its textual description, we infer the underlying scenarios and environments that could have caused the behavior. Our approach leverages large language models to construct and expand a graph that captures cause-and-effect relationships and relevant entities with properties in the environment, which is then processed to configure a robot simulation environment. Our approach supports (i) augmenting simulations based on ego-agent behaviors, (ii) controllable, counterfactual scenario generation, (iii) reasoning about agent cognition and mental states, and (iv) reasoning with distinct sensing modalities, such as braking due to faulty GPS signals. We demonstrate our method in autonomous driving and robot manipulation tasks, generating more diverse, complex simulated environments compared to existing simulations with high success rates, and enabling controllable generation for corner cases. This approach enhances the validation of robot policies and supports data or simulation augmentation, advancing scalable robot learning for improved generalization and robustness.
Phat Nguyen, Tsun-Hsuan Wang, Zhang-Wei Hong, Erfan Aasi, Andrew Silva, Guy Rosman, Sertac Karaman, Daniela Rus
ICLR1
2025 Generating Out-of-Distribution Scenarios Using Language Models
abstract
The deployment of autonomous vehicles controlled by machine learning techniques requires extensive testing in diverse real-world environments, robust handling of edge cases and out-of-distribution scenarios, and comprehensive safety validation to ensure that these systems can navigate safely and effectively under unpredictable conditions. Addressing Out-OfDistribution (OOD) driving scenarios is essential for enhancing safety, as OOD scenarios help validate the reliability of the models within the vehicle's autonomy stack. However, generating OOD scenarios is challenging due to their long-tailed distribution and rarity in urban driving datasets. Recently, Large Language Models (LLMs) have shown promise in autonomous driving, particularly for their zero-shot generalization and common-sense reasoning capabilities. In this paper, we leverage these LLM strengths to introduce a framework for generating diverse OOD driving scenarios. Our approach uses LLMs to construct a branching tree, where each branch represents a unique OOD scenario. These scenarios are then simulated in the CARLA simulator using an automated framework that aligns scene augmentation with the corresponding textual descriptions. We evaluate our framework through extensive simulations, and assess its performance via a diversity metric that measures the richness of the scenarios. Additionally, we introduce a new “OOD-ness” metric, which quantifies how much the generated scenarios deviate from typical urban driving conditions. Furthermore, we explore the capacity of modern Vision-Language Models (VLMs) to interpret and safely navigate through the simulated OOD scenarios. Our findings offer valuable insights into the reliability of language models in addressing OOD scenarios within the context of urban driving.
Erfan Aasi, Phat Nguyen, Shiva Sreeram, Guy Rosman, Sertac Karaman, Daniela Rus
ICRA2
2025 Large Language Models for Safe Minimization
abstract
Several tasks in program analysis, verification, and testing are modeled as constraint solving problems, utilizing SMT solvers as the reasoning engine. In this work, we aim to investigate the reasoning capabilities of large language models (LLMs) toward reducing the size of an infeasible string constraint system by exploiting inter-constraint interactions such that the remaining ones are still unsatisfiable. We term this safe minimization. Motivated by preliminary observations of hallucination and error propagation in LLMs, we design SafeMin, a framework leveraging an LLM and SMT solver in tandem to ensure a safe and correct minimization. We test the applicability of our approach on string benchmarks from LeetCode in the computation of minimal unsatisfiable subsets (MUSes). We observed that SafeMin helps safely minimize 94.3% of these constraints, with an average minimization ratio of 98% relative to the MUSes. In addition, we assess SafeMin's capabilities in partially enumerating non-unique MUSes, which is baked into our approach via a “sample-and-enumerate” decoding strategy. Overall, we captured 42.1% more non-unique MUSes than without such LLM-based macro-reasoning. Finally, we demonstrate SafeMin's usefulness in detecting infeasible paths in programs.
Aashish Yadavally, Xiaokai Rong, Phat Nguyen, Tien N. Nguyen
ICSE3
2024 XGA-Osteo: Towards XAI-Enabled Knee Osteoarthritis Diagnosis with Adversarial Learning
Hieu Phan Trung, Loc Le Tan, Mao Nguyen, Phat Nguyen, Sang Nguyen, Minh-Triet Tran, Thanh Tho Quan
IJCAI4
2024 Text-to-Drive: Diverse Driving Behavior Synthesis via Large Language Models
abstract
Generating varied scenarios through simulation is crucial for training and evaluating safety-critical systems, such as autonomous vehicles. Yet, the task of modeling the trajectories of other vehicles to simulate diverse and meaningful close interactions remains prohibitively costly. Adopting language descriptions to generate driving behaviors emerges as a promising strategy, offering a scalable and intuitive method for human operators to simulate a wide range of driving interactions. However, the scarcity of large-scale annotated language-trajectory data makes this approach challenging. To address this gap, we propose Text-to-Drive (T2D) to synthesize diverse driving behaviors via Large Language Models (LLMs). We introduce a knowledge-driven approach that operates in two stages. In the first stage, we employ the embedded knowledge of LLMs to generate diverse language descriptions of driving behaviors for a scene. Then, we leverage LLM’s reasoning capabilities to synthesize these behaviors in simulation. At its core, T2D employs an LLM to construct a state chart that maps low-level states to high-level abstractions. This strategy aids in downstream tasks such as summarizing low-level observations, assessing policy alignment with behavior description, and shaping the auxiliary reward, all without needing human supervision. With our knowledge-driven approach, we demonstrate that T2D generates more diverse trajectories compared to other baselines and offers a natural language interface that allows for interactive incorporation of human preference. Please check our website for more examples: here
Phat Nguyen, Tsun-Hsuan Wang, Zhang-Wei Hong, Sertac Karaman, Daniela Rus
IROS1
2023 Robust Frame-to-Frame Camera Rotation Estimation in Crowded Scenes
abstract
We present an approach to estimating camera rotation in crowded, real-world scenes from handheld monocular video. While camera rotation estimation is a well-studied problem, no previous methods exhibit both high accuracy and acceptable speed in this setting. Because the setting is not addressed well by other datasets, we provide a new dataset and benchmark, with high-accuracy, rigorously verified ground truth, on 17 video sequences. Methods developed for wide baseline stereo (e.g., 5-point methods) perform poorly on monocular video. On the other hand, methods used in autonomous driving (e.g., SLAM) leverage specific sensor setups, specific motion models, or local optimization strategies (lagging batch processing) and do not generalize well to handheld video. Finally, for dynamic scenes, commonly used robustification techniques like RANSAC require large numbers of iterations, and become prohibitively slow. We introduce a novel generalization of the Hough transform on SO(3) to efficiently and robustly find the camera rotation most compatible with optical flow. Among comparably fast methods, ours reduces error by almost 50% over the next best, and is more accurate than any method, irrespective of speed. This represents a strong new performance point for crowded scenes, an important setting for computer vision. The code and the dataset are available at https://fabiendelattre.com/robustrotation-estimation.
Fabien Delattre, David Dirnfeld, Phat Nguyen, Stephen Scarano, Michael J. Jones 0001, Pedro Miraldo, Erik G. Learned-Miller
ICCV3
2018 Experiencing an Invisible World War I Battlefield Through Narrative-Driven Redirected Walking in Virtual Reality
abstract
Redirected walking techniques have the potential to provide natural locomotion while users experience large virtual environments. However, when using redirected walking in small physical workspaces, disruptive overt resets are often required. We describe the design of an educational virtual reality experience in which users physically walk through virtual tunnels representative of the World War I battle of Vauquois. Walking in only a 15- by 5-foot tracked space, users are redirected through subtle, narrative-driven resets to walk through a tunnel nearly 50 feet in length. This work contributes approaches and lessons that can be used to provide a seamless and natural virtual reality walking experience in highly constrained physical spaces.
Run Yu 0001, Zachary Duer, J. Todd Ogle, Doug A. Bowman, Thomas W. Tucker, Dongsoo Choi, Zach Bush, Huy Ngo, Phat Nguyen, Xindi Liu
VR10