Phong Nguyen 0001

dblp:89/10562-1 · also Phong Ha Nguyen 0001, Phong Nguyen-Ha 0001 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-9678-0886ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 An End-to-End Depth-Based Pipeline for Selfie Image Rectification
abstract
Portraits or selfie images taken from a close distance typically suffer from perspective distortion. In this paper, we propose an end-to-end deep learning-based rectification pipeline to mitigate the effects of perspective distortion. We learn to predict the facial depth by training a deep CNN. The estimated depth is utilized to adjust the camera-to-subject distance by moving the camera farther, increasing the camera focal length, and reprojecting the 3D image features to the new perspective. The reprojected features are then fed to an inpainting module to fill in the missing pixels. We leverage a differentiable renderer to enable end-to-end training of our depth estimation and feature extraction nets to improve the rectified outputs. To boost the results of the inpainting module, we incorporate an auxiliary module to predict the horizontal movement of the camera which decreases the area that requires hallucination of challenging face parts such as ears. Unlike previous works, we process the full-frame input image at once without cropping the subject's face and processing it separately from the rest of the body, eliminating the need for complex post-processing steps to attach the face back to the subject's body. To train our network, we utilize the popular game engine Unreal Engine to generate a large synthetic face dataset containing various subjects, head poses, expressions, eyewear, clothes, and lighting. Quantitative and qualitative results show that our rectification pipeline outperforms previous methods, and produces comparable results with a time-consuming 3D GAN-based method while being more than 260 times faster.
Ahmed Alhawwary, Janne Mustaniemi, Phong Nguyen 0001, Janne Heikkilä
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Diverse Text-to-3D Synthesis with Augmented Text Embedding
Uy Dieu Tran, Minh Luu, Phong Nguyen 0001, Khoi Nguyen 0001, Binh-Son Hua
ECCV (75)3
2024 Cascaded and Generalizable Neural Radiance Fields for Fast View Synthesis
abstract
We present CG-NeRF, a cascade and generalizable neural radiance fields method for view synthesis. Recent generalizing view synthesis methods can render high-quality novel views using a set of nearby input views. However, the rendering speed is still slow due to the nature of uniformly-point sampling of neural radiance fields. Existing scene-specific methods can train and render novel views efficiently but can not generalize to unseen data. Our approach addresses the problems of fast and generalizing view synthesis by proposing two novel modules: a coarse radiance fields predictor and a convolutional-based neural renderer. This architecture infers consistent scene geometry based on the implicit neural fields and renders new views efficiently using a single GPU. We first train CG-NeRF on multiple 3D scenes of the DTU dataset, and the network can produce high-quality and accurate novel views on unseen real and synthetic data using only photometric losses. Moreover, our method can leverage a denser set of reference images of a single scene to produce accurate novel views without relying on additional explicit representations and still maintains the high-speed rendering of the pre-trained model. Experimental results show that CG-NeRF outperforms state-of-the-art generalizable neural rendering methods on various synthetic and real datasets.
Phong Nguyen 0001, Lam Huynh, Esa Rahtu, Jiri Matas, Janne Heikkilä
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Free-Viewpoint RGB-D Human Performance Capture and Rendering
Phong Nguyen 0001, Nikolaos Sarafianos, Christoph Lassner, Janne Heikkilä, Tony Tung
ECCV (16)1
2022 Lightweight Monocular Depth with a Novel Neural Architecture Search Method
abstract
This paper presents a novel neural architecture search method, called LiDNAS, for generating lightweight monocular depth estimation models. Unlike previous neural architecture search (NAS) approaches, where finding optimized networks is computationally demanding, the introduced novel Assisted Tabu Search leads to efficient architecture exploration. Moreover, we construct the search space on a pre-defined backbone network to balance layer diversity and search space size. The LiDNAS method outperforms the state-of-the-art NAS approach, proposed for disparity and depth estimation, in terms of search efficiency and output model performance. The LiDNAS optimized models achieve result superior to compact depth estimation state-of-the-art on NYU-Depth-v2, KITTI, and ScanNet, while being 7%-500% more compact in size, i.e the number of model parameters.
Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä
WACV2
2021 Monocular Depth Estimation Primed by Salient Point Detection and Normalized Hessian Loss
abstract
Deep neural networks have recently thrived on single image depth estimation. That being said, current developments on this topic highlight an apparent compromise between accuracy and network size. This work proposes an accurate and lightweight framework for monocular depth estimation based on a self-attention mechanism stemming from salient point detection. Specifically, we utilize a sparse set of keypoints to train a FuSaNet model that consists of two major components: Fusion-Net and Saliency-Net. In addition, we introduce a normalized Hessian loss term invariant to scaling and shear along the depth direction, which is shown to substantially improve the accuracy. The proposed method achieves state-of-the-art results on NYU-Depth-v2 and KITTI while using 3.1-38.4 times smaller model in terms of the number of parameters than baseline approaches. Experiments on the SUN-RGBD further demonstrate the generalizability of the proposed method.
Lam Huynh, Matteo Pedone, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä
3DV3
2021 RGBD-Net: Predicting Color and Depth Images for Novel Views Synthesis
abstract
We propose a new cascaded architecture for novel view synthesis, called RGBD-Net, which consists of two core components: a hierarchical depth regression network and a depth-aware generator network. The former one predicts depth maps of the target views by using adaptive depth scaling, while the latter one leverages the predicted depths and renders spatially and temporally consistent target images. In the experimental evaluation on standard datasets, RGBD-Net not only outperforms the state-of-the-art by a clear margin, but it also generalizes well to new scenes without per-scene optimization. Moreover, we show that RGBD-Net can be optionally trained without depth supervision while still retaining high-quality rendering. Thanks to the depth regression network, RGBD-Net can be also used for creating dense 3D point clouds that are more accurate than those produced by some state-of-the-art multi-view stereo methods.
Phong Nguyen 0001, Animesh Karnewar, Lam Huynh, Esa Rahtu, Jiri Matas, Janne Heikkilä
3DV1
2021 Boosting Monocular Depth Estimation with Lightweight 3D Point Fusion
abstract
In this paper, we propose enhancing monocular depth estimation by adding 3D points as depth guidance. Unlike existing depth completion methods, our approach performs well on extremely sparse and unevenly distributed point clouds, which makes it agnostic to the source of the 3D points. We achieve this by introducing a novel multi-scale 3D point fusion network that is both lightweight and efficient. We demonstrate its versatility on two different depth estimation problems where the 3D points have been acquired with conventional structure-from-motion and Li-DAR. In both cases, our network performs on par with state-of-the-art depth completion methods and achieves significantly higher accuracy when only a small number of points is used while being more compact in terms of the number of parameters. We show that our method outperforms some contemporary deep learning based multi-view stereo and structure-from-motion methods both in accuracy and in compactness.
Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä
ICCV2
2020 Sequential View Synthesis with Transformer
Phong Nguyen 0001, Lam Huynh, Esa Rahtu, Janne Heikkilä
ACCV (4)1
2020 Guiding Monocular Depth Estimation Using Depth-Attention Volume
Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä
ECCV (26)2
2018 Fuzzy-based estimation of continuous Z-distances and discrete directions of home appliances for NIR camera-based gaze tracking system
Jae Woong Jang, Hwan Heo, Jae Won Bang, Hyung Gil Hong, Rizwan Ali Naqvi, Phong Nguyen 0001, Tien Dat Nguyen, Min Beom Lee, Kang Ryoung Park
Multim. Tools Appl.6