VLDB 2026 Research / reviewers in the wild / expert
Phong Nguyen 0001
dblp:89/10562-1 · also Phong Ha Nguyen 0001, Phong Nguyen-Ha 0001
· DBLP profile ↗
11ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-9678-0886ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An End-to-End Depth-Based Pipeline for Selfie Image RectificationabstractPortraits or selfie images taken from a close distance typically suffer from perspective distortion. In this paper, we propose an end-to-end deep learning-based rectification pipeline to mitigate the effects of perspective distortion. We learn to predict the facial depth by training a deep CNN. The estimated depth is utilized to adjust the camera-to-subject distance by moving the camera farther, increasing the camera focal length, and reprojecting the 3D image features to the new perspective. The reprojected features are then fed to an inpainting module to fill in the missing pixels. We leverage a differentiable renderer to enable end-to-end training of our depth estimation and feature extraction nets to improve the rectified outputs. To boost the results of the inpainting module, we incorporate an auxiliary module to predict the horizontal movement of the camera which decreases the area that requires hallucination of challenging face parts such as ears. Unlike previous works, we process the full-frame input image at once without cropping the subject's face and processing it separately from the rest of the body, eliminating the need for complex post-processing steps to attach the face back to the subject's body. To train our network, we utilize the popular game engine Unreal Engine to generate a large synthetic face dataset containing various subjects, head poses, expressions, eyewear, clothes, and lighting. Quantitative and qualitative results show that our rectification pipeline outperforms previous methods, and produces comparable results with a time-consuming 3D GAN-based method while being more than 260 times faster. Ahmed Alhawwary, Janne Mustaniemi, Phong Nguyen 0001, Janne Heikkilä |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Diverse Text-to-3D Synthesis with Augmented Text Embedding
Uy Dieu Tran, Minh Luu, Phong Nguyen 0001, Khoi Nguyen 0001, Binh-Son Hua |
ECCV (75) | 3 |
| 2024 | Cascaded and Generalizable Neural Radiance Fields for Fast View SynthesisabstractWe present CG-NeRF, a cascade and generalizable neural radiance fields method for view synthesis. Recent generalizing view synthesis methods can render high-quality novel views using a set of nearby input views. However, the rendering speed is still slow due to the nature of uniformly-point sampling of neural radiance fields. Existing scene-specific methods can train and render novel views efficiently but can not generalize to unseen data. Our approach addresses the problems of fast and generalizing view synthesis by proposing two novel modules: a coarse radiance fields predictor and a convolutional-based neural renderer. This architecture infers consistent scene geometry based on the implicit neural fields and renders new views efficiently using a single GPU. We first train CG-NeRF on multiple 3D scenes of the DTU dataset, and the network can produce high-quality and accurate novel views on unseen real and synthetic data using only photometric losses. Moreover, our method can leverage a denser set of reference images of a single scene to produce accurate novel views without relying on additional explicit representations and still maintains the high-speed rendering of the pre-trained model. Experimental results show that CG-NeRF outperforms state-of-the-art generalizable neural rendering methods on various synthetic and real datasets. Phong Nguyen 0001, Lam Huynh, Esa Rahtu, Jiri Matas, Janne Heikkilä |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Free-Viewpoint RGB-D Human Performance Capture and Rendering
Phong Nguyen 0001, Nikolaos Sarafianos, Christoph Lassner, Janne Heikkilä, Tony Tung |
ECCV (16) | 1 |
| 2022 | Lightweight Monocular Depth with a Novel Neural Architecture Search MethodabstractThis paper presents a novel neural architecture search method, called LiDNAS, for generating lightweight monocular depth estimation models. Unlike previous neural architecture search (NAS) approaches, where finding optimized networks is computationally demanding, the introduced novel Assisted Tabu Search leads to efficient architecture exploration. Moreover, we construct the search space on a pre-defined backbone network to balance layer diversity and search space size. The LiDNAS method outperforms the state-of-the-art NAS approach, proposed for disparity and depth estimation, in terms of search efficiency and output model performance. The LiDNAS optimized models achieve result superior to compact depth estimation state-of-the-art on NYU-Depth-v2, KITTI, and ScanNet, while being 7%-500% more compact in size, i.e the number of model parameters. Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
WACV | 2 |
| 2021 | Monocular Depth Estimation Primed by Salient Point Detection and Normalized Hessian LossabstractDeep neural networks have recently thrived on single image depth estimation. That being said, current developments on this topic highlight an apparent compromise between accuracy and network size. This work proposes an accurate and lightweight framework for monocular depth estimation based on a self-attention mechanism stemming from salient point detection. Specifically, we utilize a sparse set of keypoints to train a FuSaNet model that consists of two major components: Fusion-Net and Saliency-Net. In addition, we introduce a normalized Hessian loss term invariant to scaling and shear along the depth direction, which is shown to substantially improve the accuracy. The proposed method achieves state-of-the-art results on NYU-Depth-v2 and KITTI while using 3.1-38.4 times smaller model in terms of the number of parameters than baseline approaches. Experiments on the SUN-RGBD further demonstrate the generalizability of the proposed method. Lam Huynh, Matteo Pedone, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
3DV | 3 |
| 2021 | RGBD-Net: Predicting Color and Depth Images for Novel Views SynthesisabstractWe propose a new cascaded architecture for novel view synthesis, called RGBD-Net, which consists of two core components: a hierarchical depth regression network and a depth-aware generator network. The former one predicts depth maps of the target views by using adaptive depth scaling, while the latter one leverages the predicted depths and renders spatially and temporally consistent target images. In the experimental evaluation on standard datasets, RGBD-Net not only outperforms the state-of-the-art by a clear margin, but it also generalizes well to new scenes without per-scene optimization. Moreover, we show that RGBD-Net can be optionally trained without depth supervision while still retaining high-quality rendering. Thanks to the depth regression network, RGBD-Net can be also used for creating dense 3D point clouds that are more accurate than those produced by some state-of-the-art multi-view stereo methods. Phong Nguyen 0001, Animesh Karnewar, Lam Huynh, Esa Rahtu, Jiri Matas, Janne Heikkilä |
3DV | 1 |
| 2021 | Boosting Monocular Depth Estimation with Lightweight 3D Point FusionabstractIn this paper, we propose enhancing monocular depth estimation by adding 3D points as depth guidance. Unlike existing depth completion methods, our approach performs well on extremely sparse and unevenly distributed point clouds, which makes it agnostic to the source of the 3D points. We achieve this by introducing a novel multi-scale 3D point fusion network that is both lightweight and efficient. We demonstrate its versatility on two different depth estimation problems where the 3D points have been acquired with conventional structure-from-motion and Li-DAR. In both cases, our network performs on par with state-of-the-art depth completion methods and achieves significantly higher accuracy when only a small number of points is used while being more compact in terms of the number of parameters. We show that our method outperforms some contemporary deep learning based multi-view stereo and structure-from-motion methods both in accuracy and in compactness. Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
ICCV | 2 |
| 2020 | Sequential View Synthesis with Transformer
Phong Nguyen 0001, Lam Huynh, Esa Rahtu, Janne Heikkilä |
ACCV (4) | 1 |
| 2020 | Guiding Monocular Depth Estimation Using Depth-Attention Volume
Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
ECCV (26) | 2 |
| 2018 | Fuzzy-based estimation of continuous Z-distances and discrete directions of home appliances for NIR camera-based gaze tracking system
Jae Woong Jang, Hwan Heo, Jae Won Bang, Hyung Gil Hong, Rizwan Ali Naqvi, Phong Nguyen 0001, Tien Dat Nguyen, Min Beom Lee, Kang Ryoung Park |
Multim. Tools Appl. | 6 |