VLDB 2026 Research / reviewers in the wild / expert
Anh Tuan Tran 0001
dblp:150/5269-1 · also Anh Tran 0001
· DBLP profile ↗
47ranked-venue papers
3as first author
36since 2021 · last 2026
0000-0002-3120-4036ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 3 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 3 first-author · 26 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unveiling open vocabulary relationships in images: A translation embedding approach guided by vision-language models
Nguyet Nguyen, Cong Tran, Anh Tuan Tran 0001, Cuong Pham 0001 |
Comput. Vis. Image Underst. | 3 |
| 2025 | Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Image GenerationabstractFlow matching has emerged as a promising framework for training generative models, demonstrating impressive empirical performance while offering relative ease of training compared to diffusion-based models. However, this method still requires numerous function evaluations in the sampling process. To address these limitations, we introduce a self-corrected flow distillation method that effectively integrates consistency models and adversarial training within the flow-matching framework. This work is a pioneer in achieving consistent generation quality in both few-step and one-step sampling. Our extensive experiments validate the effectiveness of our method, yielding superior results both quantitatively and qualitatively on CelebA-HQ and zero-shot benchmarks on the COCO dataset. Quan Dao, Hao Phung, Trung Tuan Dao, Dimitris N. Metaxas, Anh Tuan Tran 0001 |
AAAI | 5 |
| 2025 | Any3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask TrackingabstractExisting 3D instance segmentation methods frequently encounter issues with over-segmentation, leading to redundant and inaccurate 3D proposals that complicate downstream tasks. This challenge arises from their unsupervised merging approach, where dense 2D instance masks are lifted across frames into point clouds to form 3D candidate proposals without direct supervision. These candidates are then hierarchically merged based on heuristic criteria, often resulting in numerous redundant segments that fail to combine into precise 3D proposals. To overcome these limitations, we propose a 3D-Aware 2D Mask Tracking module that uses robust 3D priors from a 2D mask segmentation and tracking foundation model (SAM-2) to ensure consistent object masks across video frames. Rather than merging all visible superpoints across views to create a 3D mask, our 3D Mask Optimization module leverages a dynamic programming algorithm to select an optimal set of views, refining the superpoints to produce a final 3D proposal for each object. Our approach achieves comprehensive object coverage within the scene, significantly reduces unnecessary proposals by half, and lowers the average runtime by a factor of 10, mitigating potential negative impacts on downstream applications. Evaluations on ScanNet200 and ScanNet++ confirm the effectiveness of our method, with improvements across Class-Agnostic, Open-Vocabulary, and Open-Ended 3D Instance Segmentation tasks. Minh Luu, Anh Tuan Tran 0001, Cuong Pham 0001, Khoi Nguyen 0001 |
CVPR | 3 |
| 2025 | SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step DiffusionabstractRecent advances in text-guided image editing enable users to perform image edits through simple text inputs, leveraging the extensive priors of multi-step diffusion-based text-to-image models. However, these methods often fall short of the speed demands required for real-world and on-device applications due to the costly multi-step inversion and sampling process involved. In response to this, we introduce SwiftEdit, a simple yet highly efficient editing tool that achieve instant text-guided image editing (in 0.23s). The advancement of SwiftEdit lies in its two novel contributions: a one-step inversion framework that enables one-step image reconstruction via inversion and a mask-guided editing technique with our proposed attention rescaling mechanism to perform localized image editing. Extensive experiments are provided to demonstrate the effectiveness and efficiency of SwiftEdit. In particular, SwiftEdit enables instant text-guided image editing, which is extremely faster than previous multi-step methods (at least 50× times faster) while maintain a competitive performance in editing results. Our project is at https://swift-edit.github.io/. Trong-Tung Nguyen, Khoi Nguyen 0001, Anh Tuan Tran 0001, Cuong Pham 0001 |
CVPR | 4 |
| 2025 | CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
Quang-Binh Nguyen, Minh Luu, Anh Tuan Tran 0001, Khoi Nguyen 0001 |
ICCV | 4 |
| 2025 | Supercharged One-Step Text-to-Image Diffusion Models with Negative Prompts
Viet Nguyen, Trung Dao, Khoi Nguyen 0001, Cuong Pham 0001, Toan Tran 0003, Anh Tuan Tran 0001 |
ICCV | 7 |
| 2025 | SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion ModelsabstractThe rapid growth of text-to-image diffusion models has raised concerns about their potential misuse in generating harmful or unauthorized contents. To address these issues, several Concept Erasure methods have been proposed. However, most of them fail to achieve both robustness, i.e., the ability to robustly remove the target concept., and effectiveness, i.e., maintaining image quality. While few recent techniques successfully achieve these goals for NSFW concepts, none could handle narrow concepts such as copyrighted characters or celebrities. Erasing these narrow concepts is critical in addressing copyright and legal concerns. However, erasing them is challenging due to their close distances to non-target neighboring concepts, requiring finer-grained manipulation. In this paper, we introduce Subspace Mapping (SuMa), a novel method specifically designed to achieve both robustness and effectiveness in easing these narrow concepts. SuMa first derives a target subspace representing the concept to be erased and then neutralizes it by mapping it to a reference subspace that minimizes the distance between the two. This mapping ensures the target concept is robustly erased while preserving image quality. We conduct extensive experiments with SuMa across four tasks: subclass erasure, celebrity erasure, artistic style erasure, and instance erasure and compare the results with current state-of-the-art methods. Our method achieves image quality comparable to approaches focused on effectiveness, while also yielding results that are on par with methods targeting completeness. Anh Tuan Tran 0001, Cuong Pham 0001 |
ICCV | 2 |
| 2025 | Improved Training Technique for Shortcut ModelsabstractShortcut models represent a promising, non-adversarial paradigm for generative modeling, uniquely supporting one-step, few-step, and multi-step sampling from a single trained network. However, their widespread adoption has been stymied by critical performance bottlenecks. This paper tackles the five core issues that held shortcut models back: (1) the hidden flaw of compounding guidance, which we are the first to formalize, causing severe image artifacts; (2) inflexible fixed guidance that restricts inference-time control; (3) a pervasive frequency bias driven by a reliance on low-level distances in the direct domain, which biases reconstructions toward low frequencies; (4) divergent self-consistency arising from a conflict with EMA training; and (5) curvy flow trajectories that impede convergence. To address these challenges, we introduce iSM, a unified training framework that systematically resolves each limitation. Our framework is built on four key improvements: Intrinsic Guidance provides explicit, dynamic control over guidance strength, resolving both compounding guidance and inflexibility. A Multi-Level Wavelet Loss mitigates frequency bias to restore high-frequency details. Scaling Optimal Transport (sOT) reduces training variance and learns straighter, more stable generative paths. Finally, a Twin EMA strategy reconciles training stability with self-consistency. Extensive experiments on ImageNet 256x256 demonstrate that our approach yields substantial FID improvements over baseline shortcut models across one-step, few-step, and multi-step generation, making shortcut models a viable and competitive class of generative models. Viet Nguyen, Duc Vu, Trung Dao, Chi Tran, Toan Tran 0003, Anh Tuan Tran 0001 |
NeurIPS | 7 |
| 2025 | CoreNet: Leveraging context-aware representations via MLP networks for CTR prediction
Khoi N. P. Dang, Thu Thuy Tran, Ta Cong Son, Anh Tuan Tran 0001 |
Knowl. Based Syst. | 4 |
| 2024 | EFHQ: Multi-Purpose ExtremePose-Face-HQ DatasetabstractThe existing facial datasets, while having plentiful images at near frontal views, lack images with extreme head poses, leading to the downgraded performance of deep learning models when dealing with profile or pitched faces. This work aims to address this gap by introducing a novel dataset named Extreme Pose Face High-Quality Dataset (EFHQ), which includes a maximum of 450k high-quality images of faces at extreme poses. To produce such a massive dataset, we utilize a novel and meticulous dataset processing pipeline to curate two publicly available datasets, VFHQ and CelebV-HQ, which contain many high-resolution face videos captured in various settings. Our dataset can complement existing datasets on various facial-related tasks, such as facial synthesis with 2D/3D-aware GAN, diffusion-based text-to-image face generation, and face reenactment. Specifically, training with EFHQ helps models generalize well across diverse poses, significantly improving performance in scenarios involving extreme views, confirmed by extensive experiments. Additionally, we utilize EFHQ to define a challenging cross-view face verification benchmark, in which the performance of SOTA face recognition models drops 5-37% compared to frontal-to-frontal scenarios, aiming to stimulate studies on face recognition under severe pose conditions in the wild. Trung Tuan Dao, Duc Hong Vu, Cuong Pham 0001, Anh Tuan Tran 0001 |
CVPR | 4 |
| 2024 | Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask GuidanceabstractWe introduce Open3DIS, a novel solution designed to tackle the problem of Open-Vocabulary Instance Segmentation within 3D scenes. Objects within 3D environments exhibit diverse shapes, scales, and colors, making precise instance-level identification a challenging task. Recent advancements in Open-Vocabulary scene understanding have made significant strides in this area by employing class-agnostic 3D instance proposal networks for object localization and learning queryable features for each 3D mask. While these methods produce high-quality instance proposals, they struggle with identifying small-scale and geometrically ambiguous objects. The key idea of our method is a new module that aggregates 2D instance masks across frames and maps them to geometrically coherent point cloud regions as high-quality object proposals addressing the above limitations. These are then combined with 3D class-agnostic instance proposals to include a wide range of objects in the real world. To validate our approach, we conducted experiments on three prominent datasets, including ScanNet200, S3DIS, and Replica, demonstrating significant performance gains in segmenting objects with diverse categories over the state-of-the-art approaches. Phuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan 0001, Anh Tuan Tran 0001, Cuong Pham 0001, Khoi Nguyen 0001 |
CVPR | 5 |
| 2024 | Blur2Blur: Blur Conversion for Unsupervised Image Deblurring on Unknown DomainsabstractThis paper presents an innovative framework designed to train an image deblurring algorithm tailored to a specific camera device. This algorithm works by transforming a blurry input image, which is challenging to deblur, into another blurry image that is more amenable to deblurring. The transformation process, from one blurry state to another, leverages unpaired data consisting of sharp and blurry images captured by the target camera device. Learning this blur-to-blur transformation is inherently simpler than direct blur-to-sharp conversion, as it primarily involves modifying blur patterns rather than the intricate task of reconstructing fine image details. The efficacy of the proposed approach has been demonstrated through comprehensive experiments on various benchmarks, where it significantly outperforms state-of-the-art methods both quantitatively and qualitatively. Our code and data are available at https://github.com/VinAIResearch/Blur2Blur Bang-Dang Pham, Anh Tuan Tran 0001, Cuong Pham 0001, Rang Nguyen, Minh Hoai |
CVPR | 3 |
| 2024 | VOODOO 3D: Volumetric Portrait Disentanglement for One-Shot 3D Head ReenactmentabstractWe present a 3D-aware one-shot head reenactment method based on a fully volumetric neural disentanglement framework for source appearance and driver expressions. Our method is real-time and produces high-fidelity and view-consistent output, suitable for 3D teleconferencing systems based on holographic displays. Existing cutting-edge 3D-aware reenactment methods often use neural radiance fields or 3D meshes to produce view-consistent appearance encoding, but, at the same time, they rely on linear face models, such as 3DMM, to achieve its disentanglement with facial expressions. As a result, their reenactment results often exhibit identity leakage from the driver or have unnatural expressions. To address these problems, we propose a neural self-supervised disentanglement approach that lifts both the source image and driver video frame into a shared 3D volumetric representation based on tri-planes. This representation can then be freely manipu-lated with expression tri-planes extracted from the driving images and rendered from an arbitrary view using neural radiance fields. We achieve this disentanglement via self-supervised learning on a large in-the-wild video dataset. We further introduce a highly effective fine-tuning approach to improve the generalizability of the 3D lifting using the same real-world data. We demonstrate state-of-the-art performance on a wide range of datasets, and also showcase high-quality 3D-aware head reenactment on highly challenging and diverse subjects, including non-frontal head poses and complex expressions for both source and driver. Egor Zakharov, Long-Nhat Ho, Anh Tuan Tran 0001, Liwen Hu 0001, Hao Li 0015 |
CVPR | 4 |
| 2024 | Time-Efficient and Identity-Consistent Virtual Try-On Using A Variant of Altered Diffusion Models
Phuong Dam, Jihoon Jeong, Anh Tuan Tran 0001, Daeyoung Kim 0001 |
ECCV (72) | 3 |
| 2024 | SwiftBrush V2: Make Your One-Step Diffusion Model Better Than Its Teacher
Trung Tuan Dao, Thuan Hoang Nguyen, Duc Hong Vu, Khoi Nguyen 0001, Cuong Pham 0001, Anh Tuan Tran 0001 |
ECCV (82) | 7 |
| 2024 | A High-Quality Robust Diffusion Framework for Corrupted Dataset
Quan Dao, Binh Ta, Anh Tuan Tran 0001 |
ECCV (84) | 4 |
| 2024 | Data Poisoning Quantization Backdoor Attack
Tran Huynh, Anh Tuan Tran 0001, Khoa D. Doan |
ECCV (84) | 2 |
| 2024 | Flatness-Aware Sequential Learning Generates Resilient Backdoors
The-Anh Ta, Anh Tuan Tran 0001, Khoa D. Doan |
ECCV (87) | 3 |
| 2024 | DiMSUM: Diffusion Mamba - A Scalable and Unified Spatial-Frequency Method for Image GenerationabstractWe introduce a novel state-space architecture for diffusion models, effectively harnessing spatial and frequency information to enhance the inductive bias towards local features in input images for image generation tasks. While state-space networks, including Mamba, a revolutionary advancement in recurrent neural networks, typically scan input sequences from left to right, they face difficulties in designing effective scanning strategies, especially in the processing of image data. Our method demonstrates that integrating wavelet transformation into Mamba enhances the local structure awareness of visual inputs and better captures long-range relations of frequencies by disentangling them into wavelet subbands, representing both low- and high-frequency components. These wavelet-based outputs are then processed and seamlessly fused with the original Mamba outputs through a cross-attention fusion layer, combining both spatial and frequency information to optimize the order awareness of state-space models which is essential for the details and overall quality of image generation. Besides, we introduce a globally-shared transformer to supercharge the performance of Mamba, harnessing its exceptional power to capture global relationships. Through extensive experiments on standard benchmarks, our method demonstrates superior results compared to DiT and DIFFUSSM, achieving faster training convergence and delivering high-quality outputs. The codes and pretrained models are released at https://github.com/VinAIResearch/DiMSUM.git. Hao Phung, Quan Dao, Trung Tuan Dao, Viet Hoang Phan, Dimitris N. Metaxas, Anh Tuan Tran 0001 |
NeurIPS | 6 |
| 2024 | VOODOO XP: Expressive One-Shot Head Reenactment for VR TelepresenceabstractWe introduce VOODOO XP: a 3D-aware one-shot head reenactment method that can generate highly expressive facial expressions driven by an input video from a single 2D portrait. Our approach is real-time, view-consistent, and can be instantly used without calibration or fine-tuning. We demonstrate our solution in a monocular video setting and an end-to-end VR telepresence system for two-way communication. Compared to 2D head reenactment methods, 3D-aware approaches aim to preserve the identity of the subject and ensure view-consistent facial geometry for novel camera poses, which makes them suitable for immersive applications. While various facial disentanglement techniques have been introduced, cutting-edge 3D-aware neural reenactment techniques still lack expressiveness and fail to reproduce complex and fine-scale facial expressions. We present a novel cross-reenactment architecture that directly transfers the driver's facial expressions to transformer blocks of the input source's 3D lifting module. We show that highly effective disentanglement is possible using a new multi-stage self-supervision approach. It relies on a coarse-to-fine training strategy, which is combined with explicit face neutralization and 3D lifted frontalization during its initial training stage. We further integrate our novel head reenactment solution into an accessible high-fidelity VR telepresence system, where any person can instantly build a personalized neural head avatar from any photo and bring it to life using the headset. Furthermore, our proposed method demonstrates state-of-the-art expressiveness and likeness preservation on diverse subjects and capture conditions. Egor Zakharov, Long-Nhat Ho, Adilbek Karmanov, Ariana Bermudez Venegas, McLean Goldwhite, Aviral Agarwal, Liwen Hu 0001, Anh Tuan Tran 0001, Hao Li 0015 |
ACM Trans. Graph. | 9 |
| 2023 | Efficient Scale-Invariant Generator with Column-Row Entangled Pixel SynthesisabstractAny-scale image synthesis offers an efficient and scalable solution to synthesize photo-realistic images at any scale, even going beyond 2K resolution. However, existing GAN-based solutions depend excessively on convolutions and a hierarchical architecture, which introduce inconsistency and the “texture sticking” issue when scaling the output resolution. From another perspective, INR-based generators are scale-equivariant by design, but their huge memory footprint and slow inference hinder these networks from being adopted in large-scale or real-time systems. In this work, we propose Column-Row Entangled Pixel Synthesis (CREPS), a new generative model that is both efficient and scale-equivariant without using any spatial convolutions or coarse-to-fine design. To save memory footprint and make the system scalable, we employ a novel bi-line representation that decomposes layer-wise feature maps into separate “thick” column and row encodings. Experiments on various datasets, including FFHQ, LSUN-Church, MetFaces, and Flickr-Scenery, confirm CREPS’ ability to synthesize scale-consistent and alias-free images at any arbitrary resolution with proper training and inference speed. Code is available at https://github.com/VinAIResearch/CREPS. Thuan Hoang Nguyen, Thanh Van Le, Anh Tuan Tran 0001 |
CVPR | 3 |
| 2023 | HyperCUT: Video Sequence from a Single Blurry Image using Unsupervised OrderingabstractWe consider the challenging task of training models for image-to-video deblurring, which aims to recover a sequence of sharp images corresponding to a given blurry image input. A critical issue disturbing the training of an image-to-video model is the ambiguity of the frame ordering since both the forward and backward sequences are plausible solutions. This paper proposes an effective self-supervised ordering scheme that allows training highquality image-to-video deblurring models. Unlike previous methods that rely on order-invariant losses, we assign an explicit order for each video sequence, thus avoiding the order-ambiguity issue. Specifically, we map each video sequence to a vector in a latent high-dimensional space so that there exists a hyperplane such that for every video sequence, the vectors extracted from it and its reversed sequence are on different sides of the hyperplane. The side of the vectors will be used to define the order of the corresponding sequence. Last but not least, we propose a real-image dataset for the image-to-video deblurring problem that covers a variety of popular domains, including face, hand, and street. Extensive experimental results confirm the effectiveness of our method. Code and data are available at https://github.com/VinAIResearch/HyperCUT.git Bang-Dang Pham, Anh Tuan Tran 0001, Cuong Pham 0001, Rang Nguyen, Minh Hoai |
CVPR | 3 |
| 2023 | Wavelet Diffusion Models are fast and scalable Image GeneratorsabstractDiffusion models are rising as a powerful solution for high-fidelity image generation, which exceeds GANs in quality in many circumstances. However, their slow training and inference speed is a huge bottleneck, blocking them from being used in real-time applications. A recent DiffusionGAN method significantly decreases the models' running time by reducing the number of sampling steps from thousands to several, but their speeds still largely lag behind the GAN counterparts. This paper aims to reduce the speed gap by proposing a novel wavelet-based diffusion scheme. We extract low-and-high frequency components from both image and feature levels via wavelet decomposition and adaptively handle these components for faster processing while maintaining good generation quality. Furthermore, we propose to use a reconstruction term, which effectively boosts the model training convergence. Experimental results on CelebA-HQ, CIFAR-10, LSUN-Church, and STL-10 datasets prove our solution is a stepping-stone to offering real-time and high-fidelity diffusion models. Our code and pre-trained checkpoints are available at https://github.com/VinAIResearch/WaveDiff.git. Hao Phung, Quan Dao, Anh Tuan Tran 0001 |
CVPR | 3 |
| 2023 | Anti-DreamBooth: Protecting users from personalized text-to-image synthesisabstractText-to-image diffusion models are nothing but a revolution, allowing anyone, even without design skills, to create realistic images from simple text inputs. With powerful personalization tools like DreamBooth, they can generate images of a specific person just by learning from his/her few reference images. However, when misused, such a powerful and convenient tool can produce fake news or disturbing content targeting any individual victim, posing a severe negative social impact. In this paper, we explore a defense system called Anti-DreamBooth against such malicious use of DreamBooth. The system aims to add subtle noise perturbation to each user’s image before publishing in order to disrupt the generation quality of any DreamBooth model trained on these perturbed images. We investigate a wide range of algorithms for perturbation optimization and extensively evaluate them on two facial datasets over various text-to-image model versions. Despite the complicated formulation of Dream-Booth and Diffusion-based text-to-image models, our methods effectively defend users from the malicious use of those models. Their effectiveness withstands even adverse conditions, such as model or prompt/term mismatching between training and testing. Our code will be available at https://github.com/VinAIResearch/Anti-DreamBooth.git. Thanh Van Le, Hao Phung, Thuan Hoang Nguyen, Quan Dao, Ngoc N. Tran, Anh Tuan Tran 0001 |
ICCV | 6 |
| 2023 | Dataset Diffusion: Diffusion-based Synthetic Data Generation for Pixel-Level Semantic SegmentationabstractPreparing training data for deep vision models is a labor-intensive task. To address this, generative models have emerged as an effective solution for generating synthetic data. While current generative models produce image-level category labels, we propose a novel method for generating pixel-level semantic segmentation labels using the text-to-image generative model Stable Diffusion (SD). By utilizing the text prompts, cross-attention, and self-attention of SD, we introduce three new techniques: class-prompt appending, class-prompt cross-attention, and self-attention exponentiation. These techniques enable us to generate segmentation maps corresponding to synthetic images. These maps serve as pseudo-labels for training semantic segmenters, eliminating the need for labor-intensive pixel-wise annotation. To account for the imperfections in our pseudo-labels, we incorporate uncertainty regions into the segmentation, allowing us to disregard loss from those regions. We conduct evaluations on two datasets, PASCAL VOC and MSCOCO, and our approach significantly outperforms concurrent work. Our benchmarks and code will be released at https://github.com/VinAIResearch/Dataset-Diffusion. Truong Vu, Anh Tuan Tran 0001, Khoi Nguyen 0001 |
NeurIPS | 3 |
| 2023 | IBA: Towards Irreversible Backdoor Attacks in Federated LearningabstractFederated learning (FL) is a distributed learning approach that enables machine learning models to be trained on decentralized data without compromising end devices' personal, potentially sensitive data. However, the distributed nature and uninvestigated data intuitively introduce new security vulnerabilities, including backdoor attacks. In this scenario, an adversary implants backdoor functionality into the global model during training, which can be activated to cause the desired misbehaviors for any input with a specific adversarial pattern. Despite having remarkable success in triggering and distorting model behavior, prior backdoor attacks in FL often hold impractical assumptions, limited imperceptibility, and durability. Specifically, the adversary needs to control a sufficiently large fraction of clients or know the data distribution of other honest clients. In many cases, the trigger inserted is often visually apparent, and the backdoor effect is quickly diluted if the adversary is removed from the training process. To address these limitations, we propose a novel backdoor attack framework in FL, the Irreversible Backdoor Attack (IBA), that jointly learns the optimal and visually stealthy trigger and then gradually implants the backdoor into a global model. This approach allows the adversary to execute a backdoor attack that can evade both human and machine inspections. Additionally, we enhance the efficiency and durability of the proposed attack by selectively poisoning the model's parameters that are least likely updated by the main task's learning process and constraining the poisoned model update to the vicinity of the global model. Finally, we evaluate the proposed attack framework on several benchmark datasets, including MNIST, CIFAR-10, and Tiny ImageNet, and achieved high success rates while simultaneously bypassing existing backdoor defenses and achieving a more durable backdoor effect compared to other backdoor attacks. Overall, IBA offers a more effective, stealthy, and durable approach to backdoor attacks in FL. The code associated with this paper is available on [GitHub](https://github.com/sail-research/iba). Thuy Dung Nguyen, Anh Tuan Tran 0001, Khoa D. Doan, Kok-Seng Wong |
NeurIPS | 3 |
| 2022 | Self-supervised Learning with Multi-view Rendering for 3D Point Cloud Analysis
Bach Tran, Binh-Son Hua, Anh Tuan Tran 0001, Minh Hoai |
ACCV (1) | 3 |
| 2022 | HyperInverter: Improving StyleGAN Inversion via HypernetworkabstractReal-world image manipulation has achieved fantastic progress in recent years as a result of the exploration and utilization of GAN latent spaces. GAN inversion is the first step in this pipeline, which aims to map the real image to the latent code faithfully. Unfortunately, the majority of existing GAN inversion methods fail to meet at least one of the three requirements listed below: high reconstruction quality, editability, and fast inference. We present a novel two-phase strategy in this research that fits all requirements at the same time. In the first phase, we train an encoder to map the input image to StyleGAN2 W-space, which was proven to have excellent editability but lower reconstruction quality. In the second phase, we supplement the reconstruction ability in the initial phase by leveraging a series of hypernetworks to recover the missing information during inversion. These two steps complement each other to yield high reconstruction quality thanks to the hypernetwork branch and excellent editability due to the inversion done in the W-space. Our method is entirely encoder-based, resulting in extremely fast inference. Extensive experiments on two challenging datasets demonstrate the superiority of our method.11Project page: https://di-mi-ta.github.io/HyperInverter Tan M. Dinh, Anh Tuan Tran 0001, Rang Nguyen, Binh-Son Hua |
CVPR | 2 |
| 2022 | QC-StyleGAN - Quality Controllable Image Generation and ManipulationabstractThe introduction of high-quality image generation models, particularly the StyleGAN family, provides a powerful tool to synthesize and manipulate images. However, existing models are built upon high-quality (HQ) data as desired outputs, making them unfit for in-the-wild low-quality (LQ) images, which are common inputs for manipulation. In this work, we bridge this gap by proposing a novel GAN structure that allows for generating images with controllable quality. The network can synthesize various image degradation and restore the sharp image via a quality control code. Our proposed QC-StyleGAN can directly edit LQ images without altering their quality by applying GAN inversion and manipulation techniques. It also provides for free an image restoration solution that can handle various degradations, including noise, blur, compression artifacts, and their mixtures. Finally, we demonstrate numerous other applications such as image degradation synthesis, transfer, and interpolation. Dat Viet Thanh Nguyen, Phong Tran The, Tan M. Dinh, Cuong Pham 0001, Anh Tuan Tran 0001 |
NeurIPS | 5 |
| 2021 | Progressive Semantic SegmentationabstractThe objective of this work is to segment high-resolution images without overloading GPU memory usage or losing the fine details in the output segmentation map. The memory constraint means that we must either downsample the big image or divide the image into local patches for separate processing. However, the former approach would lose the fine details, while the latter can be ambiguous due to the lack of a global picture. In this work, we present MagNet, a multi-scale framework that resolves local ambiguity by looking at the image at multiple magnification levels. MagNet has multiple processing stages, where each stage corresponds to a magnification level, and the output of one stage is fed into the next stage for coarse-to-fine information propagation. Each stage analyzes the image at a higher resolution than the previous stage, recovering the previously lost details due to the lossy downsampling step, and the segmentation output is progressively refined through the processing stages. Experiments on three high-resolution datasets of urban views, aerial scenes, and medical images show that MagNet consistently outperforms the state-of-the-art methods by a significant margin. Code is available at https://github.com/VinAIResearch/MagNet. Chuong Huynh, Anh Tuan Tran 0001, Khoa Luu, Minh Hoai |
CVPR | 2 |
| 2021 | Lipstick Ain't Enough: Beyond Color Matching for In-the-Wild Makeup TransferabstractMakeup transfer is the task of applying on a source face the makeup style from a reference image. Real-life makeups are diverse and wild, which cover not only color-changing but also patterns, such as stickers, blushes, and jewelries. However, existing works overlooked the latter components and confined makeup transfer to color manipulation, focusing only on light makeup styles. In this work, we propose a holistic makeup transfer framework that can handle all the mentioned makeup components. It consists of an improved color transfer branch and a novel pattern transfer branch to learn all makeup properties, including color, shape, texture, and location. To train and evaluate such a system, we also introduce new makeup datasets for real and synthetic extreme makeup. Experimental results show that our framework achieves the state of the art performance on both light and extreme makeup styles. Code is available at https://github.com/VinAIResearch/CPM. Anh Tuan Tran 0001, Minh Hoai |
CVPR | 2 |
| 2021 | Explore Image Deblurring via Encoded Blur Kernel SpaceabstractThis paper introduces a method to encode the blur operators of an arbitrary dataset of sharp-blur image pairs into a blur kernel space. Assuming the encoded kernel space is close enough to in-the-wild blur operators, we propose an alternating optimization algorithm for blind image deblurring. It approximates an unseen blur operator by a kernel in the encoded space and searches for the corresponding sharp image. Unlike recent deep-learning-based methods, our system can handle unseen blur kernel, while avoiding using complicated handcrafted priors on the blur operator often found in classical methods. Due to the method’s design, the encoded kernel space is fully differentiable, thus can be easily adopted in deep neural network models. Moreover, our method can be used for blur synthesis by transferring existing blur operators from a given dataset into a new domain. Finally, we provide experimental results to confirm the effectiveness of the proposed method. The code is available at https://github.com/VinAIResearch/blur-kernelspace-exploring. Anh Tuan Tran 0001, Quynh Phung, Minh Hoai |
CVPR | 2 |
| 2021 | Toward Realistic Single-View 3D Object Reconstruction with Unsupervised Learning from Multiple ImagesabstractRecovering the 3D structure of an object from a single image is a challenging task due to its ill-posed nature. One approach is to utilize the plentiful photos of the same object category to learn a strong 3D shape prior for the object. This approach has successfully been demonstrated by a recent work of Wu et al. (2020), which obtained impressive 3D reconstruction networks with unsupervised learning. However, their algorithm is only applicable to symmetric objects. In this paper, we eliminate the symmetry requirement with a novel unsupervised algorithm that can learn a 3D reconstruction network from a multi-image dataset. Our algorithm is more general and covers the symmetry-required scenario as a special case. Besides, we employ a novel albedo loss that improves the reconstructed details and realisticity. Our method surpasses the previous work in both quality and robustness, as shown in experiments on datasets of various structures, including single-view, multiview, image-collection, and video sets. Code is available at: https://github.com/VinAIResearch/LeMul. Long-Nhat Ho, Anh Tuan Tran 0001, Quynh Phung, Minh Hoai |
ICCV | 2 |
| 2021 | WaNet - Imperceptible Warping-based Backdoor Attack
Anh Tuan Tran 0001 |
ICLR | 2 |
| 2021 | Exploiting Domain-Specific Features to Enhance Domain GeneralizationabstractDomain Generalization (DG) aims to train a model, from multiple observed source domains, in order to perform well on unseen target domains. To obtain the generalization capability, prior DG approaches have focused on extracting domain-invariant information across sources to generalize on target domains, while useful domain-specific information which strongly correlates with labels in individual domains and the generalization to target domains is usually ignored. In this paper, we propose meta-Domain Specific-Domain Invariant (mDSDI) - a novel theoretically sound framework that extends beyond the invariance view to further capture the usefulness of domain-specific information. Our key insight is to disentangle features in the latent space while jointly learning both domain-invariant and domain-specific features in a unified framework. The domain-specific representation is optimized through the meta-learning framework to adapt from source domains, targeting a robust generalization on unseen domains. We empirically show that mDSDI provides competitive results with state-of-the-art techniques in DG. A further ablation study with our generated dataset, Background-Colored-MNIST, confirms the hypothesis that domain-specific is essential, leading to better results when compared with only using domain-invariant. Manh-Ha Bui, Toan Tran 0003, Anh Tuan Tran 0001, Dinh Q. Phung |
NeurIPS | 3 |
| 2021 | On Learning Domain-Invariant Representations for Transfer Learning with Multiple SourcesabstractDomain adaptation (DA) benefits from the rigorous theoretical works that study its insightful characteristics and various aspects, e.g., learning domain-invariant representations and its trade-off. However, it seems not the case for the multiple source DA and domain generalization (DG) settings which are remarkably more complicated and sophisticated due to the involvement of multiple source domains and potential unavailability of target domain during training. In this paper, we develop novel upper-bounds for the target general loss which appeal us to define two kinds of domain-invariant representations. We further study the pros and cons as well as the trade-offs of enforcing learning each domain-invariant representation. Finally, we conduct experiments to inspect the trade-off of these representations for offering practical hints regarding how to use them in practice and explore other interesting properties of our developed theory. Trung Le 0001, Long Vuong, Toan Tran 0003, Anh Tuan Tran 0001, Hung Hai Bui, Dinh Q. Phung |
NeurIPS | 5 |
| 2020 | Input-Aware Dynamic Backdoor AttackabstractIn recent years, neural backdoor attack has been considered to be a potential security threat to deep learning systems. Such systems, while achieving the state-of-the-art performance on clean data, perform abnormally on inputs with predefined triggers. Current backdoor techniques, however, rely on uniform trigger patterns, which are easily detected and mitigated by current defense methods. In this work, we propose a novel backdoor attack technique in which the triggers vary from input to input. To achieve this goal, we implement an input-aware trigger generator driven by diversity loss. A novel cross-trigger test is applied to enforce trigger nonreusablity, making backdoor verification impossible. Experiments show that our method is efficient in various attack scenarios as well as multiple datasets. We further demonstrate that our backdoor can bypass the state of the art defense methods. An analysis with a famous neural network inspector again proves the stealthiness of the proposed attack. Our code is publicly available. Anh Tuan Tran 0001 |
NeurIPS | 2 |
| 2019 | Transferability and Hardness of Supervised Classification TasksabstractWe propose a novel approach for estimating the difficulty and transferability of supervised classification tasks. Unlike previous work, our approach is solution agnostic and does not require or assume trained models. Instead, we estimate these values using an information theoretic approach: treating training labels as random variables and exploring their statistics. When transferring from a source to a target task, we consider the conditional entropy between two such variables (i.e., label assignments of the two tasks). We show analytically and empirically that this value is related to the loss of the transferred model. We further show how to use this value to estimate task hardness. We test our claims extensively on three large scale data sets - CelebA (40 tasks), Animals with Attributes 2 (85 tasks), and Caltech-UCSD Birds 200 (312 tasks) - together representing 437 classification tasks. We provide results showing that our hardness and transferability estimates are strongly correlated with empirical hardness and transferability. As a case study, we transfer a learned face recognition model to CelebA attribute classification tasks, showing state of the art accuracy for tasks estimated to be highly transferable. Anh Tuan Tran 0001, Viet Cuong Nguyen, Tal Hassner |
ICCV | 1 |
| 2019 | Deep, Landmark-Free FAME: Face Alignment, Modeling, and Expression Estimation
Feng-Ju Chang, Anh Tuan Tran 0001, Tal Hassner, Iacopo Masi, Ramakant Nevatia, Gérard G. Medioni |
Int. J. Comput. Vis. | 2 |
| 2019 | Face-Specific Data Augmentation for Unconstrained Face Recognition
Iacopo Masi, Anh Tuan Tran 0001, Tal Hassner, Gozde Sahin, Gérard G. Medioni |
Int. J. Comput. Vis. | 2 |
| 2018 | Extreme 3D Face Reconstruction: Seeing Through OcclusionsabstractExisting single view, 3D face reconstruction methods can produce beautifully detailed 3D results, but typically only for near frontal, unobstructed viewpoints. We describe a system designed to provide detailed 3D reconstructions of faces viewed under extreme conditions, out of plane rotations, and occlusions. Motivated by the concept of bump mapping, we propose a layered approach which decouples estimation of a global shape from its mid-level details (e.g., wrinkles). We estimate a coarse 3D face shape which acts as a foundation and then separately layer this foundation with details represented by a bump map. We show how a deep convolutional encoder-decoder can be used to estimate such bump maps. We further show how this approach naturally extends to generate plausible details for occluded facial regions. We test our approach and its components extensively, quantitatively demonstrating the invariance of our estimated facial details. We further provide numerous qualitative examples showing that our method produces detailed 3D face shapes in viewing conditions where existing state of the art often break down. Anh Tuan Tran 0001, Tal Hassner, Iacopo Masi, Eran Paz, Yuval Nirkin, Gérard G. Medioni |
CVPR | 1 |
| 2018 | ExpNet: Landmark-Free, Deep, 3D Facial ExpressionsabstractWe describe a deep learning based method for estimating 3D facial expression coefficients. Unlike previous work, our process does not relay on facial landmark detection methods as a proxy step. Recent methods have shown that a CNN can be trained to regress accurate and discriminative 3D morphable model (3DMM) representations, directly from image intensities. By foregoing landmark detection, these methods were able to estimate shapes for occluded faces appearing in unprecedented viewing conditions. We build on those methods by showing that facial expressions can also be estimated by a robust, deep, landmark-free approach. Our ExpNet CNN is applied directly to the intensities of a face image and regresses a 29D vector of 3D expression coefficients. We propose a unique method for collecting data to train our network, leveraging on the robustness of deep networks to training label noise. We further offer a novel means of evaluating the accuracy of estimated expression coefficients: by measuring how well they capture facial emotions on the CK+ and EmotiW-17 emotion recognition benchmarks. We show that our ExpNet produces expression coefficients which better discriminate between facial emotions than those obtained using state of the art, facial landmark detectors. Moreover, this advantage grows as image scales drop, demonstrating that our ExpNet is more robust to scale changes than landmark detectors. Finally, our ExpNet is orders of magnitude faster than its alternatives. Feng-Ju Chang, Anh Tuan Tran 0001, Tal Hassner, Iacopo Masi, Ramakant Nevatia, Gérard G. Medioni |
FG | 2 |
| 2018 | On Face Segmentation, Face Swapping, and Face PerceptionabstractWe show that even when face images are unconstrained and arbitrarily paired, face swapping between them is quite simple. To this end, we make the following contributions. (a) Instead of tailoring systems for face segmentation, as others previously proposed, we show that a standard fully convolutional network (FCN) can achieve remarkably fast and accurate segmentations, provided that it is trained on a rich enough example set. For this purpose, we describe novel data collection and generation routines which provide challenging segmented face examples. (b) We use our segmentations for robust face swapping under unprecedented conditions. (c) Unlike previous work, our swapping is robust enough to allow for extensive quantitative tests. To this end, we use the Labeled Faces in the Wild (LFW) benchmark and measure the effect of intra- and inter-subject face swapping on recognition. We show that our intra-subject swapped faces remain as recognizable as their sources, testifying to the effectiveness of our method. In line with established perceptual studies, we show that better face swapping produces less recognizable inter-subject results. This is the first time this effect was quantitatively demonstrated by machine vision systems. Yuval Nirkin, Iacopo Masi, Anh Tuan Tran 0001, Tal Hassner, Gérard G. Medioni |
FG | 3 |
| 2017 | Regressing Robust and Discriminative 3D Morphable Models with a Very Deep Neural NetworkabstractThe 3D shapes of faces are well known to be discriminative. Yet despite this, they are rarely used for face recognition and always under controlled viewing conditions. We claim that this is a symptom of a serious but often overlooked problem with existing methods for single view 3D face reconstruction: when applied in the wild, their 3D estimates are either unstable and change for different photos of the same subject or they are over-regularized and generic. In response, we describe a robust method for regressing discriminative 3D morphable face models (3DMM). We use a convolutional neural network (CNN) to regress 3DMM shape and texture parameters directly from an input photo. We overcome the shortage of training data required for this purpose by offering a method for generating huge numbers of labeled examples. The 3D estimates produced by our CNN surpass state of the art accuracy on the MICC data set. Coupled with a 3D-3D face matching pipeline, we show the first competitive face recognition results on the LFW, YTF and IJB-A benchmarks using 3D face shapes as representations, rather than the opaque deep feature vectors used by other modern systems. Anh Tuan Tran 0001, Tal Hassner, Iacopo Masi, Gérard G. Medioni |
CVPR | 1 |
| 2017 | Rapid Synthesis of Massive Face Sets for Improved Face RecognitionabstractRecent work demonstrated that computer graphics techniques can be used to improve face recognition performances by synthesizing multiple new views of faces available in existing face collections. By so doing, more images and more appearance variations are available for training, thereby improving the deep models trained on these images. Similar rendering techniques were also applied at test time to align faces in 3D and reduce appearance variations when comparing faces. These previous results, however, did not consider the computational cost of rendering: At training, rendering millions of face images can be prohibitive; at test time, rendering can quickly become a bottleneck, particularly when multiple images represent a subject. This paper builds on a number of observations which, under certain circumstances, allow rendering new 3D views of faces at a computational cost which is equivalent to simple 2D image warping. We demonstrate this by showing that the run-time of an optimized OpenGL rendering engine is slower than the simple Python implementation we designed for the same purpose. The proposed rendering is used in a face recognition pipeline and tested on the challenging IJB-A and Janus CS2 benchmarks. Our results show that our rendering is not only fast, but improves recognition accuracy. Iacopo Masi, Tal Hassner, Anh Tuan Tran 0001, Gérard G. Medioni |
FG | 3 |
| 2016 | Do We Really Need to Collect Millions of Faces for Effective Face Recognition?
Iacopo Masi, Anh Tuan Tran 0001, Tal Hassner, Jatuporn Toy Leksut, Gérard G. Medioni |
ECCV (5) | 2 |
| 2014 | Real-time 3-D face tracking and modeling framework for mid-res camabstractWe present a robust, real-time 3-D face tracking and modeling system providing accurate 6 degree-of-freedom head pose in the presence of large out-of-plane motion, strong expression changes, and partial occlusions. In this paper, we have extended the previous 3-D face tracking and modeling framework [10] with automatic initialization, reacquisition, and automatic pose correction. Our system first generates a 3-D face model from a single frontal image. We then extract uniformly distributed random points and track them in 2-D. Given these correspondences, the 3-D head pose is robustly estimated using a RANSAC-PnP process. As the head moves, we dynamically add new feature points to handle a large range of poses. A measure of the accumulated error over time allows an auto-correction mechanism to recover from drift when necessary. If the tracker gets lost, due to motion blur or strong occlusions, the system re-initializes. We present live demo results, which shows excellent tracking under large motion (roll: 360°, yaw: ±90°, pitch: −60° to +90°), fast movement, occlusion and facial expression variations. The system runs at 14 fps on a laptop CPU. By experiments on different datasets, our method shows state of the art results. Jongmoo Choi, Anh Tuan Tran 0001, Yann Dumortier, Gérard G. Medioni |
WACV | 2 |