Bac Nguyen

dblp:190/4847 · DBLP profile ↗
← Back
18ranked-venue papers
16as first author
7since 2021 · last 2025
0000-0001-9193-1908ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 11 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 Improving Vector-Quantized Image Modeling with Latent Consistency-Matching Diffusion
abstract
By embedding discrete representations into a continuous latent space, we can leverage continuous-space latent diffusion models to handle generative modeling of discrete data. However, despite their initial success, most latent diffusion methods rely on fixed pretrained embeddings, limiting the benefits of joint training with the diffusion model. While jointly learning the embedding (via reconstruction loss) and the latent diffusion model (via score matching loss) could enhance performance, end-to-end training risks embedding collapse, degrading generation quality. To mitigate this issue, we introduce VQ-LCMD, a continuous-space latent diffusion framework within the embedding space that stabilizes training. VQ-LCMD uses a novel training objective combining the joint embedding-diffusion variational lower bound with a consistency-matching (CM) loss, alongside a shifted cosine noise schedule and random dropping strategy. Experiments on several benchmarks show that the proposed VQ-LCMD yields superior results on FFHQ, LSUN Churches, and LSUN Bedrooms compared to discrete-state latent diffusion models. In particular, VQ-LCMD achieves an FID of 6.81 for class-conditional image generation on ImageNet with 50 steps.
Bac Nguyen, Chieh-Hsin Lai, Yuhta Takida, Naoki Murata, Toshimitsu Uesaka, Stefano Ermon, Yuki Mitsufuji
IJCNN1
2024 SAFT: Towards Out-of-Distribution Generalization in Fine-Tuning
Bac Nguyen, Stefan Uhlich, Fabien Cardinaux, Lukas Mauch, Marzieh Edraki, Aaron C. Courville
ECCV (69)1
2024 SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision
Ankit Vani, Bac Nguyen, Samuel Lavoie-Marchildon, Ranjay Krishna, Aaron C. Courville
ECCV (66)2
2023 Autotts: End-to-End Text-to-Speech Synthesis Through Differentiable Duration Modeling
abstract
Parallel text-to-speech (TTS) models have recently enabled fast and highly-natural speech synthesis. However, they typically require external alignment models, which are not necessarily optimized for the decoder as they are not jointly trained. In this paper, we propose a differentiable duration method for learning monotonic alignments between input and output sequences. Our method is based on a soft-duration mechanism that optimizes a stochastic process in expectation. Using this differentiable duration method, we introduce AutoTTS, a direct text-to-waveform speech synthesis model. AutoTTS enables high-fidelity speech synthesis through a combination of adversarial training and matching the total ground-truth duration. Experimental results show that our model obtains competitive results while enjoying a much simpler training pipeline. Audio samples are available online1.
Bac Nguyen, Fabien Cardinaux, Stefan Uhlich
ICASSP1
2023 Improving Self-Supervised Learning for Audio Representations by Feature Diversity and Decorrelation
abstract
Self-supervised learning (SSL) has recently shown remarkable results in closing the gap between supervised and unsupervised learning. The idea is to learn robust features that are invariant to distortions of the input data. Despite its success, this idea can suffer from a collapsing issue where the network produces a constant representation. To this end, we introduce SELFIE, a novel Self-supervised Learning approach for audio representation via Feature Diversity and Decorrelation. SELFIE avoids the collapsing issue by ensuring that the representation (i) maintains a high diversity among embeddings and (ii) decorrelates the dependencies between dimensions. SELFIE is pre-trained on the large-scale AudioSet dataset and its embeddings are validated on nine audio downstream tasks, including speech, music, and sound event recognition. Experimental results show that SELFIE outperforms existing SSL methods in several tasks.
Bac Nguyen, Stefan Uhlich, Fabien Cardinaux
ICASSP1
2023 Towards Robust FastSpeech 2 by Modelling Residual Multimodality
Fabian Kögel, Bac Nguyen, Fabien Cardinaux
INTERSPEECH2
2022 NVC-Net: End-To-End Adversarial Voice Conversion
abstract
Voice conversion (VC) has gained increasing popularity in many speech synthesis applications. The idea is to change the voice identity from one speaker into another while keeping the linguistic content unchanged. Many VC approaches rely on the use of a vocoder to reconstruct the speech from acoustic features, and as a consequence, the speech quality heavily depends on such a vocoder. In this paper, we propose NVC-Net, an end-to-end adversarial network, which performs VC directly on the raw audio waveform. By disentangling the speaker identity from the speech content, NVC-Net is able to perform non-parallel traditional many-to-many VC as well as zero-shot VC from a short utterance of an unseen target speaker. Importantly, NVC-Net is non-autoregressive and fully convolutional, achieving fast inference. Objective and subjective evaluations on VC tasks show that NVC-Net obtains competitive results with significantly fewer parameters.
Bac Nguyen, Fabien Cardinaux
ICASSP1
2020 Improved deep embedding learning based on stochastic symmetric triplet loss and local sampling
Bac Nguyen, Bernard De Baets
Neurocomputing1
2020 Scalable Large-Margin Distance Metric Learning Using Stochastic Gradient Descent
abstract
The key to success of many machine learning and pattern recognition algorithms is the way of computing distances between the input data. In this paper, we propose a large-margin-based approach, called the large-margin distance metric learning (LMDML), for learning a Mahalanobis distance metric. LMDML employs the principle of margin maximization to learn the distance metric with the goal of improving k -nearest-neighbor classification. The main challenge of distance metric learning is the positive semidefiniteness constraint on the Mahalanobis matrix. Semidefinite programming is commonly used to enforce this constraint, but it becomes computationally intractable on large-scale data sets. To overcome this limitation, we develop an efficient algorithm based on a stochastic gradient descent. Our algorithm can avoid the computations of the full gradient and ensure that the learned matrix remains within the positive semidefinite cone after each iteration. Extensive experiments show that the proposed algorithm is scalable to large data sets and outperforms other state-of-the-art distance metric learning approaches regarding classification accuracy and training time.
Bac Nguyen, Carlos Morell 0001, Bernard De Baets
IEEE Trans. Cybern.1
2019 An efficient method for clustered multi-metric learning
Bac Nguyen, Francesc J. Ferri, Carlos Morell 0001, Bernard De Baets
Inf. Sci.1
2019 Kernel Distance Metric Learning Using Pairwise Constraints for Person Re-Identification
abstract
Person re-identification is a fundamental task in many computer vision and image understanding systems. Due to appearance variations from different camera views, person re-identification still poses an important challenge. In the literature, KISSME has already been introduced as an effective distance metric learning method using pairwise constraints to improve the re-identification performance. Computationally, it only requires two inverse covariance matrix estimations. However, the linear transformation induced by KISSME is not powerful enough for more complex problems. We show that KISSME can be kernelized, resulting in a nonlinear transformation, which is suitable for many real-world applications. Moreover, the proposed kernel method can be used for learning distance metrics from structured objects without having a vectorial representation. The effectiveness of our method is validated on five publicly available data sets. To further apply the proposed kernel method efficiently when data are collected sequentially, we introduce a fast incremental version that learns a dissimilarity function in the feature space without estimating the inverse covariance matrices. The experiments show that the latter variant can obtain competitive results in a computationally efficient manner.
Bac Nguyen, Bernard De Baets
IEEE Trans. Image Process.1
2019 Kernel-Based Distance Metric Learning for Supervised k-Means Clustering
abstract
Finding an appropriate distance metric that accurately reflects the (dis)similarity between examples is a key to the success of k -means clustering. While it is not always an easy task to specify a good distance metric, we can try to learn one based on prior knowledge from some available clustered data sets, an approach that is referred to as supervised clustering. In this paper, a kernel-based distance metric learning method is developed to improve the practical use of k -means clustering. Given the corresponding optimization problem, we derive a meaningful Lagrange dual formulation and introduce an efficient algorithm in order to reduce the training complexity. Our formulation is simple to implement, allowing a large-scale distance metric learning problem to be solved in a computationally tractable way. Experimental results show that the proposed method yields more robust and better performances on synthetic as well as real-world data sets compared to other state-of-the-art distance metric learning methods.
Bac Nguyen, Bernard De Baets
IEEE Trans. Neural Networks Learn. Syst.1
2018 Distance metric learning for ordinal classification based on triplet constraints
Bac Nguyen, Carlos Morell 0001, Bernard De Baets
Knowl. Based Syst.1
2018 An approach to supervised distance metric learning based on difference of convex functions programming
Bac Nguyen, Bernard De Baets
Pattern Recognit.1
2017 Distance metric learning: a two-phase approach
Bac Nguyen, Carlos Morell 0001, Bernard De Baets
ESANN1
2017 Supervised distance metric learning through maximization of the Jeffrey divergence
Bac Nguyen, Carlos Morell 0001, Bernard De Baets
Pattern Recognit.1
2017 Distance metric learning with the Universum
Bac Nguyen, Carlos Morell 0001, Bernard De Baets
Pattern Recognit. Lett.1
2016 Large-scale distance metric learning for k-nearest neighbors regression
Bac Nguyen, Carlos Morell 0001, Bernard De Baets
Neurocomputing1