Dung N. Tran

dblp:161/4474 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2024 uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
abstract
Masked Autoencoders (MAEs) learn rich low-level representations from unlabeled data but require substantial labeled data to effectively adapt to downstream tasks. Conversely, Instance Discrimination (ID) emphasizes high-level semantics, offering a potential solution to alleviate annotation requirements in MAEs. Although combining these two approaches can address downstream tasks with limited labeled data, naively integrating ID into MAEs leads to extended training times and high computational costs. To address this challenge, we introduce uaMix-MAE, an efficient ID tuning strategy that leverages unsupervised audio mixtures. Utilizing contrastive tuning, uaMix-MAE aligns the representations of pretrained MAEs, thereby facilitating effective adaptation to task-specific semantics. To optimize the model with small amounts of unlabeled data, we propose an audio mixing technique that manipulates audio samples in both input and virtual label spaces. Experiments in low/few-shot settings demonstrate that uaMix-MAE achieves 4 − 6% accuracy improvements over various benchmarks when tuned with limited unlabeled data, such as AudioSet-20K.
Afrina Tabassum, Dung N. Tran, Trung Dang 0002, Ismini Lourentzou, Kazuhito Koishida
ICASSP2
2024 Learned Image Compression With Text Quality Enhancement
abstract
Learned image compression has gained widespread popularity for their efficiency in achieving ultra-low bit-rates. Yet, images containing substantial textual content, particularly screen-content images (SCI), often suffers from text distortion at such compressed levels. To address this, we propose to minimize a novel text logit loss designed to quantify the disparity in text between the original and reconstructed images, thereby improving the perceptual quality of the reconstructed text. Through rigorous experimentation across diverse datasets and employing state-of-the-art algorithms, our findings reveal significant enhancements in the quality of reconstructed text upon integration of the proposed loss function with appropriate weighting. Notably, we achieve a Bjontegaard delta (BD) rate of $-32.64 \%$ for Character Error Rate (CER) and $-28.03 \%$ for Word Error Rate (WER) on average by applying the text logit loss for two screenshot datasets. Additionally, we present quantitative metrics tailored for evaluating text quality in image compression tasks. Our findings underscore the efficacy and potential applicability of our proposed text logit loss function across various text-aware image compression contexts.
Chih-Yu Lai, Dung N. Tran, Kazuhito Koishida
ICIP2
2024 LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
Trung Dang 0002, David Aponte, Dung N. Tran, Kazuhito Koishida
INTERSPEECH3
2024 ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
Yatong Bai, Trung Dang 0002, Dung N. Tran, Kazuhito Koishida, Somayeh Sojoudi
INTERSPEECH3
2022 Training Robust Zero-Shot Voice Conversion Models with Self-Supervised Features
abstract
Unsupervised Zero-Shot Voice Conversion (VC) aims to modify the speaker characteristic of an utterance to match an unseen target speaker without relying on parallel training data. Recently, self-supervised learning of speech representation has been shown to produce useful linguistic units without using transcripts, which can be directly passed to a VC model. In this paper, we showed that high-quality audio samples can be achieved by using a length resampling decoder, which enables the VC model to work in conjunction with different linguistic feature extractors and vocoders without requiring them to operate on the same sequence length. We showed that our method can outperform many baselines on the VCTK dataset. Without modifying the architecture, we further demonstrated that a) using pairs of different audio segments from the same speaker, b) adding a cycle consistency loss, and c) adding a speaker classification loss can help to learn a better speaker embedding. Our model trained on LibriTTS using these techniques achieves the best performance, producing audio samples transferred well to the target speaker’s voice, while preserving the linguistic content that is comparable with actual human utterances in terms of Character Error Rate.
Trung Dang 0002, Dung N. Tran, Sang (Peter) Chin, Kazuhito Koishida
ICASSP2
2021 Single-Channel Speech Enhancement Using Learnable Loss Mixup
Oscar Chang, Dung N. Tran, Kazuhito Koishida
Interspeech2
2020 Robust Pitch Regression with Voiced/Unvoiced Classification in Nonstationary Noise Environments
Dung N. Tran, Uros Batricevic, Kazuhito Koishida
INTERSPEECH1
2020 Single-Channel Speech Enhancement by Subspace Affinity Minimization
Dung N. Tran, Kazuhito Koishida
INTERSPEECH1
2018 A Greedy Pursuit Algorithm for Separating Signals from Nonlinear Compressive Observations
abstract
In this paper we study the unmixing problem which aims to separate a set of structured signals from their superposition. In this paper, we consider the scenario in which the mixture is observed via nonlinear compressive measurements. We present a fast, robust, greedy algorithm called Unmixing Matching Pursuit (UnmixMP) to solve this problem. We prove rigorously that the algorithm can recover the constituents from their noisy nonlinear compressive measurements with arbitrarily small error. We compare our algorithm to the Demixing with Hard Thresholding (DHT) algorithm [1], in a number of experiments on synthetic and real data.
Sang (Peter) Chin, Trac D. Tran, Dung N. Tran, Akshay Rangamani
ICASSP3
2017 A provable nonconvex model for factoring nonnegative matrices
abstract
We study the Nonnegative Matrix Factorization problem which approximates a nonnegative matrix by a low-rank factorization. This problem is particularly important in Machine Learning, and finds itself in a large number of applications. Unfortunately, the original formulation is ill-posed and NP-hard. In this paper, we propose a row sparse model based on Row Entropy Minimization to solve the NMF problem under separable assumption which states that each data point is a convex combination of a few distinct data columns. We utilize the concentration of the entropy function and the ℓ∞norm to concentrate the energy on the least number of latent variables. We prove that under the separability assumption, our proposed model robustly recovers data columns that generate the dataset, even when the data is corrupted by noise. We empirically justify the robustness of the proposed model and show that it is significantly more robust than the state-of-the-art separable NMF algorithms.
Dung N. Tran, Sang (Peter) Chin, Trac D. Tran
ICASSP1
2016 Low-rank matrices recovery via entropy function
abstract
The low-rank matrix recovery problem consists of reconstructing an unknown low-rank matrix from a few linear measurements, possibly corrupted by noise. One of the most popular method in low-rank matrix recovery is based on nuclear-norm minimization, which seeks to simultaneously estimate the most significant singular values of the target low-rank matrix by adding a penalizing term on its nuclear norm. In this paper, we introduce a new method that requires substantially fewer measurements needed for exact matrix recovery compared to nuclear norm minimization. The proposed optimization program utilizes a sparsity promoting regularization in the form of the entropy function of the singular values. Numerical experiments on synthetic and real data demonstrates that the proposed method outperforms stage-of-the-art nuclear norm minimization algorithms.
Dung N. Tran, Sang (Peter) Chin, Trac D. Tran
ICASSP1
2016 Sparse signal recovery based on nonconvex entropy minimization
abstract
We propose a new sparsity-promoting objective function to be used in sparse signal recovery. Specifically, the objective is an entropy function l1 defined on the sparse signal x. Compared to the conventional l1, it is a nonconvex function and the optimization problem can be solved based on the fast iterative shrinkage thresholding algorithm (FISTA). Experiments on 1-dimensional sparse signal recovery and 2-dimensional real image recovery show that minimizing lp favors sparse solutions, and that it could recover sparse signals better than the convex l1 norm minimization and the nonconvex lp-norm minimization.
Dung N. Tran, Trac D. Tran
ICIP2
2015 Nonnegative matrix factorization with gradient vertex pursuit
abstract
Nonnegative Matrix Factorization (NMF), defined as factorizing a nonnegative matrix into two nonnegative factor matrices, is a particularly important problem in machine learning. Unfortunately, it is also ill-posed and NP-hard. We propose a fast, robust, and provably correct algorithm, namely Gradient Vertex Pursuit (GVP), for solving a well-defined instance of the problem which results in a unique solution: there exists a polytope, whose vertices consist of a few columns of the original matrix, covering the entire set of remaining columns. Our algorithm is greedy: it detects, at each iteration, a correct vertex until the entire polytope is identified. We evaluate the proposed algorithm on both synthetic and real hyperspectral data, and show its superior performance compared with other state-of-the-art greedy pursuit algorithms.
Dung N. Tran, Sang (Peter) Chin, Trac D. Tran
ICASSP1
2015 Local sensing with global recovery
abstract
In this paper, we study Locally Compressed Sensing for images, where sampling process is allowed to be performed on arbitrary local regions of the images. We propose a fast and efficient reconstruction algorithm which utilizes local structures of images. Several numerical experiments on real images demonstrates that our algorithm yields better reconstruction quality than existing techniques at much lower computational complexity and memory requirement.
Dung N. Tran, Duyet N. Tran, Sang (Peter) Chin, Trac D. Tran
ICIP1
2015 An unsupervised dictionary learning algorithm for neural recordings
abstract
To meet the growing demand of wireless and power efficient neural recordings systems, we demonstrate an unsupervised dictionary learning algorithm in Compressed Sensing (CS) framework which can be implemented in VLSI systems. Without prior label information of neural spikes, we extend our previous work to unsupervised learning and construct a dictionary with discriminative structures for spike sorting. To further improve the reconstruction and classification performance, we proposed a joint prediction to determine the class of neural spikes in dictionary learning. When the neural spikes is compressed 50 times, our approach can achieve an average gain of 2 dB and 15 percentage units over state-of-the-art of CS approaches in terms of the reconstruction quality and classification accuracy respectively.
Jie Zhang 0063, Yuanming Suo, Dung N. Tran, Ralph Etienne-Cummings, Sang (Peter) Chin, Trac D. Tran
ISCAS4