Hieu Duy Nguyen

dblp:20/11471 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Adaptive Estimation and Learning under Temporal Distribution Shift
abstract
In this paper, we study the problem of estimation and learning under temporal distribution shift. Consider an observation sequence of length $n$, which is a noisy realization of a time-varying ground-truth sequence. Our focus is to develop methods to estimate the groundtruth at the final time-step while providing sharp point-wise estimation error rates. We show that, without prior knowledge on the level of temporal shift, a wavelet soft-thresholding estimator provides an optimal estimation error bound for the groundtruth. Our proposed estimation method generalizes existing researches (Mazetto and Upfal, 2023) by establishing a connection between the sequence’s non-stationarity level and the sparsity in the wavelet-transformed domain. Our theoretical findings are validated by numerical experiments. Additionally, we applied the estimator to derive sparsity-aware excess risk bounds for binary classification under distribution shift and to develop computationally efficient training objectives. As a final contribution, we draw parallels between our results and the classical signal processing problem of total-variation denoising (Mammen and van de Geer 1997; Tibshirani 2014 ), uncovering novel optimal algorithms for such task.
Dheeraj Baby, Yifei Tang, Hieu Duy Nguyen, Yu-Xiang Wang 0003, Rohit Pyati
ICML3
2024 Max-Margin Transducer Loss: Improving Sequence-Discriminative Training Using a Large-Margin Learning Strategy
abstract
In this work, we propose a novel sequence-discriminative training criterion for automatic speech recognition (ASR) based on the Conformer Transducer. Inspired by the large-margin classifier framework, we separate the "good" and the "bad" hypotheses in an N-best list produced from a pre-trained transducer model by a margin (τ), hence the term, Max-Margin Transducer (MMT) loss. It is observed that fine-tuning with the proposed loss achieves significant improvement over baseline transducer loss but does not outperform the state-of-the-art minimum word error rate (MWER) training. However, combining the proposed MMT loss with MWER surpasses the performance of either losses suggesting the complimentary nature of MWER and MMT losses. With the combined losses, we obtained 7.44% and 7.68% relative WER improvements on Librispeech test-clean and test-other sets, respectively, and up to 8.9% relative improvement on Multi-lingual Librispeech test sets.
Rupak Vignesh Swaminathan, Grant P. Strimel, Ariya Rastrow, Sri Harish Reddy Mallidi, Kai Zhen, Hieu Duy Nguyen, Nathan Susanj, Athanasios Mouchtaris
ICASSP6
2022 Sub-8-Bit Quantization Aware Training for 8-Bit Neural Network Accelerator with On-Device Speech Recognition
Kai Zhen, Hieu Duy Nguyen, Raviteja Chinta, Nathan Susanj, Athanasios Mouchtaris, Tariq Afzal, Ariya Rastrow
INTERSPEECH2
2022 Accelerator-Aware Training for Transducer-Based Speech Recognition
abstract
Machine learning model weights and activations are represented in full-precision during training. This leads to performance degradation in runtime when deployed on neural network accelerator (NNA) chips, which leverage highly parallelized fixed-point arithmetic to improve runtime memory and latency. In this work, we replicate the NNA operators during the training phase, accounting for the degradation due to low-precision inference on the NNA in back-propagation. Our proposed method efficiently emulates NNA operations, thus foregoing the need to transfer quantization error-prone data to the Central Processing Unit (CPU), ultimately reducing the user perceived latency (UPL). We apply our approach to Recurrent Neural Network-Transducer (RNN-T), an attractive architecture for on-device streaming speech recognition tasks. We train and evaluate models on 270K hours of English data and show a 5-7% improvement in engine latency while saving up to 10% relative degradation in WER.
Suhaila M. Shakiah, Rupak Vignesh Swaminathan, Hieu Duy Nguyen, Raviteja Chinta, Tariq Afzal, Nathan Susanj, Athanasios Mouchtaris, Grant P. Strimel, Ariya Rastrow
SLT3
2022 Sub-8-Bit Quantization for On-Device Speech Recognition: A Regularization-Free Approach
abstract
For on-device automatic speech recognition (ASR), quantization aware training (QAT) is ubiquitous to achieve the trade-off between model predictive performance and efficiency. Among existing QAT methods, one major drawback is that the quantization centroids have to be predetermined and fixed. To overcome this limitation, we introduce a regularization-free, “soft-to-hard” compression mechanism with self-adjustable centroids in a$\mu$-Law constrained space, resulting in a simpler yet more versatile quantization scheme, called General Quantizer (GQ). We apply GQ to ASR tasks using Recurrent Neural Network Transducer (RNN-T) and Conformer architectures on both LibriSpeech and de-identified far-field datasets. Without accuracy degradation, GQ can compress both RNN-T and Conformer into sub-8-bit, and for some RNN-T layers, to 1-bit for fast and accurate inference. We observe a 30.73% memory footprint saving and 31.75% user-perceived latency reduction compared to 8-bit QAT via physical device benchmarking.
Kai Zhen, Martin Radfar, Hieu Duy Nguyen, Grant P. Strimel, Nathan Susanj, Athanasios Mouchtaris
SLT3
2021 Sparsification via Compressed Sensing for Automatic Speech Recognition
abstract
In order to achieve high accuracy for machine learning (ML) applications, it is essential to employ models with a large number of parameters. Certain applications, such as Automatic Speech Recognition (ASR), however, require real-time interactions with users, hence compelling the model to have as low latency as possible. Deploying large scale ML applications thus necessitates model quantization and compression, especially when running ML models on resource constrained devices. For example, by forcing some of the model weight values into zero, it is possible to apply zero-weight compression, which reduces both the model size and model reading time from the memory. In the literature, such methods are referred to as sparse pruning. The fundamental questions are when and which weights should be forced to zero, i.e. be pruned. In this work, we propose a compressed sensing based pruning (CSP) approach to effectively address those questions. By reformulating sparse pruning as a sparsity inducing and compression-error reduction dual problem, we introduce the classic compressed sensing process into the ML model training process. Using ASR task as an example, we show that CSP consistently outperforms existing approaches in the literature.
Kai Zhen, Hieu Duy Nguyen, Feng-Ju Chang, Athanasios Mouchtaris, Ariya Rastrow
ICASSP2
2020 Multilingual Grapheme-To-Phoneme Conversion with Byte Representation
abstract
Grapheme-to-phoneme (G2P) models convert a written word into its corresponding pronunciation and are essential components in automatic-speech-recognition and text-to-speech systems. Recently, the use of neural encoder-decoder architectures has substantially improved G2P accuracy for mono- and multi-lingual cases. However, most multilingual G2P studies focus on sets of languages that share similar graphemes, such as European languages. Multilingual G2P for languages from different writing systems, e.g. European and East Asian, remains an understudied area. In this work, we propose a multilingual G2P model with byte-level input representation to accommodate different grapheme systems, along with an attention-based Transformer architecture. We evaluate the performance of both character-level and byte-level G2P using data from multiple European and East Asian locales. Models using byte representation yield 16.2%– 50.2% relative word error rate improvement over character-based counterparts for mono- and multi-lingual use cases. In addition, byte-level models are 15.0%–20.1% smaller in size. Our results show that byte is an efficient representation for multilingual G2P with languages having large grapheme vocabularies.
Mingzhi Yu, Hieu Duy Nguyen, Alex Sokolov, Jack Lepird, Kanthashree Mysore Sathyendra, Samridhi Choudhary, Athanasios Mouchtaris, Siegfried Kunzmann
ICASSP2
2020 Quantization Aware Training with Absolute-Cosine Regularization for Automatic Speech Recognition
Hieu Duy Nguyen, Anastasios Alexandridis, Athanasios Mouchtaris
INTERSPEECH1
2017 A High Bit-Rate Shared Key Generator with Time-Frequency Features of Wireless Channels
abstract
Although pre-shared key schemes are popularly used in securing communication systems, they are impractical in some applications such as ad-hoc communications. This paper proposes a new method for secret key generation between wireless endpoints. In contrast to previous studies, the present method utilizes a joint time-frequency multiscale features of wireless channels, thus yields higher bit-rate and/or larger secret key, as demonstrated in the numerical evaluations and experiments.
Sumei Sun, Yongdong Wu, Boon Shyang Lim, Hieu Duy Nguyen
GLOBECOM4
2017 Closed-Form Performance Bounds for Stochastic Geometry-Based Cellular Networks
abstract
In this paper, we study the performance of partial-fading and Rayleigh fading wireless networks using stochastic geometry. The aim is to provide closed-form bounds for the signal-to-interference-plus-noise ratio (SINR) distribution, average Shannon rate, and outage rate. We first characterize the SINR distribution of partial-fading channels through the Laplace transform of the inverted SINR. Since most communication systems are interference limited, we also consider the case of negligible noise power, and derive the upper and lower bounds for the signal-to-interference ratio distribution under both partial fading and fading cases. These bounds are of closed forms and thus more convenient for theoretical analysis. Based on these derivations, we obtain closed-form bounds for the average Shannon and outage rates. These results are useful for investigating the fifth-generation communication systems, for example massive multi-antenna networks as described in our illustrative example.
Hieu Duy Nguyen, Sumei Sun
IEEE Trans. Wirel. Commun.1
2016 Massive MIMO versus small-cell systems: Spectral and energy efficiency comparison
abstract
In this paper, we study the downlink performance of two important 5G network architectures, i.e. massive multiple-input multiple-output (M-MIMO) and small-cell densification. We propose a comparative modeling for the two systems, where the user and antenna/base station (BS) locations are distributed according to Poisson point processes (PPPs). We then study the SIR distribution and the outage rate of each network. By comparing these results, we observe that for user-average spectral efficiency, small-cell densification is favorable in crowded areas with moderate to high user density and M-MIMO with low user density. However, small-cell systems outperform M-MIMO in all cases when the performance metric is the energy efficiency. The results of this paper are useful for the optimal design of practical 5G networks.
Hieu Duy Nguyen, Sumei Sun
ICC1
2016 Fronthaul compression and optimization for cloud radio access networks
abstract
In the present paper, we investigate the design and optimization for fronhaul links in cloud radio access networks (C-RAN). Existing C-RAN designs rely on the instantaneous network-wide channel state information (CSI), which might impose a significant overhead due to the potential large-scale of C-RAN. To overcome this limitation, we optimize C-RAN based on the average performance metrics which only require the second-order statistics of the fading channels. Firstly, a tight upper bound of the block error rate (BLER) over Rayleigh fading channels is derived in closed-form expression, through which some insights on C-RAN are drawn: i) full diversity order, which is equal to the number of RRHs, is achievable with respect to the signal to compression plus noise ratio; and ii) the BLER is limited below by either compression or Gaussian noises. Secondly, based on the derived bound, a compression optimization is proposed to minimize the fronthaul transmission rate while satisfying some predefined BLER constraints. The premise of the proposed optimization originates from practical scenarios where most applications tolerate a non-zero BLER. Finally, a fronthaul rate allocation scheme is proposed to minimize the system BLER. It is proved that the proposed allocation scheme, which imposes uniform compression noise across the RRHs, approaches the optimal allocation as the total fronthauls' bandwidth increases.
Thang X. Vu, Hieu Duy Nguyen, Tony Q. S. Quek, Sumei Sun
ICC2
2015 Joint Decoding and Adaptive Compression with QoS Constraint for Uplinks in Cloud Radio Access Networks
abstract
Cloud Radio Access Network (C-RAN) is a promising candidate for future mobile networks to sustain the exponentially increasing demand for data rate. The centralized architecture enables C-RAN to exploit multi-cell cooperation and interference management effectively. In C-RAN, one baseband unit (BBU) communicates with users through distributed Remote Radio Heads (RRHs) which are connected to the BBU via high capacity, low latency fronthaul links and perform ``soft" relaying. However, the architecture of C-RAN imposes a shortage of fronthaul bandwidth because raw In-phase/Quadrature-phase (I/Q) samples are exchanged between the RRHs and the BBU. In this paper, we leverage on advanced signal processing to improve the compression efficiency in fronthaul uplinks. Specifically, we propose a joint decoding algorithm at the BBU that exploits the correlation among the RRHs and jointly performs decompressing and decoding. An upper bound of the Block Error Rate (BLER) of the proposed algorithm is derived using pair-wise error probability analysis. Based on the BLER upper bound, we propose an adaptive compression scheme which minimizes the fronthaul transmission rate while satisfying a target quality of service constrain on the BLER. Our proposed adaptive compressor originates from practical scenarios in which most applications tolerate certain non-zero BLER thresholds.
Thang X. Vu, Tony Q. S. Quek, Hieu Duy Nguyen
GLOBECOM3
2015 Precoder design for distributed antenna systems (DAS) with limited channel state information
abstract
A distributed antenna system (DAS) consists of multiple baseband units (BBUs) connecting to distributed antennas (DAs) via dedicated access links. In this study, we investigate a DAS with limited channel state information (CSI) and consider an average rate of users as an objective, where the expectation is taken over the channel uncertainty. We propose two distributed precoder designs that are based on a rate lower bound and a rate upper bound, respectively. As a benchmark, coordinated precoder and cooperative dirty paper coding (DPC)-based precoder with full CSI are compared with our proposed algorithms. Numerical results verifies that the rate performance of our upper-bound based scheme with limited CSI approaches tightly the maximum rates of the full CSI schemes, while that of lower-bound based scheme is relatively worse.
Hieu Duy Nguyen, Jingon Joung, Sumei Sun
ICC1
2015 Error probability minimization for MIMO systems with imperfect channel state information
abstract
Channel uncertainty degrades the performance of multiple-input multiple-output (MIMO) systems considerably. In the literature, capacity maximization for MIMO channels with imperfect channel state information (CSI) has been extensively investigated. However, the error probability minimization counterpart is less studied. In this paper, we aim to minimize the transmission error probability of MIMO systems with imperfect channel estimate and error covariance matrix available at the transmitter. We propose a precoder design which minimizes the maximum pair-wise error probability among every symbol pair. Compared with other schemes utilizing indirect alternatives, i.e., mean-square error (MSE) or approximated signal-to-noise ratio (SNR), numerical results show that the proposed design achieves a significant improvement in error performance.
Hieu Duy Nguyen, Boon Sim Thian, Sumei Sun
ICC1
2015 Improper Signaling for Symbol Error Rate Minimization in K-User Interference Channel
abstract
The rate maximization for the K-user interference channels (ICs) has been investigated extensively in the literature. However, the practical problem of minimizing the error probability with given signal modulations and/or data rates of the users is less studied. In this paper, by utilizing additional degrees of freedom from the improper signaling (versus the conventional proper signaling) , we seek to optimize the precoding matrices for the K-user single-input single-output (SISO) ICs to minimize pair-wise error probability (PEP) and symbol error rate (SER) with two proposed algorithms, respectively. Compared with conventional proper signaling and other state-of-the-art improper signaling designs, our proposed improper signaling schemes achieve notable error rate improvement in SISO-ICs under both the additive white Gaussian noise (AWGN) and cellular system setups with or without channel coding. Our study provides another viewpoint for optimizing transmissions in ICs and further justifies the practical benefit of improper signaling in interference-limited communication systems.
Hieu Duy Nguyen, Rui Zhang 0006, Sumei Sun
IEEE Trans. Commun.1
2015 Adaptive Compression and Joint Detection for Fronthaul Uplinks in Cloud Radio Access Networks
abstract
Cloud radio access network (C-RAN) has recently attracted much attention as a promising architecture for future mobile networks to sustain the exponential growth of data rate. In C-RAN, one data processing center or baseband unit (BBU) communicates with users via distributed remote radio heads (RRHs), which are connected to the BBU via high capacity, low latency fronthaul links. In this paper, we study the compression on fronthaul uplinks and propose a joint decompression algorithm at the BBU. The central premise behind the proposed algorithm is to exploit the correlation between RRHs. Our contribution is threefold. First, we propose a joint decompression and detection (JDD) algorithm which jointly performs decompressing and detecting. The JDD algorithm takes into consideration both the fading and compression effect in a single decoding step. Second, block error rate (BLER) of the proposed algorithm is analyzed in closed-form by using pair-wise error probability analysis. Third, based on the analyzed BLER, we propose adaptive compression schemes subject to quality of service (QoS) constraints to minimize the fronthaul transmission rate while satisfying the pre-defined target QoS. As a dual problem, we also propose a scheme to minimize the signal distortion subject to fronthaul rate constraint. Numerical results demonstrate that the proposed adaptive compression schemes can achieve a compression ratio of 300% in experimental setups.
Thang X. Vu, Hieu Duy Nguyen, Tony Q. S. Quek
IEEE Trans. Commun.2
2014 On design of improper signaling for ser minimization in K-user interference channel
abstract
The rate maximization for the K-user interference channels (ICs) has been investigated extensively in the literature. However, the dual problem of minimizing the error probability with given signal constellations and/or data rates of the users is less exploited. In this paper, by utilizing the additional degrees of freedom attained from the improper signaling (versus the conventional proper signaling), we optimize the precoding matrices for the K-user single-input single-output (SISO) ICs to achieve minimal transmission symbol error rate (SER). Compared to conventional proper signaling as well as other state-of-the-art improper signaling designs, our proposed improper signaling scheme is shown to achieve notable SER improvement in SISO-ICs by simulations. Our study provides another viewpoint for optimizing transmissions in ICs and further justifies the practical benefit of improper signaling in interference-limited communication systems.
Hieu Duy Nguyen, Rui Zhang 0006, Sumei Sun
GLOBECOM1
2014 Linear precoder for codeword error minimization in MIMO systems with channel estimation errors
abstract
We propose a precoder design to minimize the codeword error rate of multiple-input multiple-output (MIMO) systems in the presence of channel estimation errors. Our proposed scheme only requires knowledge of the second-order statistics of the channel, noise and the estimation errors; an estimate of the instantaneous channel state information (CSI) is not required. Compared to the instantaneous CSI requirement, our assumption is more practical since obtaining an accurate CSI is challenging if the channel fluctuates rapidly, while the channel statistics are likely to remain unchanged for a much longer period. Furthermore, the feedback overhead from the receiver to the transmitter is greatly reduced. When compared to the case without transmit preceding, our proposed design achieves a significant improvement in error performance: (i) for a real 2 × 2 MIMO system with 4-PAM modulation and at codeword error rate of 104, our proposed precoder design achieves coding gains of up to 6 dB and (ii) for a real 4 × 4 MIMO system with BPSK modulation and at codeword error rate of 104, coding gains of up to 6.5 dB can be achieved.
Boon Sim Thian, Hieu Duy Nguyen, Sumei Sun
GLOBECOM2
2013 Effect of receive spatial diversity on the degrees of freedom region in multi-cell random beamforming
abstract
The random beamforming (RBF) scheme, together with multi-user diversity based user scheduling, is able to achieve interference-free downlink transmission with only partial channel state information (CSI) at the transmitter. The impact of receive spatial diversity on RBF, however, is not fully characterized even under a single-cell setup. In this paper, we study a multi-cell multiple-input multiple-output (MIMO) broadcast system with RBF applied at each base station and either the minimum-meansquare-error (MMSE), matched filter (MF), or antenna selection (AS) based spatial receiver applied at each mobile terminal. We investigate the effect of different spatial diversity receivers on the achievable sum-rate of the multi-cell RBF system subject to both the intra- and inter-cell interferences. We focus on the high signal-to-noise ratio (SNR) regime and for a tractable analysis assume that the number of users in each cell scales in a certain order with the per-cell SNR. Under this setup, we characterize the degrees of freedom (DoF) region for the multi-cell RBF system, which constitutes all the achievable sum-rate DoF tuples of all the cells. Our results reveal significant sum-rate DoF gains with the MMSE-based spatial receiver as compared to the case without spatial diversity or with suboptimal spatial receivers (MF or AS). This observation is in sharp contrast to the existing result that spatial diversity only yields marginal sum-rate gains based on the conventional asymptotic analysis in the regime of large number of users but with fixed SNR per cell.
Hieu Duy Nguyen, Rui Zhang 0006, Hon Tat Hui
WCNC1
2012 Degrees of freedom region in multi-cell random beamforming
abstract
Random beamforming (RBF) is a practically favorable transmission scheme for multiuser multi-antenna downlink systems. This paper studies the asymptotic rates achievable with RBF in a multi-cell system subject to the inter-cell interference, by assuming that the number of users in each cell scales in a given order with the signal-to-noise ratio (SNR). In particular, we investigate the achievable degrees of freedom (DoF) for the sum-rate in each cell when the SNR goes to infinity, and characterize the achievable DoF region for all the cells. Our results show that to achieve the DoF-optimal transmission in a multi-cell system with RBF, the numbers of transmit beams in all the cells need to be assigned in a collaborative manner based on the user densities.
Hieu Duy Nguyen, Rui Zhang 0006, Hon Tat Hui
ICASSP1