Anh H. T. Nguyen

dblp:201/7474 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
7since 2021 · last 2023
0000-0003-2725-3975ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Improving Performance of Real-Time Full-Band Blind Packet-Loss Concealment with Predictive Network
abstract
Packet loss concealment (PLC) is a tool for enhancing speech degradation caused by poor network conditions or underflow/overflow in audio processing pipelines. We propose a real-time recurrent method that leverages previous outputs to mitigate artefact of lost packets without the prior knowledge of loss mask. The proposed full-band recurrent network (FRN) model operates at 48 kHz, which is suitable for high-quality telecommunication applications. Experiment results highlight the superiority of FRN over an offline non-causal baseline and a top performer in a recent PLC challenge.
Anh H. T. Nguyen, Andy W. H. Khong
ICASSP2
2023 Gridless DOA Estimation Using Complex-Valued Convolutional Neural Network With Phasor Normalization
abstract
We propose a complex LeDIM-net (C-LeDIM-net) convolutional neural network (CNN) that employs a newly-formulated complex phasor normalization for gridless direction-of-arrival (DOA) estimation. Unlike existing deep learning (DL) approaches, C-LeDIM-net extracts explicit phase information in its intermediate complex-valued feature maps to estimate unknown source DOAs. Given its explicit phase representation, the proposed complex phasor normalization leverages the phase-to-sensor relationship of the feature maps which, as a consequence, improves the robustness of C-LeDIM-net to array imperfections when operating with limited number of snapshots. Simulation results show that the proposed method outperforms the existing methods, including the subspace-based and DL-based methods.
Zhi-Wei Tan, Yuan Liu 0007, Andy W. H. Khong, Anh H. T. Nguyen
IEEE Signal Process. Lett.4
2022 Tunet: A Block-Online Bandwidth Extension Model Based On Transformers And Self-Supervised Pretraining
abstract
We introduce a block-online variant of the temporal feature-wise linear modulation (TFiLM) model to achieve bandwidth extension. The proposed architecture simplifies the UNet backbone of the TFiLM to reduce inference time and employs an efficient transformer at the bottleneck to alleviate performance degradation. We also utilize self-supervised pretraining and data augmentation to enhance the quality of bandwidth extended signals and reduce the sensitivity with respect to downsampling methods. Experiment results on the VCTK dataset show that the proposed method outperforms several recent baselines in both intrusive and non-intrusive metrics. Pretraining and filter augmentation also help stabilize and enhance the overall performance.
Anh H. T. Nguyen, Andy W. H. Khong
ICASSP2
2022 Multichannel Noise Reduction Using Dilated Multichannel U-Net and Pre-Trained Single-Channel Network
abstract
Pre-trained single-channel neural networks have become more prevalent for noise reduction in recent years. However, unlike their multichannel counterparts, these monoaural approaches do not exploit spatial information during the optimization process. Furthermore, while multichannel neural networks exploit spatial information, they are optimized for a specific microphone array configuration; extensive data collection and training are required if a new array configuration is deployed. We propose a transfer learning approach that leverages existing pre-trained single-channel neural networks for the optimization of multichannel neural networks. Simulation results on the CHiME-3 dataset show that the proposed method outperforms the state-of-the-art multichannel neural network and neural beamformer.
Zhi-Wei Tan, Anh H. T. Nguyen, Yuan Liu 0007, Andy W. H. Khong
ICASSP2
2022 Iterative Implementation Method for Robust Target Localization in a Mixed Interference Environment
abstract
For the problem of target localization under the multipath propagation environment, the existing methods are mainly restricted to the limited prior information of complex reflections, especially when the target is embedded in a mixed interference environment. They may suffer from performance degradation due to the shortage of target classification ability. To address this problem, we propose a target localization method based on iterative implementation with semiunitary constraint and eigen-decomposition technique, where a practical propagation scenario based on the spherical Earth model is considered. Compared to the previous works, the proposed method can automatically distinguish a real target from the mixed interference environment with improved localization accuracy. Neither additional decorrelation preprocessing nor prior information of the dynamic scenario is required. Both simulations and real data experiments validate the effectiveness and robustness of the proposed method.
Yuan Liu 0007, Xiang-Gen Xia 0001, Hongwei Liu 0001, Anh H. T. Nguyen, Andy W. H. Khong
IEEE Trans. Geosci. Remote. Sens.4
2021 An Adaptive Non-Linear Process for Under-Determined Virtual Microphone Beamforming
abstract
Virtual microphone beamforming techniques are attractive for devices limited by space constraints. These techniques synthesize virtual microphone signals via interpolation algorithms. We propose to extend existing virtual microphone signal interpolation by employing an adaptive non-linear (ANL) process for acoustic beamforming. The proposed ANL based interpolation utilizes a target-presence probability criteria to determine the degree of non-linearity. The beamformer output is then derived using a combination between interpolations during target inactive zones and target active zones. Such combination offers a trade-off between reducing interference and target signal distortion. We apply the proposed ANL-based interpolator to the maximum signal-to-noise ratio (MSNR) beamformer and compare its performance against conventional beamforming and virtual microphone based beamforming methods in under-determined situations.
Mehdi Bekrani, Anh H. T. Nguyen, Andy W. H. Khong
ICASSP2
2021 Directional Sparse Filtering Using Weighted Lehmer Mean for Blind Separation of Unbalanced Speech Mixtures
abstract
In blind source separation of speech signals, the inherent imbalance in the source spectrum poses a challenge for methods that rely on single-source dominance for the estimation of the mixing matrix. We propose an algorithm based on the directional sparse filtering (DSF) framework that utilizes the Lehmer mean with learnable weights to adaptively account for source imbalance. Performance evaluation in multiple real acoustic environments show improvements in source separation compared to the baseline methods.
Karn Watcharasupat, Anh H. T. Nguyen, Ching-Hui Ooi, Andy W. H. Khong
ICASSP2
2019 A Method Based on L-bfgs to Solve Constrained Complex-valued Ica
abstract
Complex-valued independent component analysis (ICA) is a celebrated method in blind separation of complex-valued signals. In this paper, we propose to transform the constrained optimization problems of complex-valued ICA into unconstrained optimization problems which can be solved by limited-memory Broyden-Fletcher-Goldfarb-Shanno update (L-BFGS). As opposed to previous approaches, the proposed method does not apply any restriction on the Hessian matrix of ICA cost function. It can separate mixed sub-Gaussian, super-Gaussian, circular, and non-circular sources. Simulations show promising results.
Anh H. T. Nguyen, V. G. Reju, Andy W. H. Khong
ICASSP1
2017 Learning complex-valued latent filters with absolute cosine similarity
abstract
We propose a new sparse coding technique based on the power mean of phase-invariant cosine distances. Our approach is a generalization of sparse filtering and K-hyperlines clustering. It offers a better sparsity enforcer than the L1/L2norm ratio that is typically used in sparse filtering. At the same time, the proposed approach scales better than the clustering counterparts for high-dimensional input. Our algorithm fully exploits the prior information obtained by preprocessing the observed data with whitening via an efficient row-wise decoupling scheme. In our simulating experiments, the algorithm produces better estimates than previous approaches do. It yields better separation of live recorded speech mixtures as well.
Anh H. T. Nguyen, V. G. Reju, Andy W. H. Khong, Ing Yann Soon
ICASSP1