Minh Nguyen 0002

dblp:83/2833-2 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
6since 2021 · last 2024
0000-0003-4762-1798ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Representation and self-supervised learning · 31% Transfer learning and domain adaptation · 25% Speech recognition and synthesis · 14%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation › domain shift
correlation shift
0.812024
Adapting to Shifting Correlations with Unlabeled Data Calibration · ECCV (87) 2024
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.812024
Adapting to Shifting Correlations with Unlabeled Data Calibration · ECCV (87) 2024
Machine learning › Representation and self-supervised learning
causal representation learning
0.712023
Learning Invariant Representations with a Nonparametric Nadaraya-Watson Head · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation
domain generalization
0.712023
Learning Invariant Representations with a Nonparametric Nadaraya-Watson Head · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning
invariant representation learning
0.712023
Learning Invariant Representations with a Nonparametric Nadaraya-Watson Head · NeurIPS 2023
Natural language and speech › Language models and text generation › text correction › spelling correction
chinese spelling check
0.512021
Domain-Shift Conditioning Using Adaptable Filtering Via Hierarchical Embeddings for Robust Chinese Spell Check · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Machine learning › Representation and self-supervised learning › text embedding
character embedding
0.412020
Hierarchical Character Embeddings: Learning Phonological and Semantic Representations in Languages of Logographic Origin Using Recursive Neural Networks · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Natural language and speech › Machine translation
transliteration
0.412019
Phonology-Augmented Statistical Framework for Machine Transliteration Using Limited Linguistic Resources · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Natural language and speech › Speech recognition and synthesis › pronunciation modeling
grapheme-to-phoneme conversion
0.312018
Multimodal neural pronunciation modeling for spoken languages with logographic origin · EMNLP 2018
Natural language and speech › Speech recognition and synthesis
pronunciation modeling
0.312018
Multimodal neural pronunciation modeling for spoken languages with logographic origin · EMNLP 2018

Methods — techniques the papers use, named apart from their topics

nonparametric regression · 0.7nadaraya-watson head · 0.7hierarchical character embeddings · 0.5confusion set filtering · 0.5recursive neural network · 0.4Tree-LSTM · 0.4statistical machine translation · 0.4neural network · 0.3geometric representation · 0.3
YearPublicationVenuePosition
2024 Adapting to Shifting Correlations with Unlabeled Data Calibration
Minh Nguyen 0002, Alan Wang 0003, Heejong Kim, Mert R. Sabuncu
ECCV (87)1
2024 Robust Learning via Conditional Prevalence Adjustment
abstract
Healthcare data often come from multiple sites in which the correlations between confounding variables can vary widely. If deep learning models exploit these unstable correlations, they might fail catastrophically in unseen sites. Although many methods have been proposed to tackle unstable correlations, each has its limitations. For example, adversarial training forces models to completely ignore unstable correlations, but doing so may lead to poor predictive performance. Other methods (e.g. Invariant Risk Minimization) try to learn domain-invariant representations that rely only on stable associations by assuming a causal data-generating process (input X causes class label Y ). Thus, they may be ineffective for anti-causal tasks (Y causes X), which are common in computer vision. We propose a method called CoPA (Conditional Prevalence-Adjustment) for anti-causal tasks. CoPA assumes that (1) generation mechanism is stable, i.e. label Y and confounding variable(s) Z generate X, and (2) the unstable conditional prevalence in each site E fully accounts for the unstable correlations between X and Y. Our crucial observation is that confounding variables are routinely recorded in healthcare settings and the prevalence can be readily estimated, for example, from a set of (Y,Z) samples (no need for corresponding samples of X). CoPA can work even if there is a single training site, a scenario which is often overlooked by existing methods. Our experiments on synthetic and real data show CoPA beating competitive baselines.
Minh Nguyen 0002, Alan Wang 0003, Heejong Kim, Mert R. Sabuncu
WACV1
2023 Learning Invariant Representations with a Nonparametric Nadaraya-Watson Head
abstract
Machine learning models will often fail when deployed in an environment with a data distribution that is different than the training distribution. When multiple environments are available during training, many methods exist that learn representations which are invariant across the different distributions, with the hope that these representations will be transportable to unseen domains. In this work, we present a nonparametric strategy for learning invariant representations based on the recently-proposed Nadaraya-Watson (NW) head. The NW head makes a prediction by comparing the learned representations of the query to the elements of a support set that consists of labeled data. We demonstrate that by manipulating the support set, one can encode different causal assumptions. In particular, restricting the support set to a single environment encourages the model to learn invariant features that do not depend on the environment. We present a causally-motivated setup for our modeling and training strategy and validate on three challenging real-world domain generalization tasks in computer vision.
Alan Wang 0003, Minh Nguyen 0002, Mert R. Sabuncu
NeurIPS2
2022 A transformer-Based neural language model that synthesizes brain activation maps from free-form text queries
Hoang Gia Ngo, Minh Nguyen 0002, Nancy F. Chen, Mert R. Sabuncu
Medical Image Anal.2
2021 Text2Brain: Synthesis of Brain Activation Maps from Free-Form Text Query
Hoang Gia Ngo, Minh Nguyen 0002, Nancy F. Chen, Mert R. Sabuncu
MICCAI (7)2
2021 Domain-Shift Conditioning Using Adaptable Filtering Via Hierarchical Embeddings for Robust Chinese Spell Check
abstract
Spell check is a useful application which processes noisy human-generated text. Spell check for Chinese poses unresolved problems due to the large number of characters, the sparse distribution of errors, and the dearth of resources with sufficient coverage of heterogeneous and shifting error domains. For Chinese spell check, filtering using confusion sets narrows the search space and makes finding corrections easier. However, most, if not all, confusion sets used to date are fixed and thus do not include new, shifting error domains. We propose a scalable adaptable filter that exploits hierarchical character embeddings to (1) obviate the need to handcraft confusion sets, and (2) resolve sparsity problems related to infrequent errors. Our approach compares favorably with competitive baselines and obtains SOTA results on the 2014 and 2015 Chinese Spelling Check Bake-off datasets.
Minh Nguyen 0002, Hoang Gia Ngo, Nancy F. Chen
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Hierarchical Character Embeddings: Learning Phonological and Semantic Representations in Languages of Logographic Origin Using Recursive Neural Networks
abstract
Logographs (Chinese characters) have recursive structures (i.e. hierarchies of sub-units in logographs) that contain phonological and semantic information, as developmental psychology literature suggests that native speakers leverage on the structures to learn how to read. Exploiting these structures could potentially lead to better embeddings that can benefit many downstream tasks. We propose building hierarchical logograph (character) embeddings from logograph recursive structures using treeLSTM, a recursive neural network. Using recursive neural network imposes a prior on the mapping from logographs to embeddings since the network must read in the sub-units in logographs according to the order specified by the recursive structures. Based on human behavior in language learning and reading, we hypothesize that modeling logographs' structures using recursive neural network should be beneficial. To verify this claim, we consider two tasks (1) predicting logographs' Cantonese pronunciation from logographic structures and (2) language modeling. Empirical results show that the proposed hierarchical embeddings outperform baseline approaches. Diagnostic analysis suggests that hierarchical embeddings constructed using treeLSTM is less sensitive to distractors, thus is more robust, especially on complex logographs.
Minh Nguyen 0002, Hoang Gia Ngo, Nancy F. Chen
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Phonology-Augmented Statistical Framework for Machine Transliteration Using Limited Linguistic Resources
abstract
Transliteration converts words in a source language (e.g., English) into words in a target language (e.g., Vietnamese). This conversion considers the phonological structure of the target language, as the transliterated output needs to be pronounceable in the target language. For example, a word in Vietnamese that begins with a consonant cluster is phonologically invalid and thus would be an incorrect output of a transliteration system. Most statistical transliteration approaches, albeit being widely adopted, do not explicitly model the target language's phonology, which often results in invalid outputs. The problem is compounded by the limited linguistic resources available when converting foreign words to transliterated words in the target language. In this paper, we present a phonology-augmented statistical framework suitable for transliteration, especially when only limited linguistic resources are available. We propose the concept of pseudo-syllables as structures representing how segments of a foreign word are organized according to the syllables of the target language's phonology. We performed transliteration experiments on Vietnamese and Cantonese. We show that the proposed framework outperforms the statistical baseline by up to 44.68% relative, when there are limited training examples (587 entries).
Hoang Gia Ngo, Minh Nguyen 0002, Nancy F. Chen
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Multimodal neural pronunciation modeling for spoken languages with logographic origin
abstract
Graphemes of most languages encode pronunciation, though some are more explicit than others.Languages like Spanish have a straightforward mapping between its graphemes and phonemes, while this mapping is more convoluted for languages like English.Spoken languages such as Cantonese present even more challenges in pronunciation modeling: (1) they do not have a standard written form, (2) the closest graphemic origins are logographic Han characters, of which only a subset of these logographic characters implicitly encodes pronunciation.In this work, we propose a multimodal approach to predict the pronunciation of Cantonese logographic characters, using neural networks with a geometric representation of logographs and pronunciation of cognates in historically related languages.The proposed framework improves performance by 18.1% and 25.0% respective to unimodal and multimodal baselines.
Minh Nguyen 0002, Hoang Gia Ngo, Nancy F. Chen
EMNLP1