Michael Saxon

dblp:222/6656 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0001-7306-5030ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Do You Know About My Nation? Investigating Multilingual Language Models' Cultural Literacy Through Factual Knowledge
abstract
Most multilingual question-answering benchmarks, while covering a diverse pool of languages, do not factor in regional diversity in the information they capture and tend to be Western-centric.This introduces a significant gap in fairly evaluating multilingual models' comprehension of factual information from diverse geographical locations.To address this, we introduce XNationQA for investigating the cultural literacy of multilingual LLMs.XNationQA encompasses a total of 49, 280 questions on the geography, culture, and history of nine countries, presented in seven languages.We benchmark eight standard multilingual LLMs on XNationQA and evaluate them using two novel transference metrics.Our analyses uncover a considerable discrepancy in the models' accessibility to culturally specific facts across languages.Notably, we often find that a model demonstrates greater knowledge of cultural information in English than in the dominant language of the respective culture.The models exhibit better performance in Western languages, although this does not necessarily translate to being more literate for Western countries, which is counterintuitive.Furthermore, we observe that models have a very limited ability to transfer knowledge across languages, particularly evident in open-source models 1 .
Eshaan Tanwar, Anwoy Chatterjee, Michael Saxon, Alon Albalak, William Yang Wang, Tanmoy Chakraborty 0002
EMNLP3
2025 VSP: Diagnosing the Dual Challenges of Perception and Reasoning in Spatial Planning Tasks for MLLMS
Qiucheng Wu, Handong Zhao, Michael Saxon, Trung Bui, William Yang Wang, Yang Zhang 0001, Shiyu Chang
ICCV3
2024 Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)
abstract
With advances in the quality of text-to-image (T2I) models has come interest in benchmarking their prompt faithfulness---the semantic coherence of generated images to the prompts they were conditioned on. A variety of T2I faithfulness metrics have been proposed, leveraging advances in cross-modal embeddings and vision-language models (VLMs). However, these metrics are not rigorously compared and benchmarked, instead presented with correlation to human Likert scores over a set of easy-to-discriminate images against seemingly weak baselines. We introduce T2IScoreScore, a curated set of semantic error graphs containing a prompt and a set of increasingly erroneous images. These allow us to rigorously judge whether a given prompt faithfulness metric can correctly order images with respect to their objective error count and significantly discriminate between different error nodes, using meta-metric scores derived from established statistical tests. Surprisingly, we find that the state-of-the-art VLM-based metrics (e.g., TIFA, DSG, LLMScore, VIEScore) we tested fail to significantly outperform simple (and supposedly worse) feature-based metrics like CLIPScore, particularly on a hard subset of naturally-occurring T2I model errors. TS2 will enable the development of better T2I prompt faithfulness metrics through more rigorous comparison of their conformity to expected orderings and separations under objective criteria.
Michael Saxon, Fatima Jahara, Mahsa Khoshnoodi, William Yang Wang
NeurIPS1
2024 Automatically Correcting Large Language Models: Surveying the Landscape of Diverse Automated Correction Strategies
abstract
Abstract While large language models (LLMs) have shown remarkable effectiveness in various NLP tasks, they are still prone to issues such as hallucination, unfaithful reasoning, and toxicity. A promising approach to rectify these flaws is correcting LLMs with feedback, where the LLM itself is prompted or guided with feedback to fix problems in its own output. Techniques leveraging automated feedback—either produced by the LLM itself (self-correction) or some external system—are of particular interest as they make LLM-based solutions more practical and deployable with minimal human intervention. This paper provides an exhaustive review of the recent advances in correcting LLMs with automated feedback, categorizing them into training-time, generation-time, and post-hoc approaches. We also identify potential challenges and future directions in this emerging field.
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang 0003, William Yang Wang
Trans. Assoc. Comput. Linguistics2
2023 Multilingual Conceptual Coverage in Text-to-Image Models
abstract
DallE mega1 Figure 1: A selection of images generated by DALLE-mega, Stable Diffusion 2, DALLE-2, and AltDiffusion, illustrating their conceptual coverage of "dog," "airplane," and "face" across English, Spanish, German, Chinese (simplified), Japanese, Hebrew, and Indonesian.Coverage of the concepts varies considerably across model and language, and can be observed in the consistency and correctness of images generated under simple prompts.
Michael Saxon, William Yang Wang
ACL (1)1
2023 PECO: Examining Single Sentence Label Leakage in Natural Language Inference Datasets through Progressive Evaluation of Cluster Outliers
abstract
Building natural language inference (NLI) benchmarks that are both challenging for modern techniques, and free from shortcut biases is difficult.Chief among these biases is single sentence label leakage, where annotatorintroduced spurious correlations yield datasets where the logical relation between (premise, hypothesis) pairs can be accurately predicted from only a single sentence, something that should in principle be impossible.We demonstrate that despite efforts to reduce this leakage, it persists in modern datasets that have been introduced since its 2018 discovery.To enable future amelioration efforts, introduce a novel model-driven technique, the progressive evaluation of cluster outliers (PECO) which enables both the objective measurement of leakage, and the automated detection of subpopulations in the data which maximally exhibit it.
Michael Saxon, Xinyi Wang 0003, Wenda Xu, William Yang Wang
EACL1
2023 Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought
abstract
Vaishnavi Himakunthala, Andy Ouyang, Daniel Rose, Ryan He, Alex Mei, Yujie Lu, Chinmay Sonar, Michael Saxon, William Wang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Vaishnavi Himakunthala, Andy Ouyang, Daniel Rose, Ryan He, Alex Mei, Chinmay Sonar, Michael Saxon, William Yang Wang
EMNLP8
2023 WikiWhy: Answering and Explaining Cause-and-Effect Questions
Matthew Ho, Michael Saxon, Sharon Levy, William Yang Wang
ICLR4
2023 Causal Balancing for Domain Generalization
Xinyi Wang 0003, Michael Saxon, Hongyang Zhang 0001, Kun Zhang 0001, William Yang Wang
ICLR2
2023 Data Augmentation for Diverse Voice Conversion in Noisy Environments
Avani Tanna, Michael Saxon, Amr El Abbadi, William Yang Wang
INTERSPEECH2
2023 Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning
abstract
In recent years, pre-trained large language models (LLMs) have demonstrated remarkable efficiency in achieving an inference-time few-shot learning capability known as in-context learning. However, existing literature has highlighted the sensitivity of this capability to the selection of few-shot demonstrations. Current understandings of the underlying mechanisms by which this capability arises from regular language model pretraining objectives remain disconnected from the real-world LLMs. This study aims to examine the in-context learning phenomenon through a Bayesian lens, viewing real-world LLMs as latent variable models. On this premise, we propose an algorithm to select optimal demonstrations from a set of annotated data with a small LM, and then directly generalize the selected demonstrations to larger LMs. We demonstrate significant improvement over baselines, averaged over eight GPT models on eight real-world text classification datasets. We also demonstrate the real-world usefulness of our algorithm on GSM8K, a math word problem dataset. Our empirical findings support our hypothesis that LLMs implicitly infer a latent variable containing task information.
Xinyi Wang 0003, Wanrong Zhu, Michael Saxon, Mark Steyvers, William Yang Wang
NeurIPS3
2022 Self-Supervised Knowledge Assimilation for Expert-Layman Text Style Transfer
abstract
Expert-layman text style transfer technologies have the potential to improve communication between members of scientific communities and the general public. High-quality information produced by experts is often filled with difficult jargon laypeople struggle to understand. This is a particularly notable issue in the medical domain, where layman are often confused by medical text online. At present, two bottlenecks interfere with the goal of building high-quality medical expert-layman style transfer systems: a dearth of pretrained medical-domain language models spanning both expert and layman terminologies and a lack of parallel corpora for training the transfer task itself. To mitigate the first issue, we propose a novel language model (LM) pretraining task, Knowledge Base Assimilation, to synthesize pretraining data from the edges of a graph of expert- and layman-style medical terminology terms into an LM during self-supervised learning. To mitigate the second issue, we build a large-scale parallel corpus in the medical expert-layman domain using a margin-based criterion. Our experiments show that transformer-based models pretrained on knowledge base assimilation and other well-established pretraining tasks fine-tuning on our new parallel corpus leads to considerable improvement against expert-layman transfer benchmarks, gaining an average relative improvement of our human evaluation, the Overall Success Rate (OSR), by 106%.
Wenda Xu, Michael Saxon, Misha Sra, William Yang Wang
AAAI2
2021 Modeling Disclosive Transparency in NLP Application Descriptions
abstract
Broader disclosive transparency-truth and clarity in communication regarding the function of AI systems-is widely considered desirable.Unfortunately, it is a nebulous concept, difficult to both define and quantify.This is problematic, as previous work has demonstrated possible trade-offs and negative consequences to disclosive transparency, such as a confusion effect, where "too much information" clouds a reader's understanding of what a system description means.Disclosive transparency's subjective nature has rendered deep study into these problems and their remedies difficult.To improve this state of affairs, We introduce neural language model-based probabilistic metrics to directly model disclosive transparency, and demonstrate that they correlate with user and expert opinions of system transparency, making them a valid objective proxy.Finally, we demonstrate the use of these metrics in a pilot study quantifying the relationships between transparency, confusion, and user perceptions in a corpus of real NLP system descriptions.
Michael Saxon, Sharon Levy, Xinyi Wang 0003, Alon Albalak, William Yang Wang
EMNLP (1)1
2021 End-to-End Spoken Language Understanding for Generalized Voice Assistants
abstract
End-to-end (E2E) spoken language understanding (SLU) systems predict utterance semantics directly from speech using a single model. Previous work in this area has focused on targeted tasks in fixed domains, where the output semantic structure is assumed a priori and the input speech is of limited complexity. In this work we present our approach to developing an E2E model for generalized SLU in commercial voice assistants (VAs). We propose a fully differentiable, transformer-based, hierarchical system that can be pretrained at both the ASR and NLU levels. This is then fine-tuned on both transcription and semantic classification losses to handle a diverse set of intent and argument combinations. This leads to an SLU system that achieves significant improvements over baselines on a complex internal generalized VA dataset with a 43% improvement in accuracy, while still meeting the 99% accuracy benchmark on the popular Fluent Speech Commands dataset. We further evaluate our model on a hard test set, exclusively containing slot arguments unseen in training, and demonstrate a nearly 20% improvement, showing the efficacy of our approach in truly demanding VA scenarios.
Michael Saxon, Samridhi Choudhary, Joseph P. McKenna, Athanasios Mouchtaris
Interspeech1
2021 Counterfactual Maximum Likelihood Estimation for Training Deep Networks
abstract
Although deep learning models have driven state-of-the-art performance on a wide array of tasks, they are prone to spurious correlations that should not be learned as predictive clues. To mitigate this problem, we propose a causality-based training framework to reduce the spurious correlations caused by observed confounders. We give theoretical analysis on the underlying general Structural Causal Model (SCM) and propose to perform Maximum Likelihood Estimation (MLE) on the interventional distribution instead of the observational distribution, namely Counterfactual Maximum Likelihood Estimation (CMLE). As the interventional distribution, in general, is hidden from the observational data, we then derive two different upper bounds of the expected negative log-likelihood and propose two general algorithms, Implicit CMLE and Explicit CMLE, for causal predictions of deep learning models using observational data. We conduct experiments on both simulated data and two real-world tasks: Natural Language Inference (NLI) and Image Captioning. The results show that CMLE methods outperform the regular MLE method in terms of out-of-domain generalization performance and reducing spurious correlations, while maintaining comparable performance on the regular evaluations.
Xinyi Wang 0003, Wenhu Chen, Michael Saxon, William Yang Wang
NeurIPS3
2020 Semantic Complexity in End-to-End Spoken Language Understanding
abstract
End-to-end spoken language understanding (SLU) models are a class of model architectures that predict semantics directly from speech. Because of their input and output types, we refer to them as speech-to-interpretation (STI) models. Previous works have successfully applied STI models to targeted use cases, such as recognizing home automation commands, however no study has yet addressed how these models generalize to broader use cases. In this work, we analyze the relationship between the performance of STI models and the difficulty of the use case to which they are applied. We introduce empirical measures of dataset semantic complexity to quantify the difficulty of the SLU tasks. We show that near-perfect performance metrics for STI models reported in the literature were obtained with datasets that have low semantic complexity values. We perform experiments where we vary the semantic complexity of a large, proprietary dataset and show that STI model performance correlates with our semantic complexity measures, such that performance increases as complexity values decrease. Our results show that it is important to contextualize an STI model's performance with the complexity values of its training dataset to reveal the scope of its applicability.
Joseph P. McKenna, Samridhi Choudhary, Michael Saxon, Grant P. Strimel, Athanasios Mouchtaris
INTERSPEECH3
2020 UncommonVoice: A Crowdsourced Dataset of Dysphonic Speech
Meredith Moore 0001, Piyush Papreja, Michael Saxon, Visar Berisha, Sethuraman Panchanathan
INTERSPEECH3
2020 Robust Estimation of Hypernasality in Dysarthria With Acoustic Model Likelihood Features
abstract
Hypernasality is a common characteristic symptom across many motor-speech disorders. For voiced sounds, hypernasality introduces an additional resonance in the lower frequencies and, for unvoiced sounds, there is reduced articulatory precision due to air escaping through the nasal cavity. However, the acoustic manifestation of these symptoms is highly variable, making hypernasality estimation very challenging, both for human specialists and automated systems. Previous work in this area relies on either engineered features based on statistical signal processing or machine learning models trained on clinical ratings. Engineered features often fail to capture the complex acoustic patterns associated with hypernasality, whereas metrics based on machine learning are prone to overfitting to the small disease-specific speech datasets on which they are trained. Here we propose a new set of acoustic features that capture these complementary dimensions. The features are based on two acoustic models trained on a large corpus of healthy speech. The first acoustic model aims to measure nasal resonance from voiced sounds, whereas the second acoustic model aims to measure articulatory imprecision from unvoiced sounds. To demonstrate that the features derived from these acoustic models are specific to hypernasal speech, we evaluate them across different dysarthria corpora. Our results show that the features generalize even when training on hypernasal speech from one disease and evaluating on hypernasal speech from another disease (e.g., training on Parkinson's disease, evaluation on Huntington's disease), and when training on neurologically disordered speech but evaluating on cleft palate speech.
Michael Saxon, Ayush Tripathi, Yishan Jiao, Julie M. Liss, Visar Berisha
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Objective Measures of Plosive Nasalization in Hypernasal Speech
abstract
Hypernasal speech is a common symptom across several neurological disorders; however it has a variable acoustic signature, making it difficult to quantify acoustically or perceptually. In this paper, we propose the nasal cognate distinctiveness features as an objective proxy for hypernasal speech. Our method is motivated by the observation that incomplete velopharyngeal closure changes the acoustics of the resultant speech such that alveolar stops /t/ and /d/ map to the alveolar nasal /n/ and bilabial stops /b/ and /p/ map to bilabial nasal /m/. We propose a new family of features based on likelihood ratios between the plosives and their respective nasal cognates. These features are based on an acoustic model that is trained only on healthy speech, and evaluated on a set of 75 speakers diagnosed with different dysarthria subtypes and exhibiting varying levels of hypernasality. Our results show that the family of features compares favorably with the clinical perception of speech-language pathologists subjectively evaluating hypernasality.
Michael Saxon, Julie M. Liss, Visar Berisha
ICASSP1
2019 Say What? A Dataset for Exploring the Error Patterns That Two ASR Engines Make
Meredith Moore 0001, Michael Saxon, Hemanth Venkateswara, Visar Berisha, Sethuraman Panchanathan
INTERSPEECH2