Yang Xu 0023

dblp:61/3906-23 · DBLP profile ↗
← Back
61ranked-venue papers
7as first author
28since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 7 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 45 · 6 first-author · 17 since 2021
YearPublicationVenuePosition
2025 Modeling word overextension in a Grey Parrot
Michal Fishkin, Shereen Chang, Yang Xu 0023
CogSci3
2025 Modelling compounding across languages with analogy and composition
Aotao Xu, Charles Kemp, Lea Frermann, Yang Xu 0023
CogSci4
2025 Visual moral inference and communication
Warren Zhu, Aida Ramezani, Yang Xu 0023
CogSci3
2025 The discordance between embedded ethics and cultural inference in large language models
abstract
Effective interactions between artificial intelligence (AI) and humans require an equitable and accurate representation of diverse cultures.It is known that current AI, particularly large language models (LLMs), possess some degrees of cultural knowledge but not without limitations.We present a framework aimed at understanding the origin of these limitations.We hypothesize that there is a fundamental discordance between embedded ethics-how LLMs represent right versus wrong, and cultural inference-how LLMs infer cultural knowledge, specifically cultural norms.We demonstrate this by extracting low-dimensional subspaces that embed ethical principles of LLMs based on established benchmarks.We then show that how LLMs make errors in culturally distinctive scenarios significantly correlates with how they represent cultural norms with respect to these embedded ethics subspaces.Furthermore, we show that coercing cultural norms to be more aligned with the embedded ethics increases LLM performance in cultural inference.Our analyses of 12 language models, two large-scale cultural benchmarks spanning 75 countries and two ethical datasets indicate that 1) the ethics-culture discordance tends to be exacerbated in instruct-tuned models, and 2) how current LLMs represent ethics can impose limitations on their adaptation to diverse cultures particularly pertaining to non-Western and low-income regions. 1
Aida Ramezani, Yang Xu 0023
EMNLP2
2024 Cognitive Factors in Word Sense Decline
Aniket Kali, Yang Xu 0023, Suzanne Stevenson
CogSci2
2024 Moral association graph: A cognitive model for moral inference
Aida Ramezani, Yang Xu 0023
CogSci2
2024 Toward Informal Language Processing: Knowledge of Slang in Large Language Models
abstract
Zhewei Sun, Qian Hu, Rahul Gupta, Richard Zemel, Yang Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhewei Sun, Rahul Gupta 0001, Richard S. Zemel, Yang Xu 0023
NAACL-HLT5
2023 Knowledge of cultural moral norms in large language models
abstract
Moral norms vary across cultures.A recent line of work suggests that English large language models contain human-like moral biases, but these studies typically do not examine moral variation in a diverse cultural setting.We investigate the extent to which monolingual English language models contain knowledge about moral norms in different countries.We consider two levels of analysis: 1) whether language models capture fine-grained moral variation across countries over a variety of topics such as "homosexuality" and "divorce"; 2) whether language models capture cultural diversity and shared tendencies in which topics people around the globe tend to diverge or agree on in their moral judgment.We perform our analyses with two public datasets from the World Values Survey (across 55 countries) and PEW global surveys (across 40 countries) on morality.We find that pre-trained English language models predict empirical moral norms across countries worse than the English moral norms reported previously.However, fine-tuning language models on the survey data improves inference across countries at the expense of a less accurate estimate of the English moral norms.We discuss the relevance and challenges of incorporating cultural knowledge into the automated inference of moral norms.
Aida Ramezani, Yang Xu 0023
ACL (1)2
2023 Word sense extension
abstract
Humans often make creative use of words to express novel senses.A long-standing effort in natural language processing has been focusing on word sense disambiguation (WSD), but little has been explored about how the sense inventory of a word may be extended toward novel meanings.We present a paradigm of word sense extension (WSE) that enables words to spawn new senses toward novel context.We develop a framework that simulates novel word sense extension by first partitioning a polysemous word type into two pseudo-tokens that mark its different senses, and then inferring whether the meaning of a pseudo-token can be extended to convey the sense denoted by the token partitioned from the same word type.Our framework combines cognitive models of chaining with a learning scheme that transforms a language model embedding space to support various types of word sense extension.We evaluate our framework against several competitive baselines and show that it is superior in predicting plausible novel senses for over 7,500 English words.Furthermore, we show that our WSE framework improves performance over a range of transformer-based WSD models in predicting rare word senses with few or zero mentions in the training data.
Yang Xu 0023
ACL (1)2
2023 Quantifying informativeness of names in visual space
Eleonora Gualdoni, Charles Kemp, Yang Xu 0023, Gemma Boleda
CogSci3
2023 Quantifying Bias in Library Classification Systems
Katie Warburton, Charles Kemp, Yang Xu 0023, Lea Frermann
CogSci3
2023 Predicting strategy choice in word formation: A case study of reuse and compounding
Aotao Xu, Charles Kemp, Lea Frermann, Yang Xu 0023
CogSci4
2022 Neural reality of argument structure constructions
abstract
In lexicalist linguistic theories, argument structure is assumed to be predictable from the meaning of verbs.As a result, the verb is the primary determinant of the meaning of a clause.In contrast, construction grammarians propose that argument structure is encoded in constructions (or form-meaning pairs) that are distinct from verbs.Decades of psycholinguistic research have produced substantial empirical evidence in favor of the construction view.Here we adapt several psycholinguistic studies to probe for the existence of argument structure constructions (ASCs) in Transformerbased language models (LMs).First, using a sentence sorting experiment, we find that sentences sharing the same construction are closer in embedding space than sentences sharing the same verb.Furthermore, LMs increasingly prefer grouping by construction with more input data, mirroring the behaviour of non-native language learners.Second, in a "Jabberwocky" priming-based experiment, we find that LMs associate ASCs with meaning, even in semantically nonsensical sentences.Our work offers the first evidence for ASCs in LMs and highlights the potential to devise novel probing methods grounded in psycholinguistic research. Transitive DitransitiveCaused-motion Resultative Throw Anita threw the hammer.Chris threw Linda the pencil.Pat threw the keys onto the roof.Lyn threw the box apart.
Zining Zhu 0001, Guillaume Thomas, Frank Rudzicz, Yang Xu 0023
ACL (1)5
2022 Communicative need modulates lexical precision across semantic domains: A domain-level account of efficient communication
Laurestine Bradford, Guillaume Thomas, Yang Xu 0023
CogSci3
2022 Gender bias in grammatical gender systems across languages
Sonali Dey, Yang Xu 0023
CogSci2
2022 The emergence of moral foundations in child language development
Aida Ramezani, Emmy Liu, Renato Ferreira Pinto Junior, Spike W. S. Lee, Yang Xu 0023
CogSci5
2022 Evolution of moral semantics through metaphorization
Aida Ramezani, Jennifer Stellar, Matthew Feinberg, Yang Xu 0023
CogSci4
2022 Word formation supports efficient communication: The case of compounds
Aotao Xu, Charles Kemp, Lea Frermann, Yang Xu 0023
CogSci4
2022 Infinite mixture chaining: Efficient temporal construction of word meaning
Yang Xu 0023
CogSci2
2022 Tracing Semantic Variation in Slang
abstract
The meaning of a slang term can vary in different communities.However, slang semantic variation is not well understood and underexplored in the natural language processing of slang.One existing view argues that slang semantic variation is driven by culture-dependent communicative needs.An alternative view focuses on slang's social functions suggesting that the desire to foster semantic distinction may have led to the historical emergence of community-specific slang senses.We explore these theories using computational models and test them against historical slang dictionary entries, with a focus on characterizing regularity in the geographical variation of slang usages attested in the US and the UK over the past two centuries.We show that our models are able to predict the regional identity of emerging slang word meanings from historical slang records.We offer empirical evidence that both communicative need and semantic distinction play a role in the variation of slang meaning yet their relative importance fluctuates over the course of history.Our work offers an opportunity for incorporating historical cultural elements into the natural language processing of slang.
Zhewei Sun, Yang Xu 0023
EMNLP2
2022 Semantically Informed Slang Interpretation
abstract
Slang is a predominant form of informal language making flexible and extended use of words that is notoriously hard for natural language processing systems to interpret.Existing approaches to slang interpretation tend to rely on context but ignore semantic extensions common in slang word usage.We propose a semantically informed slang interpretation (SSI) framework that considers jointly the contextual and semantic appropriateness of a candidate interpretation for a query slang.We perform rigorous evaluation on two large-scale online slang dictionaries and show that our approach not only achieves state-of-the-art accuracy for slang interpretation in English, but also does so in zero-shot and few-shot scenarios where training data is sparse.Furthermore, we show how the same framework can be applied to enhancing machine translation of slang from English to other languages.Our work creates opportunities for the automated interpretation and translation of informal language.
Zhewei Sun, Richard S. Zemel, Yang Xu 0023
NAACL-HLT3
2022 Noun2Verb: Probabilistic Frame Semantics for Word Class Conversion
abstract
Abstract Humans can flexibly extend word usages across different grammatical classes, a phenomenon known as word class conversion. Noun-to-verb conversion, or denominal verb (e.g., to Google a cheap flight), is one of the most prevalent forms of word class conversion. However, existing natural language processing systems are impoverished in interpreting and generating novel denominal verb usages. Previous work has suggested that novel denominal verb usages are comprehensible if the listener can compute the intended meaning based on shared knowledge with the speaker. Here we explore a computational formalism for this proposal couched in frame semantics. We present a formal framework, Noun2Verb, that simulates the production and comprehension of novel denominal verb usages by modeling shared knowledge of speaker and listener in semantic frames. We evaluate an incremental set of probabilistic models that learn to interpret and generate novel denominal verb usages via paraphrasing. We show that a model where the speaker and listener cooperatively learn the joint distribution over semantic frame elements better explains the empirical denominal verb usages than state-of-the-art language models, evaluated against data from (1) contemporary English in both adult and child speech, (2) contemporary Mandarin Chinese, and (3) the historical development of English. Our work grounds word class conversion in probabilistic frame semantics and bridges the gap between natural language processing systems and humans in lexical creativity.
Yang Xu 0023
Comput. Linguistics2
2021 How is BERT surprised? Layerwise detection of linguistic anomalies
abstract
Bai Li, Zining Zhu, Guillaume Thomas, Yang Xu, Frank Rudzicz. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zining Zhu 0001, Guillaume Thomas, Yang Xu 0023, Frank Rudzicz
ACL/IJCNLP (1)4
2021 Chaining and the formation of spatial semantic categories in childhood
Aotao Xu, Yang Xu 0023
CogSci2
2021 Predicting emergent linguistic compositions through time: Syntactic frame extension via multimodal chaining
abstract
Natural language relies on a finite lexicon to express an unbounded set of emerging ideas.One result of this tension is the formation of new compositions, such that existing linguistic units can be combined with emerging items into novel expressions.We develop a framework that exploits the cognitive mechanisms of chaining and multimodal knowledge to predict emergent compositional expressions through time.We present the syntactic frame extension model (SFEM) that draws on the theory of chaining and knowledge from "percept", "concept", and "language" to infer how verbs extend their frames to form new compositions with existing and novel nouns.We evaluate SFEM rigorously on the 1) modalities of knowledge and 2) categorization models of chaining, in a syntactically parsed English corpus over the past 150 years.We show that multimodal SFEM predicts newly emerged verb syntax and arguments substantially better than competing models using purely linguistic or unimodal knowledge.We find support for an exemplar view of chaining as opposed to a prototype view and reveal how the joint approach of multimodal chaining may be fundamental to the creation of literal and figurative language uses including metaphor and metonymy.
Yang Xu 0023
EMNLP (1)2
2021 Extended application of genomic selection to screen multiomics data for prognostic signatures of prostate cancer
abstract
Prognostic tests using expression profiles of several dozen genes help provide treatment choices for prostate cancer (PCa). However, these tests require improvement to meet the clinical need for resolving overtreatment, which continues to be a pervasive problem in PCa management. Genomic selection (GS) methodology, which utilizes whole-genome markers to predict agronomic traits, was adopted in this study for PCa prognosis. We leveraged The Cancer Genome Atlas (TCGA) database to evaluate the prediction performance of six GS methods and seven omics data combinations, which showed that the Best Linear Unbiased Prediction (BLUP) model outperformed the other methods regarding predictability and computational efficiency. Leveraging the BLUP-HAT method, an accelerated version of BLUP, we demonstrated that using expression data of a large number of disease-relevant genes and with an integration of other omics data (i.e. miRNAs) significantly increased outcome predictability when compared with panels consisting of a small number of genes. Finally, we developed a novel stepwise forward selection BLUP-HAT method to facilitate searching multiomics data for predictor variables with prognostic potential. The new method was applied to the TCGA data to derive mRNA and miRNA expression signatures for predicting relapse-free survival of PCa, which were validated in six independent cohorts. This is a transdisciplinary adoption of the highly efficient BLUP-HAT method and its derived algorithms to analyze multiomics data for PCa prognosis. The results demonstrated the efficacy and robustness of the new methodology in developing prognostic models in PCa, suggesting a potential utility in managing other types of cancer.
Ruidong Li 0002, Shibo Wang 0001, Yanru Cui, Han Qu, John M. Chater, Le Zhang 0022, Julong Wei, Meiyue Wang, Yang Xu 0023, Jianming Lu, Yuanfa Feng, Renyuan Ma, Jianguo Zhu 0003, Wei-De Zhong, Zhenyu Jia
Briefings Bioinform.9
2021 Boosting predictabilities of agronomic traits in rice using bivariate genomic selection
abstract
The multivariate genomic selection (GS) models have not been adequately studied and their potential remains unclear. In this study, we developed a highly efficient bivariate (2D) GS method and demonstrated its significant advantages over the univariate (1D) rival methods using a rice dataset, where four traditional traits (i.e. yield, 1000-grain weight, grain number and tiller number) as well as 1000 metabolomic traits were analyzed. The novelty of the method is the incorporation of the HAT methodology in the 2D BLUP GS model such that the computational efficiency has been dramatically increased by avoiding the conventional cross-validation. The results indicated that (1) the 2D BLUP-HAT GS analysis generally produces higher predictabilities for two traits than those achieved by the analysis of individual traits using 1D GS model, and (2) selected metabolites may be utilized as ancillary traits in the new 2D BLUP-HAT GS method to further boost the predictability of traditional traits, especially for agronomically important traits with low 1D predictabilities.
Shibo Wang 0001, Yang Xu 0023, Han Qu, Yanru Cui, Ruidong Li 0002, John M. Chater, Renyuan Ma, Yiru Qiao, Xuehai Hu, Weibo Xie, Zhenyu Jia
Briefings Bioinform.2
2021 A Computational Framework for Slang Generation
abstract
Abstract Slang is a common type of informal language, but its flexible nature and paucity of data resources present challenges for existing natural language systems. We take an initial step toward machine generation of slang by developing a framework that models the speaker’s word choice in slang context. Our framework encodes novel slang meaning by relating the conventional and slang senses of a word while incorporating syntactic and contextual knowledge in slang usage. We construct the framework using a combination of probabilistic inference and neural contrastive learning. We perform rigorous evaluations on three slang dictionaries and show that our approach not only outperforms state-of-the-art language models, but also better predicts the historical emergence of slang word usages from 1960s to 2000s. We interpret the proposed models and find that the contrastively learned semantic space is sensitive to the similarities between slang and conventional senses of words. Our work creates opportunities for the automated generation and interpretation of informal language.
Zhewei Sun, Richard S. Zemel, Yang Xu 0023
Trans. Assoc. Comput. Linguistics3
2020 Chaining and historical adjective extension
Karan Grewal, Yang Xu 0023
CogSci2
2020 Euphemism and Gender: A Computational Inquiry
Anna Kapron-King, Yang Xu 0023
CogSci2
2020 Chaining and the process of scientific innovation
Emmy Liu, Yang Xu 0023
CogSci2
2020 Grammatical marking and the tradeoff between code length and informativeness
Francis Mollica, Geoff Bacon, Yang Xu 0023, Terry Regier, Charles Kemp
CogSci3
2020 Tracing the Emergence of Gendered Language in Childhood
Ben Prystawski, Erin Grant, Aida Nematzadeh, Spike W. S. Lee, Suzanne Stevenson, Yang Xu 0023
CogSci6
2020 The Typology of Polysemy: A Multilingual Distributional Framework
Ella Rabinovich, Yang Xu 0023, Suzanne Stevenson
CogSci2
2020 Gender convergence in the expressions of love: A computational analysis of lyrics
Lana El Sanyoura, Yang Xu 0023
CogSci2
2020 The Emergence and Propagation of Online Slang
Zhewei Sun, Yizhan Jiang, Yang Xu 0023
CogSci3
2020 Prototype theory and emotion semantic change
Aotao Xu, Jennifer Stellar, Yang Xu 0023
CogSci3
2020 How nouns surface as verbs: Inference and generation in word class conversion
Lana El Sanyoura, Yang Xu 0023
CogSci3
2020 Word class flexibility: A deep contextualized approach
abstract
Word class flexibility refers to the phenomenon whereby a single word form is used across different grammatical categories.Extensive work in linguistic typology has sought to characterize word class flexibility across languages, but quantifying this phenomenon accurately and at scale has been fraught with difficulties.We propose a principled methodology to explore regularity in word class flexibility.Our method builds on recent work in contextualized word embeddings to quantify semantic shift between word classes (e.g., noun-to-verb, verb-to-noun), and we apply this method to 37 languages 1 .We find that contextualized embeddings not only capture human judgment of class variation within words in English, but also uncover shared tendencies in class flexibility across languages.Specifically, we find greater semantic variation when flexible lemmas are used in their dominant word class, supporting the view that word class flexibility is a directional process.Our work highlights the utility of deep contextualized models in linguistic typology.
Guillaume Thomas, Yang Xu 0023, Frank Rudzicz
EMNLP (1)3
2019 Children's overextension as communication by multimodal chaining
Renato Ferreira Pinto Junior, Yang Xu 0023
CogSci2
2019 Rapid information gain explains cross-linguistic tendencies in numeral ordering
Emmy Liu, Yang Xu 0023
CogSci2
2019 Slang Generation as Categorization
Zhewei Sun, Richard S. Zemel, Yang Xu 0023
CogSci3
2019 A predictability-distinctiveness trade-off in the historical emergence of word forms
Aotao Xu, Christian Ramiro, Yang Xu 0023
CogSci3
2019 Slang Detection and Identification
abstract
The prevalence of informal language such as slang presents challenges for natural language systems, particularly in the automatic discovery of flexible word usages.Previous work has explored slang in terms of dictionary construction, sentiment analysis, word formation, and interpretation, but scarce research has attempted the basic problem of slang detection and identification.We examine the extent to which deep learning methods support automatic detection and identification of slang from natural sentences using a combination of bidirectional recurrent neural networks, conditional random field, and multilayer perceptron.We test these models based on a comprehensive set of linguistic features in sentencelevel detection and token-level identification of slang.We found that a prominent feature of slang is the surprising use of words across syntactic categories or syntactic shift (e.g., verb→noun).Our best models detect the presence of slang at the sentence level with an F1-score of 0.80 and identify its exact position at the token level with an F1-Score of 0.50.
Zhengqi Pei, Zhewei Sun, Yang Xu 0023
CoNLL3
2019 Text-based inference of moral sentiment change
abstract
Jing Yi Xie, Renato Ferreira Pinto Junior, Graeme Hirst, Yang Xu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jing Yi Xie, Renato Ferreira Pinto Junior, Graeme Hirst, Yang Xu 0023
EMNLP/IJCNLP (1)4
2018 Stability in the temporal dynamics of word meanings
Yiwei Luo, Yang Xu 0023
CogSci2
2018 Lexical evolution, cognition, and computation
Yang Xu 0023, Barbara Malt, Mahesh Srinivasan
CogSci1
2018 Is Nike female? Exploring the role of sound symbolism in predicting brand name gender
abstract
Are brand names such as Nike female or male?Previous research suggests that the sound of a person's first name is associated with the person's gender, but no research has tried to use this knowledge to assess the gender of brand names.We present a simple computational approach that uses sound symbolism to address this open issue.Consistent with previous research, a model trained on various linguistic features of name endings predicts human gender with high accuracy.Applying this model to a data set of over a thousand commerciallytraded brands in 17 product categories, our results reveal an overall bias toward male names, cutting across both male-oriented product categories as well as female-oriented categories.In addition, we find variation within categories, suggesting that firms might be seeking to imbue their brands with differentiating characteristics as part of their competitive strategy.
Sridhar Moorthy, Ruth Pogacar, Samin Khan, Yang Xu 0023
EMNLP4
2017 Mental Algorithms in the Historical Emergence of Word Meanings
Christian Ramiro, Barbara Malt, Mahesh Srinivasan, Yang Xu 0023
CogSci4
2016 The Sapir-Whorf Hypothesis and Probabilistic Inference: Evidence from the Domain of Color
Emily Cibelli, Yang Xu 0023, Joseph L. Austerweil, Thomas L. Griffiths 0001, Terry Regier
CogSci2
2016 A computational investigation of the Sapir-Whorf hypothesis: The case of spatial relations
Christine Tseng, Alexandra Carstensen, Terry Regier, Yang Xu 0023
CogSci4
2016 The Pragmatics of Spatial Language
Tomer D. Ullman, Yang Xu 0023, Noah D. Goodman
CogSci2
2016 Evolution of polysemous word senses from metaphorical mappings
Yang Xu 0023, Barbara Malt, Mahesh Srinivasan
CogSci1
2015 Tense systems across languages support efficient communication
Geoff Bacon, Yang Xu 0023, Terry Regier
CogSci2
2015 The space of spatial relations: An extended stimulus set
Alexandra Carstensen, Yang Xu 0023, Charles Kemp, Terry Regier
CogSci2
2015 A Computational Evaluation of Two Laws of Semantic Change
Yang Xu 0023, Charles Kemp
CogSci1
2015 Semantic chaining and efficient communication: The case of container names
Yang Xu 0023, Terry Regier, Barbara Malt
CogSci1
2015 An adaptive cue combination model of spatial reorientation
Yang Xu 0023, Terry Regier, Nora S. Newcombe
CogSci1
2014 Numeral systems across languages support efficient communication: From approximate numerosity to recursion
Yang Xu 0023, Terry Regier
CogSci1
2010 Inference and communication in the game of Password
abstract
Communication between a speaker and hearer will be most efficient when both parties make accurate inferences about the other. We study inference and communication in a television game called Password, where speakers must convey secret words to hearers by providing one-word clues. Our working hypothesis is that human communication is relatively efficient, and we use game show data to examine three predictions. First, we predict that speakers and hearers are both considerate, and that both take the other’s perspective into account. Second, we predict that speakers and hearers are calibrated, and that both make accurate assumptions about the strategy used by the other. Finally, we predict that speakers and hearers are collaborative, and that they tend to share the cognitive burden of communication equally. We find evidence in support of all three predictions, and demonstrate in addition that efficient communication tends to break down when speakers and hearers are placed under time pressure.
Yang Xu 0023, Charles Kemp
NIPS1
2009 R/BHC: fast Bayesian hierarchical clustering for microarray data
abstract
BACKGROUND: Although the use of clustering methods has rapidly become one of the standard computational approaches in the literature of microarray gene expression data analysis, little attention has been paid to uncertainty in the results obtained. RESULTS: We present an R/Bioconductor port of a fast novel algorithm for Bayesian agglomerative hierarchical clustering and demonstrate its use in clustering gene expression microarray data. The method performs bottom-up hierarchical clustering, using a Dirichlet Process (infinite mixture) to model uncertainty in the data and Bayesian model selection to decide at each step which clusters to merge. CONCLUSION: Biologically plausible results are presented from a well studied data set: expression profiles of A. thaliana subjected to a variety of biotic and abiotic stresses. Our method avoids several limitations of traditional methods, for example how many clusters there should be and how to choose a principled distance metric.
Richard S. Savage, Katherine A. Heller, Yang Xu 0023, Zoubin Ghahramani, William M. Truman, Murray Grant, Katherine J. Denby, David L. Wild
BMC Bioinform.3