EDBT 2026 Demo / reviewers in the wild / expert
Tyler A. Chang
dblp:265/6105
· DBLP profile ↗
18ranked-venue papers
12as first author
18since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 12 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Goldfish: Monolingual Language Models for 350 LanguagesabstractFor many low-resource languages, the only available language models are large multilingual models trained on many languages simultaneously. Despite state-of-the-art performance on reasoning tasks, we find that these models still struggle with basic grammatical text generation in many languages. First, large multilingual models perform worse than bigrams for many languages (e.g. 24% of languages in XGLM 4.5B; 43% in BLOOM 7.1B) using FLORES perplexity as an evaluation metric. Second, when we train small monolingual models with only 125M parameters on 1GB or less data for 350 languages, these small models outperform large multilingual models both in perplexity and on a massively multilingual grammaticality benchmark. To facilitate future work on low-resource language modeling, we release Goldfish, a suite of over 1,000 small monolingual language models trained comparably for 350 languages. These models represent the first publicly-available monolingual language models for 215 of the languages included. Tyler A. Chang, Catherine Arnett, Zhuowen Tu, Ben Bergen 0001 |
LREC | 1 |
| 2025 | On the Acquisition of Shared Grammatical Representations in Bilingual Language ModelsabstractWhile crosslingual transfer is crucial to contemporary language models’ multilingual capabilities, how it occurs is not well understood. In this paper, we ask what happens to a monolingual language model when it begins to be trained on a second language. Specifically, we train small bilingual models for which we control the amount of data for each language and the order of language exposure. To find evidence of shared multilingual representations, we turn to structural priming, a method used to study grammatical representations in humans. We first replicate previous crosslingual structural priming results and find that after controlling for training data quantity and language exposure, there are asymmetrical effects across language pairs and directions. We argue that this asymmetry may shape hypotheses about human structural priming effects. We also find that structural priming effects are less robust for less similar language pairs, highlighting potential limitations of crosslingual transfer learning and shared representations for typologically diverse languages. Catherine Arnett, Tyler A. Chang, James A. Michaelov, Ben Bergen 0001 |
ACL (1) | 2 |
| 2025 | Scalable Influence and Fact Tracing for Large Language Model PretrainingabstractTraining data attribution (TDA) methods aim to attribute model outputs back to specific training examples, and the application of these methods to large language model (LLM) outputs could significantly advance model transparency and data curation. However, it has been challenging to date to apply these methods to the full scale of LLM pretraining. In this paper, we refine existing gradient-based methods to work effectively at scale, allowing us to retrieve influential examples for an 8B-parameter language model from a pretraining corpus of over 160B tokens with no need for subsampling or pre-filtering. Our method combines several techniques, including optimizer state correction, a task-specific Hessian approximation, and normalized encodings, which we find to be critical for performance at scale. In quantitative evaluations on a fact tracing task, our method performs best at identifying examples that influence model predictions, but classical, model-agnostic retrieval methods such as BM25 still perform better at finding passages which explicitly contain relevant facts. These results demonstrate a misalignment between factual *attribution* and causal *influence*. With increasing model size and training tokens, we find that influence more closely aligns with factual attribution. Finally, we examine different types of examples identified as influential by our method, finding that while many directly entail a particular fact, others support the same output by reinforcing priors on relation types, common entities, and names. We release our prompt set and model outputs, along with a web-based visualization tool to explore influential examples for factual predictions, commonsense reasoning, arithmetic, and open-ended generation for an 8B-parameter LLM. Tyler A. Chang, Dheeraj Rajagopal, Tolga Bolukbasi, Lucas Dixon, Ian Tenney |
ICLR | 1 |
| 2025 | Explaining and Mitigating Crosslingual Tokenizer InequitiesabstractThe number of tokens it takes to encode parallel text in different languages is known to vary. These disparities are called *token premiums*. Having high token premiums leads to less throughput during training and increases costs at inference.
In this paper, we show that even after controlling for dataset size, vocabulary size, and data content, monolingual tokenizers exhibit a wide range of token premiums across languages. To understand the cross-linguistic differences that cause these token premiums,
we train a suite of approximately 7,000 comparable monolingual tokenizers for 97 languages, manipulating tokenization algorithm vocabulary size, and dataset size. We measure token premiums and test for a relationship between factors such as data similarity (between tokenizer training and evaluation), vocabulary size, and pre-tokenization. We also investigate the role of language-specific features such as writing system and word length. We find that similarity between training and test data does not impact token premiums, but vocabulary size and pre-tokenization do. While simply increasing vocabulary size does not lead to reduced token premium effects, we can determine an "optimal" vocabulary size for each language to achieve significantly reduced token premium effects. We also train superword tokenizers which allow merges over whitespaces, and we find that they both reduce token premium effects and improve compression overall. Thus, intervening on the vocabulary size or the pre-tokenizer significantly reduces crosslingual token premium effects. Catherine Arnett, Tyler A. Chang, Stella Biderman, Ben Bergen 0001 |
NeurIPS | 2 |
| 2025 | Bigram Subnetworks: Mapping to Next Tokens in Transformer Language ModelsabstractIn Transformer language models, activation vectors transform from current token embeddings to next token predictions as they pass through the model. To isolate a minimal form of this transformation, we identify language model subnetworks that make bigram predictions, naive next token predictions based only on the current token. We find that bigram subnetworks can be found in fully trained language models up to 1B parameters, and these subnetworks are critical for model performance even when they consist of less than 0.2% of model parameters. Bigram subnetworks are concentrated in the first Transformer MLP layer, and they overlap significantly with subnetworks trained to optimally prune a given model. Mechanistically, the bigram subnetworks often recreate a pattern from the full models where the first layer induces a sharp change that aligns activations with next token predictions rather than current token representations. Our results demonstrate that bigram subnetworks comprise a minimal subset of parameters that are both necessary and sufficient for basic next token predictions in language models, and they help drive the transformation from current to next token activations in the residual stream. These subnetworks can lay a foundation for studying more complex language model circuits by building up from a minimal circuit. Tyler A. Chang, Ben Bergen 0001 |
NeurIPS | 1 |
| 2024 | Detecting Hallucination and Coverage Errors in Retrieval Augmented Generation for Controversial TopicsabstractWe explore a strategy to handle controversial topics in LLM-based chatbots based on Wikipedia’s Neutral Point of View (NPOV) principle: acknowledge the absence of a single true answer and surface multiple perspectives. We frame this as retrieval augmented generation, where perspectives are retrieved from a knowledge base and the LLM is tasked with generating a fluent and faithful response from the given perspectives. As a starting point, we use a deterministic retrieval system and then focus on common LLM failure modes that arise during this approach to text generation, namely hallucination and coverage errors. We propose and evaluate three methods to detect such errors based on (1) word-overlap, (2) salience, and (3) LLM-based classifiers. Our results demonstrate that LLM-based classifiers, even when trained only on synthetic errors, achieve high error detection performance, with ROC AUC scores of 95.3% for hallucination and 90.5% for coverage error detection on unambiguous error cases. We show that when no training data is available, our other methods still yield good results on hallucination (84.0%) and coverage error (85.2%) detection. Tyler A. Chang, Katrin Tomanek, Jessica Hoffmann, Nithum Thain, Erin van Liemt, Kathy Meier-Hellstern, Lucas Dixon |
LREC/COLING | 1 |
| 2024 | Correlations between Multilingual Language Model Geometry and Crosslingual Transfer PerformanceabstractA common approach to interpreting multilingual language models is to evaluate their internal representations. For example, studies have found that languages occupy distinct subspaces in the models’ representation spaces, and geometric distances between languages often reflect linguistic properties such as language families and typological features. In our work, we investigate whether geometric distances between language representations correlate with zero-shot crosslingual transfer performance for POS-tagging and NER in three multilingual language models. We consider four distance metrics, including new metrics that identify a basis for a multilingual representation space that sorts axes based on their language-separability. We find that each distance metric either only moderately correlates or does not correlate with crosslingual transfer performance, and metrics do not generalize well across models, layers, and tasks. Although pairwise language separability is a reasonable predictor of crosslingual transfer, representational geometry overall is an inconsistent predictor for the crosslingual performance of multilingual language models. Cheril Shah, Yashashree Chandak, Atharv Mahesh Mane, Ben Bergen 0001, Tyler A. Chang |
LREC/COLING | 5 |
| 2024 | When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource LanguagesabstractMultilingual language models are widely used to extend NLP systems to low-resource languages.However, concrete evidence for the effects of multilinguality on language modeling performance in individual languages remains scarce.Here, we pre-train over 10,000 monolingual and multilingual language models for over 250 languages, including multiple language families that are under-studied in NLP.We assess how language modeling performance in each language varies as a function of (1) monolingual dataset size, (2) added multilingual dataset size, (3) linguistic similarity of the added languages, and (4) model size (up to 45M parameters).We find that in moderation, adding multilingual data improves low-resource language modeling performance, similar to increasing low-resource dataset sizes by up to 33%.Improvements depend on the syntactic similarity of the added multilingual data, with marginal additional effects of vocabulary overlap.However, high-resource languages consistently perform worse in multilingual pre-training scenarios.As dataset sizes increase, adding multilingual data begins to hurt performance for both low-resource and highresource languages, likely due to limited model capacity (the "curse of multilinguality").These results suggest that massively multilingual pretraining may not be optimal for any languages involved, but that more targeted models can significantly improve performance. Tyler A. Chang, Catherine Arnett, Zhuowen Tu, Ben Bergen 0001 |
EMNLP | 1 |
| 2024 | Language Model Behavior: A Comprehensive SurveyabstractAbstract Transformer language models have received widespread public attention, yet their generated text is often surprising even to NLP researchers. In this survey, we discuss over 250 recent studies of English language model behavior before task-specific fine-tuning. Language models possess basic capabilities in syntax, semantics, pragmatics, world knowledge, and reasoning, but these capabilities are sensitive to specific inputs and surface features. Despite dramatic increases in generated text quality as models scale to hundreds of billions of parameters, the models are still prone to unfactual responses, commonsense errors, memorized text, and social biases. Many of these weaknesses can be framed as over-generalizations or under-generalizations of learned patterns in text. We synthesize recent results to highlight what is currently known about large language model capabilities, thus providing a resource for applied work and for research in adjacent fields that use language models. Tyler A. Chang, Ben Bergen 0001 |
Comput. Linguistics | 1 |
| 2024 | Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and StabilityabstractAbstract How do language models learn to make predictions during pre-training? To study this, we extract learning curves from five autoregressive English language model pre-training runs, for 1M unseen tokens in context. We observe that the language models generate short repetitive phrases before learning to generate longer and more coherent text. We also find that individual tokens often exhibit sudden increases or decreases in loss that are surprisingly consistent across pre-training runs. To better understand these fluctuations, we quantify the final surprisal, within-run variability, age of acquisition, forgettability, and cross-run variability of learning curves for individual tokens in context. More frequent tokens reach lower final surprisals, exhibit less variability within and across pre-training runs, are learned earlier, and are less likely to be “forgotten” during pre-training. Higher n-gram probabilities further accentuate these effects. Independent of the target token, shorter and more frequent contexts correlate with marginally more stable and quickly acquired predictions. Based on our results, we argue for the existence of sequential learning dependencies between different model capabilities, and we characterize language model learning as early n-gram learning before gradual refinement of tail n-gram predictions. Tyler A. Chang, Zhuowen Tu, Ben Bergen 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2023 | Characterizing and Measuring Linguistic Dataset DriftabstractTyler A. Chang, Kishaloy Halder, Neha Anna John, Yogarshi Vyas, Yassine Benajiba, Miguel Ballesteros, Dan Roth. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Tyler A. Chang, Kishaloy Halder, Neha Anna John, Yogarshi Vyas, Yassine Benajiba, Miguel Ballesteros, Dan Roth 0001 |
ACL (1) | 1 |
| 2023 | Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language Modelsabstractgrammatical knowledge-of parts of speech and grammatical patterns-is key to the capacity for linguistic generalization in humans.But how abstract is grammatical knowledge in large language models?In the human literature, compelling evidence for grammatical abstraction comes from structural priming.A sentence that shares the same grammatical structure as a preceding sentence is processed and produced more readily.Because confounds exist when using stimuli in a single language, evidence of abstraction is even more compelling from crosslingual structural priming, where use of a syntactic structure in one language primes an analogous structure in another language.We measure crosslingual structural priming in large language models, comparing model behavior to human experimental results from eight crosslingual experiments covering six languages, and four monolingual structural priming experiments in three non-English languages.We find evidence for abstract monolingual and crosslingual grammatical representations in the models that function similarly to those found in humans.These results demonstrate that grammatical representations in multilingual language models are not only similar across languages, but they can causally influence text produced in different languages. James A. Michaelov, Catherine Arnett, Tyler A. Chang, Ben Bergen 0001 |
EMNLP | 3 |
| 2022 | Does Contextual Diversity Hinder Early Word Acquisition?
Tyler A. Chang, Ben Bergen 0001 |
CogSci | 1 |
| 2022 | Distrubutional Semantics Still Can't Account for Affordances
Cameron R. Jones, Tyler A. Chang, Seana Coulson, James A. Michaelov, Sean Trott, Ben Bergen 0001 |
CogSci | 2 |
| 2022 | The Geometry of Multilingual Language Model RepresentationsabstractWe assess how multilingual language models maintain a shared multilingual representation space while still encoding language-sensitive information in each language.Using XLM-R as a case study, we show that languages occupy similar linear subspaces after mean-centering, evaluated based on causal effects on language modeling performance and direct comparisons between subspaces for 88 languages.The subspace means differ along language-sensitive axes that are relatively stable throughout middle layers, and these axes encode information such as token vocabularies.Shifting representations by language means is sufficient to induce token predictions in different languages.However, we also identify stable languageneutral axes that encode information such as token positions and part-of-speech.We visualize representations projected onto languagesensitive and language-neutral axes, identifying language family and part-of-speech clusters, along with spirals, toruses, and curves representing token position information.These results demonstrate that multilingual language models encode information along orthogonal language-sensitive and language-neutral axes, allowing the models to extract a variety of features for downstream tasks and cross-lingual transfer learning. Tyler A. Chang, Zhuowen Tu, Ben Bergen 0001 |
EMNLP | 1 |
| 2022 | Word Acquisition in Neural Language ModelsabstractAbstract We investigate how neural language models acquire individual words during training, extracting learning curves and ages of acquisition for over 600 words on the MacArthur-Bates Communicative Development Inventory (Fenson et al., 2007). Drawing on studies of word acquisition in children, we evaluate multiple predictors for words’ ages of acquisition in LSTMs, BERT, and GPT-2. We find that the effects of concreteness, word length, and lexical class are pointedly different in children and language models, reinforcing the importance of interaction and sensorimotor experience in child language acquisition. Language models rely far more on word frequency than children, but, like children, they exhibit slower learning of words in longer utterances. Interestingly, models follow consistent patterns during training for both unidirectional and bidirectional models, and for both LSTM and Transformer architectures. Models predict based on unigram token frequencies early in training, before transitioning loosely to bigram probabilities, eventually converging on more nuanced predictions. These results shed light on the role of distributional learning mechanisms in children, while also providing insights for more human-like language acquisition in language models. Tyler A. Chang, Ben Bergen 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | Convolutions and Self-Attention: Re-interpreting Relative Positions in Pre-trained Language ModelsabstractTyler Chang, Yifan Xu, Weijian Xu, Zhuowen Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Tyler A. Chang, Yifan Xu 0009, Weijian Xu, Zhuowen Tu |
ACL/IJCNLP (1) | 1 |
| 2021 | Co-Scale Conv-Attentional Image TransformersabstractIn this paper, we present Co-scale conv-attentional image Transformers (CoaT), a Transformer-based image classifier equipped with co-scale and conv-attentional mechanisms. First, the co-scale mechanism maintains the integrity of Transformers’ encoder branches at individual scales, while allowing representations learned at different scales to effectively communicate with each other; we design a series of serial and parallel blocks to realize the co-scale mechanism. Second, we devise a conv-attentional mechanism by realizing a relative position embedding formulation in the factorized attention module with an efficient convolution-like implementation. CoaT empowers image Transformers with enriched multi-scale and contextual modeling capabilities. On ImageNet, relatively small CoaT models attain superior classification results compared with similar-sized convolutional neural networks and image/vision Transformers. The effectiveness of CoaT’s backbone is also illustrated on object detection and instance segmentation, demonstrating its applicability to downstream computer vision tasks. Weijian Xu, Yifan Xu 0009, Tyler A. Chang, Zhuowen Tu |
ICCV | 3 |