Tom McCoy 0001

dblp:205/9005 · also R. Thomas McCoy, Richard Thomas McCoy · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0001-9383-4079ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 7 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Collocational bootstrapping: A hypothesis about the learning of subject-verb agreement in humans and neural networks
abstract
In what ways might statistical signals in linguistic input assist with the acquisition of syntax? Here we hypothesize a mechanism called collocational bootstrapping, in which regularities in word co-occurrence patterns can provide cues to syntactic dependencies. We investigate whether this mechanism can support the acquisition of English subject-verb agreement. First, we simulate language acquisition by training neural networks on synthetic datasets that vary in how predictable their subject-verb pairings are. We find that there is a range of variability levels at which these statistical learners robustly learn subject-verb agreement. We then analyze the variability of subject-verb pairings in child-directed language, and we find that the variability in such data falls within the range that supported robust generalization in our computational simulations. Taken together, these results suggest that collocational bootstrapping is a viable learning strategy for the type of input that children receive.
Claire Hobbs, Tom McCoy 0001
CoNLL2
2026 What Exactly do Children Receive in Language Acquisition? A Case Study on CHILDES with Automated Detection of Filler-Gap Dependencies
abstract
Children's acquisition of filler-gap dependencies has been argued by some to depend on innate grammatical knowledge, while others suggest that the distributional evidence available in child-directed speech suffices.Unfortunately, the relevant input is difficult to quantify at scale with fine granularity, making this question difficult to resolve.We present a system that identifies three core filler-gap constructions in spoken English corpora -matrix whquestions, embedded wh-questions, and relative clauses -and further identifies the extraction site (i.e., subject vs. object vs. adjunct).Our approach combines constituency and dependency parsing, leveraging their complementary strengths for construction classification and extraction site identification.We validate the system on human-annotated data and find that it scores well across most categories.Applying the system to 57 English CHILDES corpora, we are able to characterize children's fillergap input and their filler-gap production trajectories over the course of development, including construction-specific frequencies and extraction-site asymmetries.The resulting finegrained labels enable future work in both acquisition and computational studies, which we demonstrate with a case study using filtered corpus training with language models.
Zhenghao Zhou, William Dai, Maya Viswanathan, Simon Charlow, Tom McCoy 0001, Robert Frank 0001
CoNLL5
2025 Teasing Apart Architecture and Initial Weights as Sources of Inductive Bias in Neural Networks
Gianluca M. Bencomo, Max Gupta, Ioana Marinescu, Tom McCoy 0001, Thomas L. Griffiths 0001
CogSci4
2025 Convolutional Neural Networks Can (Meta-)Learn the Same-Different Relation
Max Gupta, Sunayana Rane, Tom McCoy 0001, Thomas L. Griffiths 0001
CogSci3
2025 Additive Analogies Reveal Compositional Structure in Neural Network Weights
Abi Tenenbaum, Tom McCoy 0001
CogSci2
2025 Is In-Context Learning a Type of Error-Driven Learning? Evidence from the Inverse Frequency Effect in Structural Priming
abstract
Zhenghao Zhou, Robert Frank, R. Thomas McCoy. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zhenghao Zhou, Robert Frank 0001, Tom McCoy 0001
NAACL (Long Papers)3
2025 Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language Models
abstract
Accuracy remains a standard metric for evaluating AI systems, but it offers limited insight into how models arrive at their solutions. In this work, we introduce a benchmark based on brainteasers written in long narrative form to probe more deeply into the types of reasoning strategies that models use. Brainteasers are well-suited for this goal because they can be solved with multiple approaches, such as a few-step solution that uses a creative insight or a longer solution that uses more brute force. We investigate large language models (LLMs) across multiple layers of reasoning, focusing not only on correctness but also on the quality and creativity of their solutions. We investigate many aspects of the reasoning process: (1) semantic parsing of the brainteasers into precise mathematical competition style formats; (2) self-correcting solutions based on gold solutions; (3) producing step-by-step sketches of solutions; and (4) making use of hints. We find that LLMs are in many cases able to find creative, insightful solutions to brainteasers, suggesting that they capture some of the capacities needed to solve novel problems in creative ways. Nonetheless, there also remain situations where they rely on brute force despite the availability of more efficient, creative solutions, highlighting a potential direction for improvement in the reasoning abilities of LLMs.
Simeng Han, Howard Dai, Stephen Xia, Grant Zhang, Chen Liu 0020, Lichang Chen, Hongyuan Mei, Jiayuan Mao, Tom McCoy 0001
NeurIPS10
2024 Distilling Symbolic Priors for Concept Learning into Neural Networks
Ioana Marinescu, Tom McCoy 0001, Thomas L. Griffiths 0001
CogSci2
2023 How poor is the stimulus? Evaluating hierarchical generalization in neural networks trained on child-directed speech
abstract
When acquiring syntax, children consistently choose hierarchical rules over competing nonhierarchical possibilities.Is this preference due to a learning bias for hierarchical structure, or due to more general biases that interact with hierarchical cues in children's linguistic input?We explore these possibilities by training LSTMs and Transformers-two types of neural networks without a hierarchical biason data similar in quantity and content to children's linguistic input: text from the CHILDES corpus.We then evaluate what these models have learned about English yes/no questions, a phenomenon for which hierarchical structure is crucial.We find that, though they perform well at capturing the surface statistics of childdirected speech (as measured by perplexity), both model types generalize in a way more consistent with an incorrect linear rule than the correct hierarchical rule.These results suggest that human-like generalization from text alone requires stronger biases than the general sequence-processing biases of standard neural network architectures.
Aditya Yedetore, Tal Linzen, Robert Frank 0001, Tom McCoy 0001
ACL (1)4
2023 How Much Do Language Models Copy From Their Training Data? Evaluating Linguistic Novelty in Text Generation Using RAVEN
abstract
Abstract Current language models can generate high-quality text. Are they simply copying text they have seen before, or have they learned generalizable linguistic abstractions? To tease apart these possibilities, we introduce RAVEN, a suite of analyses for assessing the novelty of generated text, focusing on sequential structure (n-grams) and syntactic structure. We apply these analyses to four neural language models trained on English (an LSTM, a Transformer, Transformer-XL, and GPT-2). For local structure—e.g., individual dependencies—text generated with a standard sampling scheme is substantially less novel than our baseline of human-generated text from each model’s test set. For larger-scale structure—e.g., overall sentence structure—model-generated text is as novel or even more novel than the human-generated baseline, but models still sometimes copy substantially, in some cases duplicating passages over 1,000 words long from the training set. We also perform extensive manual analysis, finding evidence that GPT-2 uses both compositional and analogical generalization mechanisms and showing that GPT-2’s novel text is usually well-formed morphologically and syntactically but has reasonably frequent semantic issues (e.g., being self-contradictory).
Tom McCoy 0001, Paul Smolensky, Tal Linzen, Jianfeng Gao 0001, Asli Celikyilmaz
Trans. Assoc. Comput. Linguistics1
2021 Infinite use of finite means? Evaluating the generalization of center embedding learned from an artificial grammar
Tom McCoy 0001, Jennifer Culbertson, Paul Smolensky, Geraldine Legendre
CogSci1
2020 Representations of Syntax [MASK] Useful: Effects of Constituency and Dependency Structure in Recursive LSTMs
abstract
Sequence-based neural networks show significant sensitivity to syntactic structure, but they still perform less well on syntactic tasks than tree-based networks.Such tree-based networks can be provided with a constituency parse, a dependency parse, or both.We evaluate which of these two representational schemes more effectively introduces biases for syntactic structure that increase performance on the subject-verb agreement prediction task.We find that a constituency-based network generalizes more robustly than a dependencybased one, and that combining the two types of structure does not yield further improvement.Finally, we show that the syntactic robustness of sequential models can be substantially improved by fine-tuning on a small amount of constructed data, suggesting that data augmentation is a viable alternative to explicit constituency structure for imparting the syntactic biases that sequential models are lacking.
Michael A. Lepori, Tal Linzen, Tom McCoy 0001
ACL3
2020 Syntactic Data Augmentation Increases Robustness to Inference Heuristics
abstract
Pretrained neural models such as BERT, when fine-tuned to perform natural language inference (NLI), often show high accuracy on standard datasets, but display a surprising lack of sensitivity to word order on controlled challenge sets.We hypothesize that this issue is not primarily caused by the pretrained model's limitations, but rather by the paucity of crowdsourced NLI examples that might convey the importance of syntactic structure at the finetuning stage.We explore several methods to augment standard training sets with syntactically informative examples, generated by applying syntactic transformations to sentences from the MNLI corpus.The best-performing augmentation method, subject/object inversion, improved BERT's accuracy on controlled examples that diagnose sensitivity to word order from 0.28 to 0.73, without affecting performance on the MNLI test set.This improvement generalized beyond the particular construction used for data augmentation, suggesting that augmentation causes BERT to recruit abstract syntactic representations.
Junghyun Min, Tom McCoy 0001, Dipanjan Das 0001, Emily Pitler, Tal Linzen
ACL2
2020 Universal linguistic inductive biases via meta-learning
Tom McCoy 0001, Erin Grant, Paul Smolensky, Thomas L. Griffiths 0001, Tal Linzen
CogSci1
2020 Picking BERT's Brain: Probing for Linguistic Dependencies in Contextualized Embeddings Using Representational Similarity Analysis
abstract
As the name implies, contextualized representations of language are typically motivated by their ability to encode context.Which aspects of context are captured by such representations?We introduce an approach to address this question using Representational Similarity Analysis (RSA).As case studies, we investigate the degree to which a verb embedding encodes the verb's subject, a pronoun embedding encodes the pronoun's antecedent, and a full-sentence representation encodes the sentence's head word (as determined by a dependency parse).In all cases, we show that BERT's contextualized embeddings reflect the linguistic dependency being studied, and that BERT encodes these dependencies to a greater degree than it encodes less linguistically-salient controls.These results demonstrate the ability of our approach to adjudicate between hypotheses about which aspects of context are encoded in representations of language.
Michael A. Lepori, Tom McCoy 0001
COLING2
2020 Does Syntax Need to Grow on Trees? Sources of Hierarchical Inductive Bias in Sequence-to-Sequence Networks
abstract
Learners that are exposed to the same training data might generalize differently due to differing inductive biases. In neural network models, inductive biases could in theory arise from any aspect of the model architecture. We investigate which architectural factors affect the generalization behavior of neural sequence-to-sequence models trained on two syntactic tasks, English question formation and English tense reinflection. For both tasks, the training set is consistent with a generalization based on hierarchical structure and a generalization based on linear order. All architectural factors that we investigated qualitatively affected how models generalized, including factors with no clear connection to hierarchical structure. For example, LSTMs and GRUs displayed qualitatively different inductive biases. However, the only factor that consistently contributed a hierarchical bias across tasks was the use of a tree-structured model rather than a model with sequential recurrence, suggesting that human-like syntactic generalization requires architectural syntactic structure.
Tom McCoy 0001, Robert Frank 0001, Tal Linzen
Trans. Assoc. Comput. Linguistics1
2019 Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
abstract
A machine learning system can score well on a given test set by relying on heuristics that are effective for frequent example types but break down in more challenging cases.We study this issue within natural language inference (NLI), the task of determining whether one sentence entails another.We hypothesize that statistical NLI models may adopt three fallible syntactic heuristics: the lexical overlap heuristic, the subsequence heuristic, and the constituent heuristic.To determine whether models have adopted these heuristics, we introduce a controlled evaluation set called HANS (Heuristic Analysis for NLI Systems), which contains many examples where the heuristics fail.We find that models trained on MNLI, including BERT, a state-of-the-art model, perform very poorly on HANS, suggesting that they have indeed adopted these heuristics.We conclude that there is substantial room for improvement in NLI systems, and that the HANS dataset can motivate and measure progress in this area.
Tom McCoy 0001, Ellie Pavlick, Tal Linzen
ACL (1)1
2019 Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language Modeling
abstract
Alex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Jan Hula, Patrick Xia 0002, Raghavendra Pappagari, Tom McCoy 0001, Roma Patel, Najoung Kim, Ian Tenney, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman
ACL (1)5
2019 RNNs implicitly implement tensor-product representations
Tom McCoy 0001, Tal Linzen, Ewan Dunbar, Paul Smolensky
ICLR (Poster)1
2019 What do you learn from context? Probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia 0002, Berlin Chen, Adam Poliak, Tom McCoy 0001, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das 0001, Ellie Pavlick
ICLR (Poster)6
2018 Revisiting the poverty of the stimulus: hierarchical generalization without a hierarchical bias in recurrent neural networks
Tom McCoy 0001, Robert Frank 0001, Tal Linzen
CogSci1
2018 Parser combinators for Tigrinya and Oromo morphology
Patrick Littell, Tom McCoy 0001, Na-Rae Han, Shruti Rijhwani, Zaid Sheikh, David R. Mortensen, Teruko Mitamura, Lori S. Levin
LREC2
2017 TAG Parsing with Neural Networks and Vector Representations of Supertags
abstract
We present supertagging-based models for Tree Adjoining Grammar parsing that use neural network architectures and dense vector representation of supertags (elementary trees) to achieve state-of-the-art performance in unlabeled and labeled attachment scores.The shift-reduce parsing model eschews lexical information entirely, and uses only the 1-best supertags to parse a sentence, providing further support for the claim that supertagging is "almost parsing."We demonstrate that the embedding vector representations the parser induces for supertags possess linguistically interpretable structure, supporting analogies between grammatical structures like those familiar from recent work in distributional semantics.This dense representation of supertags overcomes the drawbacks for statistical models of TAG as compared to CCG parsing, raising the possibility that TAG is a viable alternative for NLP tasks that require the assignment of richer structural descriptions to sentences.
Jungo Kasai, Robert Frank 0001, Tom McCoy 0001, Owen Rambow, Alexis Nasr
EMNLP3