VLDB 2026 Research / reviewers in the wild / expert
Tom McCoy 0001
dblp:205/9005 · also R. Thomas McCoy, Richard Thomas McCoy
· DBLP profile ↗
23ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0001-9383-4079ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 7 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Collocational bootstrapping: A hypothesis about the learning of subject-verb agreement in humans and neural networksabstractIn what ways might statistical signals in linguistic input assist with the acquisition of syntax? Here we hypothesize a mechanism called collocational bootstrapping, in which regularities in word co-occurrence patterns can provide cues to syntactic dependencies. We investigate whether this mechanism can support the acquisition of English subject-verb agreement. First, we simulate language acquisition by training neural networks on synthetic datasets that vary in how predictable their subject-verb pairings are. We find that there is a range of variability levels at which these statistical learners robustly learn subject-verb agreement. We then analyze the variability of subject-verb pairings in child-directed language, and we find that the variability in such data falls within the range that supported robust generalization in our computational simulations. Taken together, these results suggest that collocational bootstrapping is a viable learning strategy for the type of input that children receive. Claire Hobbs, Tom McCoy 0001 |
CoNLL | 2 |
| 2026 | What Exactly do Children Receive in Language Acquisition? A Case Study on CHILDES with Automated Detection of Filler-Gap DependenciesabstractChildren's acquisition of filler-gap dependencies has been argued by some to depend on innate grammatical knowledge, while others suggest that the distributional evidence available in child-directed speech suffices.Unfortunately, the relevant input is difficult to quantify at scale with fine granularity, making this question difficult to resolve.We present a system that identifies three core filler-gap constructions in spoken English corpora -matrix whquestions, embedded wh-questions, and relative clauses -and further identifies the extraction site (i.e., subject vs. object vs. adjunct).Our approach combines constituency and dependency parsing, leveraging their complementary strengths for construction classification and extraction site identification.We validate the system on human-annotated data and find that it scores well across most categories.Applying the system to 57 English CHILDES corpora, we are able to characterize children's fillergap input and their filler-gap production trajectories over the course of development, including construction-specific frequencies and extraction-site asymmetries.The resulting finegrained labels enable future work in both acquisition and computational studies, which we demonstrate with a case study using filtered corpus training with language models. Zhenghao Zhou, William Dai, Maya Viswanathan, Simon Charlow, Tom McCoy 0001, Robert Frank 0001 |
CoNLL | 5 |
| 2025 | Teasing Apart Architecture and Initial Weights as Sources of Inductive Bias in Neural Networks
Gianluca M. Bencomo, Max Gupta, Ioana Marinescu, Tom McCoy 0001, Thomas L. Griffiths 0001 |
CogSci | 4 |
| 2025 | Convolutional Neural Networks Can (Meta-)Learn the Same-Different Relation
Max Gupta, Sunayana Rane, Tom McCoy 0001, Thomas L. Griffiths 0001 |
CogSci | 3 |
| 2025 | Additive Analogies Reveal Compositional Structure in Neural Network Weights
Abi Tenenbaum, Tom McCoy 0001 |
CogSci | 2 |
| 2025 | Is In-Context Learning a Type of Error-Driven Learning? Evidence from the Inverse Frequency Effect in Structural PrimingabstractZhenghao Zhou, Robert Frank, R. Thomas McCoy. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zhenghao Zhou, Robert Frank 0001, Tom McCoy 0001 |
NAACL (Long Papers) | 3 |
| 2025 | Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language ModelsabstractAccuracy remains a standard metric for evaluating AI systems, but it offers limited insight into how models arrive at their solutions.
In this work, we introduce a benchmark based on brainteasers written in long narrative form to probe more deeply into the types of reasoning strategies that models use. Brainteasers are well-suited for this goal because they can be solved with multiple approaches, such as a few-step solution that uses a creative insight or a longer solution that uses more brute force.
We investigate large language models (LLMs) across multiple layers of reasoning, focusing not only on correctness but also on the quality and creativity of their solutions.
We investigate many aspects of the reasoning process: (1) semantic parsing of the brainteasers into precise mathematical competition style formats; (2) self-correcting solutions based on gold solutions; (3) producing step-by-step sketches of solutions; and (4) making use of hints.
We find that LLMs are in many cases able to find creative, insightful solutions to brainteasers, suggesting that they capture some of the capacities needed to solve novel problems in creative ways. Nonetheless, there also remain situations where they rely on brute force despite the availability of more efficient, creative solutions, highlighting a potential direction for improvement in the reasoning abilities of LLMs. Simeng Han, Howard Dai, Stephen Xia, Grant Zhang, Chen Liu 0020, Lichang Chen, Hongyuan Mei, Jiayuan Mao, Tom McCoy 0001 |
NeurIPS | 10 |
| 2024 | Distilling Symbolic Priors for Concept Learning into Neural Networks
Ioana Marinescu, Tom McCoy 0001, Thomas L. Griffiths 0001 |
CogSci | 2 |
| 2023 | How poor is the stimulus? Evaluating hierarchical generalization in neural networks trained on child-directed speechabstractWhen acquiring syntax, children consistently choose hierarchical rules over competing nonhierarchical possibilities.Is this preference due to a learning bias for hierarchical structure, or due to more general biases that interact with hierarchical cues in children's linguistic input?We explore these possibilities by training LSTMs and Transformers-two types of neural networks without a hierarchical biason data similar in quantity and content to children's linguistic input: text from the CHILDES corpus.We then evaluate what these models have learned about English yes/no questions, a phenomenon for which hierarchical structure is crucial.We find that, though they perform well at capturing the surface statistics of childdirected speech (as measured by perplexity), both model types generalize in a way more consistent with an incorrect linear rule than the correct hierarchical rule.These results suggest that human-like generalization from text alone requires stronger biases than the general sequence-processing biases of standard neural network architectures. Aditya Yedetore, Tal Linzen, Robert Frank 0001, Tom McCoy 0001 |
ACL (1) | 4 |
| 2023 | How Much Do Language Models Copy From Their Training Data? Evaluating Linguistic Novelty in Text Generation Using RAVENabstractAbstract Current language models can generate high-quality text. Are they simply copying text they have seen before, or have they learned generalizable linguistic abstractions? To tease apart these possibilities, we introduce RAVEN, a suite of analyses for assessing the novelty of generated text, focusing on sequential structure (n-grams) and syntactic structure. We apply these analyses to four neural language models trained on English (an LSTM, a Transformer, Transformer-XL, and GPT-2). For local structure—e.g., individual dependencies—text generated with a standard sampling scheme is substantially less novel than our baseline of human-generated text from each model’s test set. For larger-scale structure—e.g., overall sentence structure—model-generated text is as novel or even more novel than the human-generated baseline, but models still sometimes copy substantially, in some cases duplicating passages over 1,000 words long from the training set. We also perform extensive manual analysis, finding evidence that GPT-2 uses both compositional and analogical generalization mechanisms and showing that GPT-2’s novel text is usually well-formed morphologically and syntactically but has reasonably frequent semantic issues (e.g., being self-contradictory). Tom McCoy 0001, Paul Smolensky, Tal Linzen, Jianfeng Gao 0001, Asli Celikyilmaz |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | Infinite use of finite means? Evaluating the generalization of center embedding learned from an artificial grammar
Tom McCoy 0001, Jennifer Culbertson, Paul Smolensky, Geraldine Legendre |
CogSci | 1 |
| 2020 | Representations of Syntax [MASK] Useful: Effects of Constituency and Dependency Structure in Recursive LSTMsabstractSequence-based neural networks show significant sensitivity to syntactic structure, but they still perform less well on syntactic tasks than tree-based networks.Such tree-based networks can be provided with a constituency parse, a dependency parse, or both.We evaluate which of these two representational schemes more effectively introduces biases for syntactic structure that increase performance on the subject-verb agreement prediction task.We find that a constituency-based network generalizes more robustly than a dependencybased one, and that combining the two types of structure does not yield further improvement.Finally, we show that the syntactic robustness of sequential models can be substantially improved by fine-tuning on a small amount of constructed data, suggesting that data augmentation is a viable alternative to explicit constituency structure for imparting the syntactic biases that sequential models are lacking. Michael A. Lepori, Tal Linzen, Tom McCoy 0001 |
ACL | 3 |
| 2020 | Syntactic Data Augmentation Increases Robustness to Inference HeuristicsabstractPretrained neural models such as BERT, when fine-tuned to perform natural language inference (NLI), often show high accuracy on standard datasets, but display a surprising lack of sensitivity to word order on controlled challenge sets.We hypothesize that this issue is not primarily caused by the pretrained model's limitations, but rather by the paucity of crowdsourced NLI examples that might convey the importance of syntactic structure at the finetuning stage.We explore several methods to augment standard training sets with syntactically informative examples, generated by applying syntactic transformations to sentences from the MNLI corpus.The best-performing augmentation method, subject/object inversion, improved BERT's accuracy on controlled examples that diagnose sensitivity to word order from 0.28 to 0.73, without affecting performance on the MNLI test set.This improvement generalized beyond the particular construction used for data augmentation, suggesting that augmentation causes BERT to recruit abstract syntactic representations. Junghyun Min, Tom McCoy 0001, Dipanjan Das 0001, Emily Pitler, Tal Linzen |
ACL | 2 |
| 2020 | Universal linguistic inductive biases via meta-learning
Tom McCoy 0001, Erin Grant, Paul Smolensky, Thomas L. Griffiths 0001, Tal Linzen |
CogSci | 1 |
| 2020 | Picking BERT's Brain: Probing for Linguistic Dependencies in Contextualized Embeddings Using Representational Similarity AnalysisabstractAs the name implies, contextualized representations of language are typically motivated by their ability to encode context.Which aspects of context are captured by such representations?We introduce an approach to address this question using Representational Similarity Analysis (RSA).As case studies, we investigate the degree to which a verb embedding encodes the verb's subject, a pronoun embedding encodes the pronoun's antecedent, and a full-sentence representation encodes the sentence's head word (as determined by a dependency parse).In all cases, we show that BERT's contextualized embeddings reflect the linguistic dependency being studied, and that BERT encodes these dependencies to a greater degree than it encodes less linguistically-salient controls.These results demonstrate the ability of our approach to adjudicate between hypotheses about which aspects of context are encoded in representations of language. Michael A. Lepori, Tom McCoy 0001 |
COLING | 2 |
| 2020 | Does Syntax Need to Grow on Trees? Sources of Hierarchical Inductive Bias in Sequence-to-Sequence NetworksabstractLearners that are exposed to the same training data might generalize differently due to differing inductive biases. In neural network models, inductive biases could in theory arise from any aspect of the model architecture. We investigate which architectural factors affect the generalization behavior of neural sequence-to-sequence models trained on two syntactic tasks, English question formation and English tense reinflection. For both tasks, the training set is consistent with a generalization based on hierarchical structure and a generalization based on linear order. All architectural factors that we investigated qualitatively affected how models generalized, including factors with no clear connection to hierarchical structure. For example, LSTMs and GRUs displayed qualitatively different inductive biases. However, the only factor that consistently contributed a hierarchical bias across tasks was the use of a tree-structured model rather than a model with sequential recurrence, suggesting that human-like syntactic generalization requires architectural syntactic structure. Tom McCoy 0001, Robert Frank 0001, Tal Linzen |
Trans. Assoc. Comput. Linguistics | 1 |
| 2019 | Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language InferenceabstractA machine learning system can score well on a given test set by relying on heuristics that are effective for frequent example types but break down in more challenging cases.We study this issue within natural language inference (NLI), the task of determining whether one sentence entails another.We hypothesize that statistical NLI models may adopt three fallible syntactic heuristics: the lexical overlap heuristic, the subsequence heuristic, and the constituent heuristic.To determine whether models have adopted these heuristics, we introduce a controlled evaluation set called HANS (Heuristic Analysis for NLI Systems), which contains many examples where the heuristics fail.We find that models trained on MNLI, including BERT, a state-of-the-art model, perform very poorly on HANS, suggesting that they have indeed adopted these heuristics.We conclude that there is substantial room for improvement in NLI systems, and that the HANS dataset can motivate and measure progress in this area. Tom McCoy 0001, Ellie Pavlick, Tal Linzen |
ACL (1) | 1 |
| 2019 | Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language ModelingabstractAlex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Jan Hula, Patrick Xia 0002, Raghavendra Pappagari, Tom McCoy 0001, Roma Patel, Najoung Kim, Ian Tenney, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman |
ACL (1) | 5 |
| 2019 | RNNs implicitly implement tensor-product representations
Tom McCoy 0001, Tal Linzen, Ewan Dunbar, Paul Smolensky |
ICLR (Poster) | 1 |
| 2019 | What do you learn from context? Probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia 0002, Berlin Chen, Adam Poliak, Tom McCoy 0001, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das 0001, Ellie Pavlick |
ICLR (Poster) | 6 |
| 2018 | Revisiting the poverty of the stimulus: hierarchical generalization without a hierarchical bias in recurrent neural networks
Tom McCoy 0001, Robert Frank 0001, Tal Linzen |
CogSci | 1 |
| 2018 | Parser combinators for Tigrinya and Oromo morphology
Patrick Littell, Tom McCoy 0001, Na-Rae Han, Shruti Rijhwani, Zaid Sheikh, David R. Mortensen, Teruko Mitamura, Lori S. Levin |
LREC | 2 |
| 2017 | TAG Parsing with Neural Networks and Vector Representations of SupertagsabstractWe present supertagging-based models for Tree Adjoining Grammar parsing that use neural network architectures and dense vector representation of supertags (elementary trees) to achieve state-of-the-art performance in unlabeled and labeled attachment scores.The shift-reduce parsing model eschews lexical information entirely, and uses only the 1-best supertags to parse a sentence, providing further support for the claim that supertagging is "almost parsing."We demonstrate that the embedding vector representations the parser induces for supertags possess linguistically interpretable structure, supporting analogies between grammatical structures like those familiar from recent work in distributional semantics.This dense representation of supertags overcomes the drawbacks for statistical models of TAG as compared to CCG parsing, raising the possibility that TAG is a viable alternative for NLP tasks that require the assignment of richer structural descriptions to sentences. Jungo Kasai, Robert Frank 0001, Tom McCoy 0001, Owen Rambow, Alexis Nasr |
EMNLP | 3 |