Thomas Hikaru Clark

dblp:301/9504 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-1720-0450ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 48% Probabilistic and Bayesian machine learning · 45% Knowledge representation and reasoning · 7%
Databases, data mining, and information retrieval
1 paper
Knowledge graphs · 100%

Topics — the 4 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning
probabilistic inference
0.912025
Resource-Rational Noisy-Channel Language Processing: Testing the Effect of Algorithmic Constraints on Inferences · EMNLP 2025
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
sequential monte carlo
0.912025
Resource-Rational Noisy-Channel Language Processing: Testing the Effect of Algorithmic Constraints on Inferences · EMNLP 2025
Natural language and speech › Language models and text generation › large language model › knowledge in language models
knowledge injection
0.512021
Mixture-of-Partitions: Infusing Large Biomedical Knowledge Graphs into BERT · EMNLP (1) 2021
Natural language and speech › Language models and text generation
pre-trained language model
0.512021
Mixture-of-Partitions: Infusing Large Biomedical Knowledge Graphs into BERT · EMNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

mixture of experts · 1.0adapter · 1.0sequential monte carlo · 0.9rejuvenation · 0.9language model · 0.9
YearPublicationVenuePosition
2026 Readers make targeted regressions to plausible errors in reanalysis of "noisy-channel garden-path" sentences
abstract
A key question in psycholinguistics is how inferences about the meaning of linguistic input unfold incrementally a comprehender's mind.In this work, we study reading dynamics for "noisy-channel garden-path" sentences, which temporarily appear well-formed but feature late-appearing violations of expectation that can be resolved not by inferring an alternative syntactic structure, but by inferring the presence of an error.We find evidence for targeted regressions -eye movements towards regions that are promising loci of possible errors in light of later-arriving information, showing patterns consistent with the posterior inferences of a model of noisy-channel processing with reanalysis.We discuss the implications of these findings for theories of noisy-channel language comprehension and information-theoretic explanations of reading dynamics.
Thomas Hikaru Clark, Roger Levy, Edward Gibson
CoNLL1
2025 A Model of Approximate and Incremental Noisy-Channel Language Processing
Thomas Hikaru Clark, Jacob Hoover Vigly, Edward Gibson, Roger Levy
CogSci1
2025 Resource-Rational Noisy-Channel Language Processing: Testing the Effect of Algorithmic Constraints on Inferences
abstract
Human language use is robust to errors: comprehenders can and do mentally correct utterances that are implausible or anomalous.How are humans able to solve these problems in real time, picking out alternatives from an unbounded space of options using limited cognitive resources?And can language models trained on next-word prediction for typical language be augmented to handle language anomalies in a human-like way?Using a language model as a prior and an error model to encode likelihoods, we use Sequential Monte Carlo with optional rejuvenation to perform incremental and approximate probabilistic inference over intended sentences and production errors.We demonstrate that the model captures previously established patterns in human sentence processing, and that a trade-off between human-like noisy-channel inferences and computational resources falls out of this model.From a psycholinguistic perspective, our results offer a candidate algorithmic model of rational inference in language processing.From an NLP perspective, our results showcase how to elicit human-like noisy-channel inference behavior from a relatively small LLM while controlling the amount of computation available during inference.Our model is implemented in the Gen.jl probabilistic programming language, and our code is available at https://github. com/thomashikaru/noisy_channel_model.
Thomas Hikaru Clark, Jacob Hoover Vigly, Edward Gibson, Roger Levy
EMNLP1
2025 Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models
Anna A. Ivanova, Aalok Sathe, Benjamin Lipkin, Unnathi U. Kumar, Setayesh Radkani, Thomas Hikaru Clark, Carina Kauf, Jennifer Hu 0001, R. T. Pramod, Gabriel Grand, Vivian C. Paulun, Maria Ryskina, Ekin Akyürek, Ethan Wilcox, Nafisa Rashid, Leshem Choshen, Roger Levy, Evelina Fedorenko, Josh Tenenbaum, Jacob Andreas
Trans. Assoc. Comput. Linguistics6
2024 Inferring errors and intended meanings with a generative model of language production in aphasia
Thomas Hikaru Clark, Edward Gibson, Roger Levy
CogSci1
2023 Context-sensitive features predict sentence memorability in the absence of memorable words
Thomas Hikaru Clark, Greta Tuckute, Bryan Medina, Evelina Fedorenko
CogSci1
2023 A Cross-Linguistic Pressure for Uniform Information Density in Word Order
abstract
Abstract While natural languages differ widely in both canonical word order and word order flexibility, their word orders still follow shared cross-linguistic statistical patterns, often attributed to functional pressures. In the effort to identify these pressures, prior work has compared real and counterfactual word orders. Yet one functional pressure has been overlooked in such investigations: The uniform information density (UID) hypothesis, which holds that information should be spread evenly throughout an utterance. Here, we ask whether a pressure for UID may have influenced word order patterns cross-linguistically. To this end, we use computational models to test whether real orders lead to greater information uniformity than counterfactual orders. In our empirical study of 10 typologically diverse languages, we find that: (i) among SVO languages, real word orders consistently have greater uniformity than reverse word orders, and (ii) only linguistically implausible counterfactual orders consistently exceed the uniformity of real orders. These findings are compatible with a pressure for information uniformity in the development and usage of natural languages.1
Thomas Hikaru Clark, Clara Meister, Tiago Pimentel, Michael Hahn 0001, Ryan Cotterell, Richard Futrell, Roger Levy
Trans. Assoc. Comput. Linguistics1
2022 Evidence for Availability Effects on Speaker Choice in the Russian Comparative Alternation
Thomas Hikaru Clark, Ethan Wilcox, Edward Gibson, Roger Levy
CogSci1
2021 Mixture-of-Partitions: Infusing Large Biomedical Knowledge Graphs into BERT
abstract
Infusing factual knowledge into pretrained models is fundamental for many knowledgeintensive tasks.In this paper, we propose Mixture-of-Partitions (MoP), an infusion approach that can handle a very large knowledge graph (KG) by partitioning it into smaller subgraphs and infusing their specific knowledge into various BERT models using lightweight adapters.To leverage the overall factual knowledge for a target task, these sub-graph adapters are further fine-tuned along with the underlying BERT through a mixture layer.We evaluate our MoP with three biomedical BERTs (SciBERT, BioBERT, PubmedBERT) on six downstream tasks (inc.NLI, QA, Classification), and the results show that our MoP consistently enhances the underlying BERTs in task performance, and achieves new SOTA performances on five evaluated datasets.1
Zaiqiao Meng, Fangyu Liu 0001, Thomas Hikaru Clark, Ehsan Shareghi, Nigel Collier
EMNLP (1)3