VLDB 2026 Research / reviewers in the wild / expert
Thomas Hikaru Clark
dblp:301/9504
· DBLP profile ↗
9ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-1720-0450ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 48% Probabilistic and Bayesian machine learning · 45% Knowledge representation and reasoning · 7% | |
| Databases, data mining, and information retrieval
1 paper |
Knowledge graphs · 100% |
Topics — the 4 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning
probabilistic inference |
0.9 | 1 | 2025 | Resource-Rational Noisy-Channel Language Processing: Testing the Effect of Algorithmic Constraints on Inferences · EMNLP 2025 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
sequential monte carlo |
0.9 | 1 | 2025 | Resource-Rational Noisy-Channel Language Processing: Testing the Effect of Algorithmic Constraints on Inferences · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model › knowledge in language models
knowledge injection |
0.5 | 1 | 2021 | Mixture-of-Partitions: Infusing Large Biomedical Knowledge Graphs into BERT · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.5 | 1 | 2021 | Mixture-of-Partitions: Infusing Large Biomedical Knowledge Graphs into BERT · EMNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
mixture of experts · 1.0adapter · 1.0sequential monte carlo · 0.9rejuvenation · 0.9language model · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Readers make targeted regressions to plausible errors in reanalysis of "noisy-channel garden-path" sentencesabstractA key question in psycholinguistics is how inferences about the meaning of linguistic input unfold incrementally a comprehender's mind.In this work, we study reading dynamics for "noisy-channel garden-path" sentences, which temporarily appear well-formed but feature late-appearing violations of expectation that can be resolved not by inferring an alternative syntactic structure, but by inferring the presence of an error.We find evidence for targeted regressions -eye movements towards regions that are promising loci of possible errors in light of later-arriving information, showing patterns consistent with the posterior inferences of a model of noisy-channel processing with reanalysis.We discuss the implications of these findings for theories of noisy-channel language comprehension and information-theoretic explanations of reading dynamics. Thomas Hikaru Clark, Roger Levy, Edward Gibson |
CoNLL | 1 |
| 2025 | A Model of Approximate and Incremental Noisy-Channel Language Processing
Thomas Hikaru Clark, Jacob Hoover Vigly, Edward Gibson, Roger Levy |
CogSci | 1 |
| 2025 | Resource-Rational Noisy-Channel Language Processing: Testing the Effect of Algorithmic Constraints on InferencesabstractHuman language use is robust to errors: comprehenders can and do mentally correct utterances that are implausible or anomalous.How are humans able to solve these problems in real time, picking out alternatives from an unbounded space of options using limited cognitive resources?And can language models trained on next-word prediction for typical language be augmented to handle language anomalies in a human-like way?Using a language model as a prior and an error model to encode likelihoods, we use Sequential Monte Carlo with optional rejuvenation to perform incremental and approximate probabilistic inference over intended sentences and production errors.We demonstrate that the model captures previously established patterns in human sentence processing, and that a trade-off between human-like noisy-channel inferences and computational resources falls out of this model.From a psycholinguistic perspective, our results offer a candidate algorithmic model of rational inference in language processing.From an NLP perspective, our results showcase how to elicit human-like noisy-channel inference behavior from a relatively small LLM while controlling the amount of computation available during inference.Our model is implemented in the Gen.jl probabilistic programming language, and our code is available at https://github. com/thomashikaru/noisy_channel_model. Thomas Hikaru Clark, Jacob Hoover Vigly, Edward Gibson, Roger Levy |
EMNLP | 1 |
| 2025 | Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models
Anna A. Ivanova, Aalok Sathe, Benjamin Lipkin, Unnathi U. Kumar, Setayesh Radkani, Thomas Hikaru Clark, Carina Kauf, Jennifer Hu 0001, R. T. Pramod, Gabriel Grand, Vivian C. Paulun, Maria Ryskina, Ekin Akyürek, Ethan Wilcox, Nafisa Rashid, Leshem Choshen, Roger Levy, Evelina Fedorenko, Josh Tenenbaum, Jacob Andreas |
Trans. Assoc. Comput. Linguistics | 6 |
| 2024 | Inferring errors and intended meanings with a generative model of language production in aphasia
Thomas Hikaru Clark, Edward Gibson, Roger Levy |
CogSci | 1 |
| 2023 | Context-sensitive features predict sentence memorability in the absence of memorable words
Thomas Hikaru Clark, Greta Tuckute, Bryan Medina, Evelina Fedorenko |
CogSci | 1 |
| 2023 | A Cross-Linguistic Pressure for Uniform Information Density in Word OrderabstractAbstract While natural languages differ widely in both canonical word order and word order flexibility, their word orders still follow shared cross-linguistic statistical patterns, often attributed to functional pressures. In the effort to identify these pressures, prior work has compared real and counterfactual word orders. Yet one functional pressure has been overlooked in such investigations: The uniform information density (UID) hypothesis, which holds that information should be spread evenly throughout an utterance. Here, we ask whether a pressure for UID may have influenced word order patterns cross-linguistically. To this end, we use computational models to test whether real orders lead to greater information uniformity than counterfactual orders. In our empirical study of 10 typologically diverse languages, we find that: (i) among SVO languages, real word orders consistently have greater uniformity than reverse word orders, and (ii) only linguistically implausible counterfactual orders consistently exceed the uniformity of real orders. These findings are compatible with a pressure for information uniformity in the development and usage of natural languages.1 Thomas Hikaru Clark, Clara Meister, Tiago Pimentel, Michael Hahn 0001, Ryan Cotterell, Richard Futrell, Roger Levy |
Trans. Assoc. Comput. Linguistics | 1 |
| 2022 | Evidence for Availability Effects on Speaker Choice in the Russian Comparative Alternation
Thomas Hikaru Clark, Ethan Wilcox, Edward Gibson, Roger Levy |
CogSci | 1 |
| 2021 | Mixture-of-Partitions: Infusing Large Biomedical Knowledge Graphs into BERTabstractInfusing factual knowledge into pretrained models is fundamental for many knowledgeintensive tasks.In this paper, we propose Mixture-of-Partitions (MoP), an infusion approach that can handle a very large knowledge graph (KG) by partitioning it into smaller subgraphs and infusing their specific knowledge into various BERT models using lightweight adapters.To leverage the overall factual knowledge for a target task, these sub-graph adapters are further fine-tuned along with the underlying BERT through a mixture layer.We evaluate our MoP with three biomedical BERTs (SciBERT, BioBERT, PubmedBERT) on six downstream tasks (inc.NLI, QA, Classification), and the results show that our MoP consistently enhances the underlying BERTs in task performance, and achieves new SOTA performances on five evaluated datasets.1 Zaiqiao Meng, Fangyu Liu 0001, Thomas Hikaru Clark, Ehsan Shareghi, Nigel Collier |
EMNLP (1) | 3 |