VLDB 2026 Research / reviewers in the wild / expert
Marten van Schijndel
dblp:127/0199
· DBLP profile ↗
21ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0002-9858-5881ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 8 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 75% Information extraction and text analysis · 13% Representation and self-supervised learning · 10% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
neural language model |
0.7 | 2 | 2019 | Quantity doesn't buy quality syntax with neural language models · EMNLP/IJCNLP (1) 2019 A Neural Model of Adaptation in Reading · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.6 | 1 | 2022 | Discourse Context Predictability Effects in Hindi Word Order · EMNLP 2022 |
Machine learning › Representation and self-supervised learning › representation analysis
representational similarity analysis |
0.5 | 1 | 2021 | All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › linguistic generalization
syntactic generalization |
0.4 | 1 | 2019 | Quantity doesn't buy quality syntax with neural language models · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Language models and text generation › neural language model
adaptive language models |
0.3 | 1 | 2018 | A Neural Model of Adaptation in Reading · EMNLP 2018 |
Natural language and speech › Language models and text generation › language acquisition
language acquisition modeling |
0.2 | 1 | 2014 | Bootstrapping into Filler-Gap: An Acquisition Story · ACL (1) 2014 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
cognitive modeling |
0.1 | 1 | 2018 | A Neural Model of Adaptation in Reading · EMNLP 2018 |
Natural language and speech › Language models and text generation › cognitive modeling of language
reading time prediction |
0.1 | 1 | 2018 | A Neural Model of Adaptation in Reading · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis › natural language semantics › computational semantics
semantic role assignment |
0.1 | 1 | 2014 | Bootstrapping into Filler-Gap: An Acquisition Story · ACL (1) 2014 |
Methods — techniques the papers use, named apart from their topics
syntactic priming · 0.6classifier-based prediction · 0.6LSTM · 0.6fine-tuning · 0.5euclidean distance · 0.5cosine similarity · 0.5recurrent neural network · 0.4adaptation mechanism · 0.3part-of-speech tagging · 0.2gaussian mixture model · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generating Representations In Space with GRIS
John R. Starr, Ashlyn Winship, Marten van Schijndel |
CogSci | 3 |
| 2025 | Experimentally extracting implicit instruments
Ashlyn Winship, Zander Lynch, Marten van Schijndel |
CogSci | 3 |
| 2025 | Disentangling language change: sparse autoencoders quantify the semantic evolution of indigeneity in FrenchabstractJacob A. Matthews, Laurent Dubreuil, Imane Terhmina, Yunci Sun, Matthew Wilkens, Marten Van Schijndel. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jacob Matthews, Laurent Dubreuil, Imane Terhmina, Yunci Sun, Matthew Wilkens, Marten van Schijndel |
NAACL (Long Papers) | 6 |
| 2024 | Does Dependency Locality Predict Non-canonical Word Order in Hindi?
Sidharth Ranjan, Marten van Schijndel |
CogSci | 2 |
| 2022 | Discourse Context Predictability Effects in Hindi Word OrderabstractWe test the hypothesis that discourse predictability influences Hindi syntactic choice.While prior work has shown that a number of factors (e.g., information status, dependency length, and syntactic surprisal) influence Hindi word order preferences, the role of discourse predictability is underexplored in the literature.Inspired by prior work on syntactic priming, we investigate how the words and syntactic structures in a sentence influence the word order of the following sentences.Specifically, we extract sentences from the Hindi-Urdu Treebank corpus (HUTB), permute the preverbal constituents of those sentences, and build a classifier to predict which sentences actually occurred in the corpus against artificially generated distractors.The classifier uses a number of discourse-based features and cognitive features to make its predictions, including dependency length, surprisal, and information status.We find that information status and LSTM-based discourse predictability influence word order choices, especially for non-canonical objectfronted orders.We conclude by situating our results within the broader syntactic priming literature. Sidharth Ranjan, Marten van Schijndel, Sumeet Agarwal, Rajakrishnan Rajkumar |
EMNLP | 2 |
| 2021 | Uncovering Constraint-Based Behavior in Neural Models via Targeted Fine-TuningabstractForrest Davis, Marten van Schijndel. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Forrest Davis, Marten van Schijndel |
ACL/IJCNLP (1) | 2 |
| 2021 | All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational QualityabstractSimilarity measures are a vital tool for understanding how language models represent and process language.Standard representational similarity measures such as cosine similarity and Euclidean distance have been successfully used in static word embedding models to understand how words cluster in semantic space.Recently, these measures have been applied to embeddings from contextualized models such as BERT and GPT-2.In this work, we call into question the informativity of such measures for contextualized language models.We find that a small number of rogue dimensions, often just 1-3, dominate these measures.Moreover, we find a striking mismatch between the dimensions that dominate similarity measures and those which are important to the behavior of the model.We show that simple postprocessing techniques such as standardization are able to correct for rogue dimensions and reveal underlying representational quality.We argue that accounting for rogue dimensions is essential for any similarity-based analysis of contextual language models. William Timkey, Marten van Schijndel |
EMNLP (1) | 2 |
| 2020 | Recurrent Neural Network Language Models Always Learn English-Like Relative Clause AttachmentabstractA standard approach to evaluating language models analyzes how models assign probabilities to valid versus invalid syntactic constructions (i.e. is a grammatical sentence more probable than an ungrammatical sentence).Our work uses ambiguous relative clause attachment to extend such evaluations to cases of multiple simultaneous valid interpretations, where stark grammaticality differences are absent.We compare model performance in English and Spanish to show that non-linguistic biases in RNN LMs advantageously overlap with syntactic structure in English but not Spanish.Thus, English models may appear to acquire human-like syntactic preferences, while models trained on Spanish fail to acquire comparable human-like preferences.We conclude by relating these results to broader concerns about the relationship between comprehension (i.e.typical language model use cases) and production (which generates the training data for language models), suggesting that necessary linguistic biases are not present in the training signal at all. Forrest Davis, Marten van Schijndel |
ACL | 2 |
| 2020 | Interaction with Context During Recurrent Neural Network Sentence Processing
Forrest Davis, Marten van Schijndel |
CogSci | 2 |
| 2020 | Filler-gaps that neural networks fail to generalizeabstractIt can be difficult to separate abstract linguistic knowledge in recurrent neural networks (RNNs) from surface heuristics.In this work, we probe for highly abstract syntactic constraints that have been claimed to govern the behavior of filler-gap dependencies across different surface constructions.For models to generalize abstract patterns in expected ways to unseen data, they must share representational features in predictable ways.We use cumulative priming to test for representational overlap between disparate filler-gap constructions in English and find evidence that the models learn a general representation for the existence of filler-gap dependencies.However, we find no evidence that the models learn any of the shared underlying grammatical constraints we tested.Our work raises questions about the degree to which RNN language models learn abstract linguistic representations. Debasmita Bhattacharya, Marten van Schijndel |
CoNLL | 2 |
| 2020 | Discourse structure interacts with reference but not syntax in neural language modelsabstractLanguage models (LMs) trained on large quantities of text have been claimed to acquire abstract linguistic representations. Our work tests the robustness of these abstractions by focusing on the ability of LMs to learn interactions between different linguistic representations. In particular, we utilized stimuli from psycholinguistic studies showing that humans can condition reference (i.e. coreference resolution) and syntactic processing on the same discourse structure (implicit causality). We compared both transformer and long short-term memory LMs to find that, contrary to humans, implicit causality only influences LM behavior for reference, not syntax, despite model representations that encode the necessary discourse information. Our results further suggest that LM behavior can contradict not only learned representations of discourse but also syntactic agreement, pointing to shortcomings of standard language modeling. Forrest Davis, Marten van Schijndel |
CoNLL | 2 |
| 2019 | Using Priming to Uncover the Organization of Syntactic Representations in Neural Language ModelsabstractNeural language models (LMs) perform well on tasks that require sensitivity to syntactic structure.Drawing on the syntactic priming paradigm from psycholinguistics, we propose a novel technique to analyze the representations that enable such success.By establishing a gradient similarity metric between structures, this technique allows us to reconstruct the organization of the LMs' syntactic representational space.We use this technique to demonstrate that LSTM LMs' representations of different types of sentences with relative clauses are organized hierarchically in a linguistically interpretable manner, suggesting that the LMs track abstract properties of the sentence. Grusha Prasad, Marten van Schijndel, Tal Linzen |
CoNLL | 2 |
| 2019 | Quantity doesn't buy quality syntax with neural language modelsabstractMarten van Schijndel, Aaron Mueller, Tal Linzen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Marten van Schijndel, Aaron Mueller, Tal Linzen |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Modeling garden path effects without explicit hierarchical syntax
Marten van Schijndel, Tal Linzen |
CogSci | 1 |
| 2018 | A Neural Model of Adaptation in ReadingabstractIt has been argued that humans rapidly adapt their lexical and syntactic expectations to match the statistics of the current linguistic context.We provide further support to this claim by showing that the addition of a simple adaptation mechanism to a neural language model improves our predictions of human reading times compared to a non-adaptive model.We analyze the performance of the model on controlled materials from psycholinguistic experiments and show that it adapts not only to lexical items but also to abstract syntactic structures. Marten van Schijndel, Tal Linzen |
EMNLP | 1 |
| 2017 | Approximations of Predictive Entropy Correlate with Reading Times
Marten van Schijndel, William Schuler |
CogSci | 1 |
| 2015 | Hierarchic syntax improves reading time predictionabstractPrevious work has debated whether humans make use of hierarchic syntax when processing language (Frank and Bod, 2011;Fossum and Levy, 2012).This paper uses an eye-tracking corpus to demonstrate that hierarchic syntax significantly improves reading time prediction over a strong n-gram baseline.This study shows that an interpolated 5-gram baseline can be made stronger by combining n-gram statistics over entire eye-tracking regions rather than simply using the last n-gram in each region, but basic hierarchic syntactic measures are still able to achieve significant improvements over this improved baseline. Marten van Schijndel, William Schuler |
HLT-NAACL | 1 |
| 2014 | Bootstrapping into Filler-Gap: An Acquisition StoryabstractAnalyses of filler-gap dependencies usu-ally involve complex syntactic rules or heuristics; however recent results suggest that filler-gap comprehension begins ear-lier than seemingly simpler constructions such as ditransitives or passives. Therefore, this work models filler-gap acquisition as a byproduct of learning word orderings (e.g. SVO vs OSV), which must be done at a very young age anyway in order to extract meaning from language. Specifically, this model, trained on part-of-speech tags, rep-resents the preferred locations of semantic roles relative to a verb as Gaussian mix-tures over real numbers. This approach learns role assignment in filler-gap constructions in a manner con-sistent with current developmental findings and is extremely robust to initialization variance. Additionally, this model is shown to be able to account for a characteristic er-ror made by learners during this period (A and B gorped interpreted as A gorped B). 1 Marten van Schijndel, Micha Elsner |
ACL (1) | 1 |
| 2014 | Frequency effects in the processing of unbounded dependencies
Marten van Schijndel, William Schuler, Peter W. Culicover |
CogSci | 1 |
| 2013 | An Analysis of Frequency- and Memory-Based Processing Costs
Marten van Schijndel, William Schuler |
HLT-NAACL | 1 |
| 2012 | Accurate Unbounded Dependency Recovery using Generalized Categorial Grammars
Luan Nguyen, Marten van Schijndel, William Schuler |
COLING | 2 |