Sebastian Schuster 0001

dblp:80/2843-1 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0002-8066-0142ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 7 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 RExBench: Can coding agents autonomously implement AI research extensions?
abstract
Nicholas Edwards, Yukyung Lee, Yujun Audrey Mao, Yulu Qin, Sebastian Schuster, Najoung Kim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Nicholas Edwards, Yukyung Lee, Yujun Audrey Mao, Yulu Qin, Sebastian Schuster 0001, Najoung Kim
ACL (1)5
2026 There Is No Spoon: Existential Presupposition in Large Language Models
Marie-Léontine Wörgötter, Shikai Lai, Sebastian Schuster 0001
LREC3
2025 Semantic-Pragmatic Adaptation to Variable Use of Temporal Expressions
Sebastian Schuster 0001
CogSci2
2024 SIGA: A Naturalistic NLI Dataset of English Scalar Implicatures with Gradable Adjectives
abstract
Many utterances convey meanings that go beyond the literal meaning of a sentence. One class of such meanings is scalar implicatures, a phenomenon by which a speaker conveys the negation of a more informative utterance by producing a less informative utterance. This paper introduces a Natural Language Inference (NLI) dataset designed to investigate the ability of language models to interpret utterances with scalar implicatures. Our dataset is comprised of text extracted from the C4 English text corpus and annotated with both crowd-sourced and expert annotations. We evaluate NLI models based on DeBERTa to investigate 1) whether NLI models can learn to predict pragmatic inferences involving gradable adjectives and 2) whether models generalize to utterances involving unseen adjectives. We find that fine-tuning NLI models on our dataset significantly improves their performance to derive scalar implicatures, both for in-domain and for out-of domain examples. At the same time, we find that the investigated models still perform considerably worse on examples with scalar implicatures than on other types of NLI examples, highlighting that pragmatic inferences still pose challenges for current models.
Rashid Nizamani, Sebastian Schuster 0001, Vera Demberg
LREC/COLING2
2024 SpreadNaLa: A Naturalistic Code Generation Evaluation Dataset of Spreadsheet Formulas
abstract
Automatic generation of code from natural language descriptions has emerged as one of the main use cases of large language models (LLMs). This has also led to a proliferation of datasets to track progress in the reliability of code generation models, including domains such as programming challenges and common data science tasks. However, existing datasets primarily target the use of code generation models to aid expert programmers in writing code. In this work, we consider a domain of code generation which is more frequently used by users without sophisticated programming skills: translating English descriptions to spreadsheet formulas that can be used to do everyday data processing tasks. We extract naturalistic instructions from StackOverflow posts and manually verify and standardize the corresponding spreadsheet formulas. We use this dataset to evaluate an off-the-shelf code generation model (GPT 3.5 text-davinci-003) as well as recently proposed pragmatic code generation procedures and find that Code Reviewer reranking (Zhang et al., 2022) performs best among the evaluated methods but still frequently generates formulas that differ from human-generated ones.
Sebastian Schuster 0001, Ayesha Ansar, Om Agarwal, Vera Demberg
LREC/COLING1
2024 Scope Ambiguities in Large Language Models
abstract
Abstract Sentences containing multiple semantic operators with overlapping scope often create ambiguities in interpretation, known as scope ambiguities. These ambiguities offer rich insights into the interaction between semantic structure and world knowledge in language processing. Despite this, there has been little research into how modern large language models treat them. In this paper, we investigate how different versions of certain autoregressive language models—GPT-2, GPT-3/3.5, Llama 2, and GPT-4—treat scope ambiguous sentences, and compare this with human judgments. We introduce novel datasets that contain a joint total of almost 1,000 unique scope-ambiguous sentences, containing interactions between a range of semantic operators, and annotated for human judgments. Using these datasets, we find evidence that several models (i) are sensitive to the meaning ambiguity in these sentences, in a way that patterns well with human judgments, and (ii) can successfully identify human-preferred readings at a high level of accuracy (over 90% in some cases).1
Gaurav Kamath, Sebastian Schuster 0001, Sowmya Vajjala, Siva Reddy
Trans. Assoc. Comput. Linguistics2
2023 Entity Tracking in Language Models
abstract
Keeping track of how states of entities change as a text or dialog unfolds is a key prerequisite to discourse understanding.Yet, there have been few systematic investigations into the ability of large language models (LLMs) to track discourse entities.In this work, we present a task probing to what extent a language model can infer the final state of an entity given an English description of the initial state and a series of state-changing operations.We use this task to first investigate whether Flan-T5, GPT-3 and GPT-3.5 can track the state of entities, and find that only GPT-3.5 models, which have been pretrained on large amounts of code, exhibit this ability.We then investigate whether smaller models pretrained primarily on text can learn to track entities, through finetuning T5 on several training/evaluation splits.While performance degrades for more complex splits, we find that even when evaluated on a different set of entities from training or longer operation sequences, a finetuned model can perform nontrivial entity tracking.Taken together, these results suggest that language models can learn to track entities but pretraining on text corpora alone does not make this capacity surface.
Najoung Kim, Sebastian Schuster 0001
ACL (1)2
2023 Working memory updating modulates adaptation to speaker-specific use of uncertainty expressions
Sebastian Schuster 0001, Alexandra Mayn, Vera Demberg
CogSci1
2023 Expectations over Unspoken Alternatives Predict Pragmatic Inferences
abstract
Abstract Scalar inferences (SI) are a signature example of how humans interpret language based on unspoken alternatives. While empirical studies have demonstrated that human SI rates are highly variable—both within instances of a single scale, and across different scales—there have been few proposals that quantitatively explain both cross- and within-scale variation. Furthermore, while it is generally assumed that SIs arise through reasoning about unspoken alternatives, it remains debated whether humans reason about alternatives as linguistic forms, or at the level of concepts. Here, we test a shared mechanism explaining SI rates within and across scales: context-driven expectations about the unspoken alternatives. Using neural language models to approximate human predictive distributions, we find that SI rates are captured by the expectedness of the strong scalemate as an alternative. Crucially, however, expectedness robustly predicts cross-scale variation only under a meaning-based view of alternatives. Our results suggest that pragmatic inferences arise from context-driven expectations over alternatives, and these expectations operate at the level of concepts.1
Jennifer Hu 0001, Roger Levy, Judith Degen, Sebastian Schuster 0001
Trans. Assoc. Comput. Linguistics4
2022 When a sentence does not introduce a discourse entity, Transformer-based models still sometimes refer to it
abstract
based models still sometimes refer to it
Sebastian Schuster 0001, Tal Linzen
NAACL-HLT1
2021 NOPE: A Corpus of Naturally-Occurring Presuppositions in English
abstract
Alicia Parrish, Sebastian Schuster, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R. Bowman, Tal Linzen. Proceedings of the 25th Conference on Computational Natural Language Learning. 2021.
Alicia Parrish, Sebastian Schuster 0001, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R. Bowman, Tal Linzen
CoNLL2
2020 Harnessing the linguistic signal to predict scalar inferences
abstract
Pragmatic inferences often subtly depend on the presence or absence of linguistic features.For example, the presence of a partitive construction (of the) increases the strength of a so-called scalar inference: listeners perceive the inference that Chris did not eat all of the cookies to be stronger after hearing "Chris ate some of the cookies" than after hearing the same utterance without a partitive, "Chris ate some cookies".In this work, we explore to what extent neural network sentence encoders can learn to predict the strength of scalar inferences.We first show that an LSTM-based sentence encoder trained on an English dataset of human inference strength ratings is able to predict ratings with high accuracy (r = 0.78).We then probe the model's behavior using manually constructed minimal sentence pairs and corpus data.We find that the model inferred previously established associations between linguistic features and inference strength, suggesting that the model learns to use linguistic features to predict pragmatic inferences.
Sebastian Schuster 0001, Judith Degen
ACL1
2020 Semantic Adaptation in Quantifier Meanings in Preschool Aged Children
Sophie Regan, Sebastian Schuster 0001, Judith Degen, Michael C. Frank
CogSci2
2020 Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection
abstract
Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework. The annotation consists in a linguistically motivated word segmentation; a morphological layer comprising lemmas, universal part-of-speech tags, and standardized morphological features; and a syntactic layer focusing on syntactic relations between predicates, arguments and modifiers. In this paper, we describe version 2 of the universal guidelines (UD v2), discuss the major changes from UD v1 to UD v2, and give an overview of the currently available treebanks for 90 languages.
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajic 0001, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster 0001, Francis M. Tyers, Daniel Zeman
LREC7
2019 Speaker-specific adaptation to variable use of uncertainty expressions
Sebastian Schuster 0001, Judith Degen
CogSci1
2018 Sentences with Gapping: Parsing and Reconstructing Elided Predicates
abstract
Sebastian Schuster, Joakim Nivre, Christopher D. Manning. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Sebastian Schuster 0001, Joakim Nivre, Christopher D. Manning
NAACL-HLT1
2016 Enhanced English Universal Dependencies: An Improved Representation for Natural Language Understanding Tasks
Sebastian Schuster 0001, Christopher D. Manning
LREC1
2014 Human Effort and Machine Learnability in Computer Aided Translation
abstract
Spence Green, Sida I. Wang, Jason Chuang, Jeffrey Heer, Sebastian Schuster, Christopher D. Manning. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
Spence Green, Sida I. Wang, Jason Chuang, Jeffrey Heer, Sebastian Schuster 0001, Christopher D. Manning
EMNLP5
2014 Speaker-independent detection of child-directed speech
abstract
Identifying the distinct register that adults use when speaking to children is an important task for child development research. We present a fully automatic, speaker-independent system that detects child-directed speech. The two-stage system uses diarization-style voice activation techniques to extract speech segments followed by a supervised ν-SVM classifier trained on 1582 prosodic and log Mel energy features. The system significantly improves the state of the art, detecting child-directed speech with F1 of .66 (exact boundary) and .83 (within 1 second). A feature analysis confirms the importance of F0 features (especially 3rd quartile and range) as well as new features like the variance, kurtosis, and min of log Mel energy within a frequency band.
Sebastian Schuster 0001, Stephanie Pancoast, Milind Ganjoo, Michael C. Frank, Daniel Jurafsky
SLT1