Tatsuya Aoyama

dblp:324/3845 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021
YearPublicationVenuePosition
2025 Language Models Grow Less Humanlike beyond Phase Transition
abstract
LMs' alignment with human reading behavior (i.e.psychometric predictive power; PPP) is known to improve during pretraining up to a tipping point, beyond which it either plateaus or degrades.Various factors, such as word frequency, recency bias in attention, and context size, have been theorized to affect PPP, yet there is no current account that explains why such a tipping point exists, and how it interacts with LMs' pretraining dynamics more generally.We hypothesize that the underlying factor is a pretraining phase transition, characterized by the rapid emergence of specialized attention heads.We conduct a series of correlational and causal experiments to show that such a phase transition is responsible for the tipping point in PPP.We then show that, rather than producing attention patterns that contribute to the degradation in PPP, phase transitions alter the subsequent learning dynamics of the model, such that further training keeps damaging PPP.
Tatsuya Aoyama, Ethan Wilcox
ACL (1)1
2025 Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs
abstract
Do language models (LMs) offer insights into human language learning?A common argument against this idea is that because their architecture and training paradigm are so vastly different from humans, LMs can learn arbitrary inputs as easily as natural languages.We test this claim by training LMs to model impossible and typologically unattested languages.Unlike previous work, which has focused exclusively on English, we conduct experiments on 12 languages from 4 language families with two newly constructed parallel corpora.Our results show that while GPT-2 small can largely distinguish attested languages from their impossible counterparts, it does not achieve perfect separation between all the attested languages and all the impossible ones.We further test whether GPT-2 small distinguishes typologically attested from unattested languages with different NP orders by manipulating word order based on Greenberg's Universal 20.We find that the model's perplexity scores do not distinguish attested vs. unattested word orders, while its performance on the generalization test does.These findings suggest that LMs exhibit some human-like inductive biases, though these biases are weaker than those found in human learners.
Xiulin Yang, Tatsuya Aoyama, Yuekun Yao, Ethan Wilcox
ACL (1)2
2025 Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaning
abstract
Humans have a remarkable ability to acquire and understand grammatical phenomena that are seen rarely, if ever, during childhood.Recent evidence suggests that language models with human-scale pretraining data may possess a similar ability by generalizing from frequent to rare constructions.However, it remains an open question how widespread this generalization ability is, and to what extent this knowledge extends to meanings of rare constructions, as opposed to just their forms.We fill this gap by testing human-scale transformer language models on their knowledge of both the form and meaning of the (rare and quirky) English LET-ALONE construction.To evaluate our LMs we construct a bespoke synthetic benchmark that targets syntactic and semantic properties of the construction.We find that human-scale LMs are sensitive to form, even when related constructions are filtered from the dataset.However, human-scale LMs do not make correct generalizations about LET-ALONE's meaning.These results point to an asymmetry in the current architectures' sample efficiency between language form and meaning, something which is not present in human language learners.1
Wesley Scivetti, Tatsuya Aoyama, Ethan Wilcox, Nathan Schneider 0001
EMNLP2
2025 Distributionally Robust Active Learning for Gaussian Process Regression
abstract
Gaussian process regression (GPR) or kernel ridge regression is a widely used and powerful tool for nonlinear prediction. Therefore, active learning (AL) for GPR, which actively collects data labels to achieve an accurate prediction with fewer data labels, is an important problem. However, existing AL methods do not theoretically guarantee prediction accuracy for target distribution. Furthermore, as discussed in the distributionally robust learning literature, specifying the target distribution is often difficult. Thus, this paper proposes two AL methods that effectively reduce the worst-case expected error for GPR, which is the worst-case expectation in target distribution candidates. We show an upper bound of the worst-case expected squared error, which suggests that the error will be arbitrarily small by a finite number of data labels under mild conditions. Finally, we demonstrate the effectiveness of the proposed methods through synthetic and real-world datasets.
Shion Takeno, Yoshito Okura, Yu Inatsu, Tatsuya Aoyama, Tomonari Tanaka, Satoshi Akahane, Hiroyuki Hanada, Noriaki Hashimoto, Taro Murayama, Hanju Lee, Shinya Kojima, Ichiro Takeuchi
ICML4
2025 eRST: A Signaled Graph Theory of Discourse Relations and Organization
abstract
Abstract In this article we present Enhanced Rhetorical Structure Theory (eRST), a new theoretical framework for computational discourse analysis, based on an expansion of Rhetorical Structure Theory (RST). The framework encompasses discourse relation graphs with tree-breaking, non-projective and concurrent relations, as well as implicit and explicit signals which give explainable rationales to our analyses. We survey shortcomings of RST and other existing frameworks, such as Segmented Discourse Representation Theory, the Penn Discourse Treebank, and Discourse Dependencies, and address these using constructs in the proposed theory. We provide annotation, search, and visualization tools for data, and present and evaluate a freely available corpus of English annotated according to our framework, encompassing 12 spoken and written genres with over 200K tokens. Finally, we discuss automatic parsing, evaluation metrics, and applications for data in our framework.
Amir Zeldes, Tatsuya Aoyama, Yang Janet Liu, Siyao Peng, Debopam Das, Luke Gessler
Comput. Linguistics2
2024 J-SNACS: Adposition and Case Supersenses for Japanese Joshi
abstract
Many languages use adpositions (prepositions or postpositions) to mark a variety of semantic relations, with different languages exhibiting both commonalities and idiosyncrasies in the relations grouped under the same lexeme. We present the first Japanese extension of the SNACS framework (Schneider et al., 2018), which has served as the basis for annotating adpositions in corpora from several languages. After establishing which of the set of particles (joshi) in Japanese qualify as case markers and adpositions as defined in SNACS, we annotate 10 chapters (≈10k tokens) of the Japanese translation of Le Petit Prince (The Little Prince), achieving high inter-annotator agreement. We find that, while a majority of the particles and their uses are captured by the existing and extended SNACS annotation guidelines from the previous work, some unique cases were observed. We also conduct experiments investigating the cross-lingual similarity of adposition and case marker supersenses, showing that the language-agnostic SNACS framework captures similarities not clearly observed in multilingual embedding space.
Tatsuya Aoyama, Chihiro Taguchi, Nathan Schneider 0001
LREC/COLING1
2024 DISRPT: A Multilingual, Multi-domain, Cross-framework Benchmark for Discourse Processing
abstract
This paper presents DISRPT, a multilingual, multi-domain, and cross-framework benchmark dataset for discourse processing, covering the tasks of discourse unit segmentation, connective identification, and relation classification. DISRPT includes 13 languages, with data from 24 corpora covering about 4 millions tokens and around 250,000 discourse relation instances from 4 discourse frameworks: RST, SDRT, PDTB, and Discourse Dependencies. We present an overview of the data, its development across three NLP shared tasks on discourse processing carried out in the past five years, and the latest modifications and added extensions. We also carry out an evaluation of state-of-the-art multilingual systems trained on the data for each task, showing plateau performance on segmentation, but important room for improvement for connective identification and relation classification. The DISRPT benchmark employs a unified format that we make available on GitHub and HuggingFace in order to encourage future work on discourse processing across languages, domains, and frameworks.
Chloé Braud, Amir Zeldes, Laura Rivière, Yang Janet Liu, Philippe Muller, Damien Sileo, Tatsuya Aoyama
LREC/COLING7
2024 Modeling Nonnative Sentence Processing with L2 Language Models
abstract
We study LMs pretrained sequentially on two languages ("L2LMs") for modeling nonnative sentence processing.In particular, we pretrain GPT2 on 6 different first languages (L1s), followed by English as the second language (L2).We examine the effect of the choice of pretraining L1 on the model's ability to predict human reading times, evaluating on English readers from a range of L1 backgrounds.Experimental results show that, while all of the LMs' word surprisals improve prediction of L2 reading times, especially for human L1s distant from English, there is no reliable effect of the choice of L2LM's L1.We also evaluate the learning trajectory of a monolingual English LM: for predicting L2 as opposed to L1 reading, it peaks much earlier and immediately falls off, possibly mirroring the difference in proficiency between the native and nonnative populations.Lastly, we provide examples of L2LMs' surprisals, which could potentially generate hypotheses about human L2 reading.
Tatsuya Aoyama, Nathan Schneider 0001
EMNLP1
2024 GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and Domains
abstract
Yang Janet Liu, Tatsuya Aoyama, Wesley Scivetti, Yilun Zhu, Shabnam Behzad, Lauren Elizabeth Levine, Jessica Lin, Devika Tiwari, Amir Zeldes. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yang Janet Liu, Tatsuya Aoyama, Wesley Scivetti, Yilun Zhu 0001, Shabnam Behzad, Lauren Levine, Jessica Lin 0004, Devika Tiwari, Amir Zeldes
EMNLP2
2023 What's Hard in English RST Parsing? Predictive Models for Error Analysis
abstract
Despite recent advances in Natural Language Processing (NLP), hierarchical discourse parsing in the framework of Rhetorical Structure Theory remains challenging, and our understanding of the reasons for this are as yet limited.In this paper, we examine and model some of the factors associated with parsing difficulties in previous work: the existence of implicit discourse relations, challenges in identifying long-distance relations, out-of-vocabulary items, and more.In order to assess the relative importance of these variables, we also release two annotated English test-sets with explicit correct and distracting discourse markers associated with gold standard RST relations.Our results show that as in shallow discourse parsing, the explicit/implicit distinction plays a role, but that long-distance dependencies are the main challenge, while lack of lexical overlap is less of a problem, at least for in-domain parsing.Our final model is able to predict where errors will occur with an accuracy of 76.3% for the bottom-up parser and 76.6% for the top-down parser.
Yang Janet Liu, Tatsuya Aoyama, Amir Zeldes
SIGDIAL2