Tatsuki Kuribayashi

dblp:228/5787 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0001-7762-5576ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 8 first-author · 22 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages
abstract
Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina, Ashwath Rao B, Parameswari Krishnamurthy, Muhammad Cendekia Airlangga, Rifo Ahmad Genadi, Nguyen Phan Gia Bao, Amir Hossein Yari, Hawau Olamide Toyin, Nurdaulet Mukhituly, Mena Attia, Besher Hassan, Ahmad Fathan Hidayatullah, Tatsuki Kuribayashi, Haonan Li, Suma Bhat, Fajri Koto. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina, Ashwath Rao, Parameswari Krishnamurthy, Muhammad Cendekia Airlangga, Rifo Ahmad Genadi, Nguyen Phan Gia Bao, Amir Hossein Yari, Hawau Olamide Toyin, Nurdaulet Mukhituly, Mena Attia, Besher Hassan, Ahmad Fathan Hidayatullah, Tatsuki Kuribayashi, Haonan Li 0002, Suma Bhat, Fajri Koto
ACL (1)16
2026 Dual Alignment Between Language Model Layers and Human Sentence Processing
abstract
A recent study (Kuribayashi et al., 2025) has shown that human sentence processing behavior, typically measured on syntactically unchallenging constructions, can be effectively modeled using surprisal from early layers of large language models (LLMs).This raises the question of whether such advantages of internal layers extend to more syntactically challenging constructions, where surprisal has been reported to underestimate human cognitive effort.In this paper, we begin by exploring internal layers that better estimate human cognitive effort observed in syntactic ambiguity processing in English.Our experiments show that, in contrast to naturalistic reading, later layers better estimate such a cognitive effort, but still underestimate the human data.This dual alignment sheds light on different modes of sentence processing in humans and LMs: naturalistic reading employs a somewhat weak prediction akin to earlier layers of LMs, while syntactically challenging processing requires more fully-contextualized representations, better modeled by later layers of LMs.Motivated by these findings, we also explore several probability-update measures using shallow and deep layers of LMs, showing a complementary advantage to single-layer's surprisal in reading time modeling. https://github.com/kuribayashi4/ internal_surprisal_targeted_assessmentPhenomena Example MVRR D + : The girl fed the lamb remained relatively calm before the sunset in silence.D -: The girl who was fed the lamb remained relatively calm before the sunset in silence.NPS D + : The girl found the lamb remained relatively calm near the wooden fence.D -: The girl found that the lamb remained relatively calm near the wooden fence.NPZ D + : When the girl attacked the lamb remained relatively calm despite the sudden noise.D -: When the girl attacked, the lamb remained relatively calm despite the sudden noise.RC D + : The bus driver that the kids followed waited patiently at dawn.D -: The bus driver that followed the kids waited patiently at dawn.Attachment D + : Janet charmed the executive of the assistants who decides almost everything during long weekly meetings.
Tatsuki Kuribayashi, Alex Warstadt, Yohei Oseki, Ethan Wilcox
ACL (1)1
2026 On the Effect of Hyperparameters in Language Modeling for Computational Linguistics
abstract
Training language models and examining their linguistic behaviors have been a common protocol in computational linguistics for studying linguistic phenomena and modeling human language processing.However, work in this area is often limited to proof-of-concept demonstrations with arbitrary model configurations, without considering hyperparameter sensitivity, an important source of variation in model performance.In this work, we replicate three prior studies (Chang and Bergen, 2022; Hu et al., 2020b;Kuribayashi et al., 2024) with hyperparameters varied within a practical range, and show that modest hyperparameter changes can alter some qualitative conclusions about models' linguistic abilities and even reverse the ranking of model performance.Our results highlight the risk that prior work may have reflected optimization artifacts rather than the genuine inductive biases of model classes, and that hyperparameter sensitivity should receive more attention as a factor that can meaningfully influence model behavior.We suggest future work to report the variation of performance across the configuration space to enhance the reliability and generalizability of conclusions.
Ruoxi Ning, Yongpeng Zhu, Qingcheng Zeng, Tatsuki Kuribayashi, Freda Shi
ACL (1)4
2026 An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal
abstract
Surprisal theory hypothesizes that the difficulty of human sentence processing increases linearly with surprisal, the negative logprobability of a word given its context.Computational psycholinguistics has tested this hypothesis using language models (LMs) as proxies for human prediction.While surprisal derived from recent neural LMs generally captures human processing difficulty on naturalistic corpora that predominantly consist of simple sentences, it severely underestimates processing difficulty on sentences that require syntactic disambiguation (garden-path effects).This leads to the claim that the processing difficulty of such sentences cannot be reduced to surprisal, although it remains possible that neural LMs simply differ from humans in nextword prediction.In this paper, we investigate whether it is truly impossible to construct a neural LM that can explain garden-path effects via surprisal.Specifically, instead of evaluating offthe-shelf neural LMs, we fine-tune these LMs on garden-path sentences so as to better align surprisal-based reading-time estimates with actual human reading times.Our results show that fine-tuned LMs do not overfit and successfully capture human reading slowdowns on held-out garden-path items; they even improve predictive power for human reading times on naturalistic corpora and preserve their general LM capabilities.These results provide an existence proof for a neural LM that can explain both garden-path effects and naturalistic reading times via surprisal, but also raise a theoretical question: what kind of evidence can truly falsify surprisal theory?
Ryo Yoshida, Shinnosuke Isono, Taiga Someya, Yohei Oseki, Tatsuki Kuribayashi
ACL (1)5
2026 Can Language Models Learn Typologically Implausible Languages?
abstract
Abstract Grammatical features across human languages exhibit intriguing correlations, often attributed to learning biases in humans. Language models (LMs) provide a scalable and naturalistic framework for studying artificial language learning—one not available in human research. We investigate how learnability varies across typologically plausible and implausible languages that closely follow the word order universals identified by linguistic typologists. Our study trains LMs on highly naturalistic counterfactual versions of English (head-initial) and Japanese (head-final). Compared to prior work, our datasets more precisely target the boundary between typological plausibility and implausibility. Our experiments show that LMs learn subtly implausible languages more slowly, though they eventually reach similar performance on some metrics regardless of typological plausibility. These findings suggest that LMs exhibit typologically aligned learning preferences and that certain typological patterns may emerge from general learning biases. https://github.com/sally-xu-42/Typological_Universals.
Tianyang Xu 0002, Tatsuki Kuribayashi, Yohei Oseki, Ryan Cotterell, Alex Warstadt
Trans. Assoc. Comput. Linguistics2
2025 Can LLMs Simulate L2-English Dialogue? An Information-Theoretic Analysis of L1-Dependent Biases
abstract
Rena Gao, Xuetong Wu, Tatsuki Kuribayashi, Mingrui Ye, Siya Qi, Carsten Roever, Yuanxing Liu, Zheng Yuan, Jey Han Lau. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Rena Gao, Xuetong Wu, Tatsuki Kuribayashi, Mingrui Ye, Siya Qi, Carsten Roever, Yuanxing Liu 0001, Jey Han Lau
ACL (1)3
2025 Does Vision Accelerate Hierarchical Generalization in Neural Language Learners?
abstract
Neural language models (LMs) are arguably less data-efficient than humans from a language acquisition perspective. One fundamental question is why this human–LM gap arises. This study explores the advantage of grounded language acquisition, specifically the impact of visual information — which humans can usually rely on but LMs largely do not have access to during language acquisition — on syntactic generalization in LMs. Our experiments, following the poverty of stimulus paradigm under two scenarios (using artificial vs. naturalistic images), demonstrate that if the alignments between the linguistic and visual components are clear in the input, access to vision data does help with the syntactic generalization of LMs, but if not, visual input does not help. This highlights the need for additional biases or signals, such as mutual gaze, to enhance cross-modal alignment and enable efficient syntactic generalization in multimodal LMs.
Tatsuki Kuribayashi, Timothy Baldwin
COLING1
2025 Which Word Orders Facilitate Length Generalization in LMs? An Investigation with GCG-Based Artificial Languages
abstract
Whether language models (LMs) have inductive biases that favor typologically frequent grammatical properties over rare, implausible ones has been investigated, typically using artificial languages (ALs) (White and Cotterell, 2021;Kuribayashi et al., 2024).In this paper, we extend these works from two perspectives.First, we extend their context-free AL formalization by adopting Generalized Categorial Grammar (GCG) (Wood, 2014), which allows ALs to cover attested but previously overlooked constructions, such as unbounded dependency and mildly context-sensitive structures.Second, our evaluation focuses more on the generalization ability of LMs to process unseen longer test sentences.Thus, our ALs better capture features of natural languages and our experimental paradigm leads to clearer conclusionstypologically plausible word orders tend to be easier for LMs to productively generalize.
Nadine El-Naggar, Tatsuki Kuribayashi, Ted Briscoe
EMNLP2
2025 Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders
abstract
Recent interpretability work on large language models (LLMs) has been increasingly dominated by a feature-discovery approach with the help of proxy modules. Then, the quality of features learned by, e.g., sparse auto-encoders (SAEs), is evaluated. This paradigm naturally raises a critical question: do such learned features have better properties than those already represented within the original model parameters, and unfortunately, only a few studies have made such comparisons systematically so far. In this work, we revisit the interpretability of feature vectors stored in feed-forward (FF) layers, given the perspective of FF as key-value memories, with modern interpretability benchmarks. Our extensive evaluation revealed that SAE and FFs exhibits a similar range of interpretability, although SAEs displayed an observable but minimal improvement in some aspects. Furthermore, in certain aspects, surprisingly, even vanilla FFs yielded better interpretability than the SAEs, and features discovered in SAEs and FFs diverged. These bring questions about the advantage of SAEs from both perspectives of feature quality and faithfulness, compared to directly interpreting FF feature vectors, and FF key-value parameters serve as a strong baseline in modern interpretability research.
Mengyu Ye, Jun Suzuki 0001, Tatsuro Inaba, Tatsuki Kuribayashi
NeurIPS4
2025 Large Language Models Are Human-Like Internally
abstract
Abstract Recent cognitive modeling studies have reported that larger language models (LMs) exhibit a poorer fit to human reading behavior (Oh and Schuler, 2023b; Shain et al., 2024; Kuribayashi et al., 2024), leading to claims of their cognitive implausibility. In this paper, we revisit this argument through the lens of mechanistic interpretability and argue that prior conclusions were skewed by an exclusive focus on the final layers of LMs. Our analysis reveals that next-word probabilities derived from internal layers of larger LMs align with human sentence processing data as well as, or better than, those from smaller LMs. This alignment holds consistently across behavioral (self-paced reading times, gaze durations, MAZE task processing times) and neurophysiological (N400 brain potentials) measures, challenging earlier mixed results and suggesting that the cognitive plausibility of larger LMs has been underestimated. Furthermore, we first identify an intriguing relationship between LM layers and human measures: Earlier layers correspond more closely with fast gaze durations, while later layers better align with relatively slower signals such as N400 potentials and MAZE processing times. Our work opens new avenues for interdisciplinary research at the intersection of mechanistic interpretability and cognitive modeling.1
Tatsuki Kuribayashi, Yohei Oseki, Souhaib Ben Taieb, Kentaro Inui, Timothy Baldwin
Trans. Assoc. Comput. Linguistics1
2024 Emergent Word Order Universals from Cognitively-Motivated Language Models
abstract
Tatsuki Kuribayashi, Ryo Ueda, Ryo Yoshida, Yohei Oseki, Ted Briscoe, Timothy Baldwin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Tatsuki Kuribayashi, Ryo Ueda, Ryo Yoshida, Yohei Oseki, Ted Briscoe, Timothy Baldwin
ACL (1)1
2024 To Drop or Not to Drop? Predicting Argument Ellipsis Judgments: A Case Study in Japanese
abstract
Speakers sometimes omit certain arguments of a predicate in a sentence; such omission is especially frequent in pro-drop languages. This study addresses a question about ellipsis—what can explain the native speakers’ ellipsis decisions?—motivated by the interest in human discourse processing and writing assistance for this choice. To this end, we first collect large-scale human annotations of whether and why a particular argument should be omitted across over 2,000 data points in the balanced corpus of Japanese, a prototypical pro-drop language. The data indicate that native speakers overall share common criteria for such judgments and further clarify their quantitative characteristics, e.g., the distribution of related linguistic factors in the balanced corpus. Furthermore, the performance of the language model–based argument ellipsis judgment model is examined, and the gap between the systems’ prediction and human judgments in specific linguistic aspects is revealed. We hope our fundamental resource encourages further studies on natural human ellipsis judgment.
Yukiko Ishizuki, Tatsuki Kuribayashi, Yuichiroh Matsubayashi, Ryohei Sasano, Kentaro Inui
LREC/COLING2
2024 First Heuristic Then Rational: Dynamic Use of Heuristics in Language Model Reasoning
abstract
Yoichi Aoki, Keito Kudo, Tatsuki Kuribayashi, Shusaku Sone, Masaya Taniguchi, Keisuke Sakaguchi, Kentaro Inui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yoichi Aoki, Keito Kudo, Tatsuki Kuribayashi, Shusaku Sone, Masaya Taniguchi, Keisuke Sakaguchi, Kentaro Inui
EMNLP3
2024 Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention Maps
abstract
Transformers are ubiquitous in wide tasks. Interpreting their internals is a pivotal goal. Nevertheless, their particular components, feed-forward (FF) blocks, have typically been less analyzed despite their substantial parameter amounts. We analyze the input contextualization effects of FF blocks by rendering them in the attention maps as a human-friendly visualization scheme. Our experiments with both masked- and causal-language models reveal that FF networks modify the input contextualization to emphasize specific types of linguistic compositions. In addition, FF and its surrounding components tend to cancel out each other's effects, suggesting potential redundancy in the processing of the Transformer layer.
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro Inui
ICLR2
2024 CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
abstract
Visual Question Answering~(VQA) is an important task in multimodal AI, which requires models to understand and reason on knowledge present in visual and textual data. However, most of the current VQA datasets and models are primarily focused on English and a few major world languages, with images that are Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, some datasets extend the text to other languages, either via translation or some other approaches, but usually keep the same images, resulting in narrow cultural representation. To address these limitations, we create CVQA, a new Culturally-diverse Multilingual Visual Question Answering benchmark dataset, designed to cover a rich set of languages and regions, where we engage native speakers and cultural experts in the data collection process. CVQA includes culturally-driven images and questions from across 28 countries in four continents, covering 26 languages with 11 scripts, providing a total of 9k questions. We benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and we show that the dataset is challenging for the current state-of-the-art models. This benchmark will serve as a probing evaluation suite for assessing the cultural bias of multimodal models and hopefully encourage more research efforts towards increasing cultural awareness and linguistic diversity in this field.
Chenyang Lyu, Haryo Akbarianto Wibowo, Santiago Góngora, Aishik Mandal, Sukannya Purkayastha, Jesús-Germán Ortiz-Barajas, Emilio Villa-Cueva, Jinheon Baek, Soyeong Jeong, Injy Hamed, Zheng Wei Lim, Paula Mónica Silva, Jocelyn Dunstan, Mélanie Jouitteau, David Le Meur, Joan Nwatu, Ganzorig Batnasan, Munkh-Erdene Otgonbold, Munkhjargal Gochoo, Guido Ivetta, Luciana Benotti, Laura Alonso Alemany, Hernán Maina, Jiahui Geng, Tiago Timponi Torrent, Frederico Belcavello, Marcelo Viridiano, Jan Christian Blaise Cruz, Dan John Velasco, Oana Ignat, Zara Burzo, Chenxi Whitehouse, Artem Abzaliev, Teresa Clifford, Grainne Caulfield, Teresa Lynn, Christian Salamea Palacios, Vladimir Araujo, Yova Kementchedjhieva, Mihail Mihaylov, Israel Abebe Azime, Henok Biadglign Ademtew, Bontu Fufa Balcha, Naome A. Etori, David Ifeoluwa Adelani, Rada Mihalcea, Atnafu Lambebo Tonja, Maria Camila Buitrago Cabrera, Gisela Vallejo, Holy Lovenia, Ruochen Zhang 0001, Marcos Estecha-Garitagoitia, Mario Rodríguez-Cantelar, Toqeer Ehsan, Rendi Chevi, Muhammad Farid Adilazuarda, Ryandito Diandaru, Samuel Cahyawijaya, Fajri Koto, Tatsuki Kuribayashi, Haiyue Song, Aditya Khandavally, Thanmay Jayakumar, Raj Dabre, Mohamed Fazli Mohamed Imam, Kumaranage Ravindu Yasas Nagasinghe, Alina Dragonetti, Luis Fernando D'Haro, Olivier Niyomugisha, Jay Gala, Pranjal A. Chitale, Fauzan Farooqui, Thamar Solorio, Alham Fikri Aji
NeurIPS62
2023 Do Deep Neural Networks Capture Compositionality in Arithmetic Reasoning?
abstract
Keito Kudo, Yoichi Aoki, Tatsuki Kuribayashi, Ana Brassard, Masashi Yoshikawa, Keisuke Sakaguchi, Kentaro Inui. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Keito Kudo, Yoichi Aoki, Tatsuki Kuribayashi, Ana Brassard, Masashi Yoshikawa, Keisuke Sakaguchi, Kentaro Inui
EACL3
2023 Assessing Step-by-Step Reasoning against Lexical Negation: A Case Study on Syllogism
abstract
Large language models (LLMs) take advantage of step-by-step reasoning instructions, e.g., chain-of-thought (CoT) prompting.Building on this, their ability to perform CoT-style reasoning robustly is of interest from a probing perspective.In this study, we inspect the stepby-step reasoning ability of LLMs with a focus on negation, which is a core linguistic phenomenon that is difficult to process.In particular, we introduce several controlled settings (e.g., reasoning on fictional entities) to evaluate the logical reasoning abilities of the models.We observed that dozens of modern LLMs were not robust against lexical negation (e.g., plausi-ble→implausible) when performing CoT-style reasoning, and the results highlight unique limitations in each LLM family.https://github.com/muyo8692/ stepbystep-reasoning-vs-negation Setting Few-shot exemplars Target example If fails at this setting BASE Is a sentence "A does B" plausible?A is a C player.B happens in C/X.So the answer is yes/no.Is a sentence "D does E" plausible?D is a F player.E happens in F/Y.So the answer is __ CoT-style reasoning fails. FIC Is a sentence "A does B" plausible?A is a C player.B happens in C/X.So the answer is yes/no.Is a sentence "α does β" plausible?α is a γ player.β happens in γ/χ.So the answer is __ Reasoning cannot be abstracted to fictional texts. FICNEG Is a sentence "A does B" implausible?A is a C player.B happens in C/X.So the answer is yes/no.Is a sentence "α does β" implausible?α is a γ player.β happens in γ/χ.So the answer is __ Abstract CoT-style reasoning is only achieved on the affirmative domain. FICNEG-O Is a sentence "A does B" plausible?A is a C player.B happens in C/X.So the answer is yes/no.Is a sentence "α does β" implausible?α is a γ player.β happens in γ/χ.So the answer is __ Model cannot handle domain shift in terms of negation.Table 1: General task format in each setting.Few-shot exemplars are first shown to a model, and then the model answers to the target example given its question and reasoning chain.Symbols, e.g., A and α, are replaced with certain real or fictional entities in the actual input.The REAL setting indicates that the entity choices reflect the factual reality, and FIC.indicates that the entity choices do not reflect factual reality, e.g., Is "Judy Tate was safe at first." plausible?Judy Tate is a turboglide player.Getting out at first happens in turboglide.So the answer is yes.Refer to Appendix A for the exact input. Format: SyllogismWe evaluated the LLMs' ability to judge the validity of particular types of syllogisms.Here, we utilized three settings to ensure the robustness of the results (Section 3); however, we consider the following SPORTS TASK (SP) format as an example to explain the settings.The base format of the syllogism is as follows:Premise1: PERSON is a SPORT player.Premise2: ACTION happens in the SPORT. Conclusion: PERSON does ACTION.Here, the transitivity of reasoning (A>B, B>C, then A>C) is targeted. C.3 Answer Distribution
Mengyu Ye, Tatsuki Kuribayashi, Jun Suzuki 0001, Goro Kobayashi, Hiroaki Funayama
EMNLP2
2023 Use of an AI-powered Rewriting Support Software in Context with Other Tools: A Study of Non-Native English Speakers
abstract
Academic writing in English can be challenging for non-native English speakers (NNESs). AI-powered rewriting tools can potentially improve NNESs’ writing outcomes at a low cost. However, whether and how NNESs make valid assessments of the revisions provided by these algorithmic recommendations remains unclear. We report a study where NNESs leverage an AI-powered rewriting tool, Langsmith, to polish their drafted academic essays. We examined the participants’ interactions with the tool via user studies and interviews. Our data reveal that most participants used Langsmith in combination with other tools, such as machine translation (MT), and those who used MT had different ways of understanding and evaluating Langsmith’s suggestions than those who did not. Based on these findings, we assert that NNESs’ quality assessment in AI-powered rewriting tools is influenced by the simultaneous use of multiple tools, offering valuable insights into the design of future rewriting tools for NNESs.
Takumi Ito, Naomi Yamashita, Tatsuki Kuribayashi, Masatoshi Hidaka, Jun Suzuki 0001, Ge Gao 0001, Jack Jamieson, Kentaro Inui
UIST3
2022 Topicalization in Language Models: A Case Study on Japanese
abstract
Humans use different wordings depending on the context to facilitate efficient communication. For example, instead of completely new information, information related to the preceding context is typically placed at the sentence-initial position. In this study, we analyze whether neural language models (LMs) can capture such discourse-level preferences in text generation. Specifically, we focus on a particular aspect of discourse, namely the topic-comment structure. To analyze the linguistic knowledge of LMs separately, we chose the Japanese language, a topic-prominent language, for designing probing tasks, and we created human topicalization judgment data by crowdsourcing. Our experimental results suggest that LMs have different generalizations from humans; LMs exhibited less context-dependent behaviors toward topicalization judgment. These results highlight the need for the additional inductive biases to guide LMs to achieve successful discourse-level generalization.
Riki Fujihara, Tatsuki Kuribayashi, Kaori Abe, Ryoko Tokuhisa, Kentaro Inui
COLING2
2022 Context Limitations Make Neural Language Models More Human-Like
abstract
Language models (LMs) have been used in cognitive modeling as well as engineering studiesthey compute information-theoretic complexity metrics that simulate humans' cognitive load during reading.This study highlights a limitation of modern neural LMs as the model of choice for this purpose: there is a discrepancy between their context access capacities and that of humans.Our results showed that constraining the LMs' context access improved their simulation of human reading behavior.We also showed that LM-human gaps in context access were associated with specific syntactic constructions; incorporating syntactic biases into LMs' context access might enhance their cognitive plausibility.1
Tatsuki Kuribayashi, Yohei Oseki, Ana Brassard, Kentaro Inui
EMNLP1
2021 Lower Perplexity is Not Always Human-Like
abstract
Tatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida, Masayuki Asahara, Kentaro Inui. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Tatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida, Masayuki Asahara, Kentaro Inui
ACL/IJCNLP (1)1
2021 Incorporating Residual and Normalization Layers into Analysis of Masked Language Models
abstract
Transformer architecture has become ubiquitous in the natural language processing field.To interpret the Transformer-based models, their attention patterns have been extensively analyzed.However, the Transformer architecture is not only composed of the multihead attention; other components can also contribute to Transformers' progressive performance.In this study, we extended the scope of the analysis of Transformers from solely the attention patterns to the whole attention block, i.e., multi-head attention, residual connection, and layer normalization.Our analysis of Transformer-based masked language models shows that the token-to-token interaction performed via attention has less impact on the intermediate representations than previously assumed.These results provide new intuitive explanations of existing reports; for example, discarding the learned attention patterns tends not to adversely affect the performance.The codes of our experiments are publicly available.
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro Inui
EMNLP (1)2
2021 Instance-Based Neural Dependency Parsing
abstract
Abstract Interpretable rationales for model predictions are crucial in practical applications. We develop neural models that possess an interpretable inference process for dependency parsing. Our models adopt instance-based inference, where dependency edges are extracted and labeled by comparing them to edges in a training set. The training edges are explicitly used for the predictions; thus, it is easy to grasp the contribution of each edge to the predictions. Our experiments show that our instance-based models achieve competitive accuracy with standard neural models and have the reasonable plausibility of instance-based explanations.
Hiroki Ouchi, Jun Suzuki 0001, Sosuke Kobayashi, Sho Yokoi, Tatsuki Kuribayashi, Masashi Yoshikawa, Kentaro Inui
Trans. Assoc. Comput. Linguistics5
2020 Language Models as an Alternative Evaluator of Word Order Hypotheses: A Case Study in Japanese
abstract
We examine a methodology using neural language models (LMs) for analyzing the word order of language.This LM-based method has the potential to overcome the difficulties existing methods face, such as the propagation of preprocessor errors in count-based methods.In this study, we explore whether the LMbased method is valid for analyzing the word order.As a case study, this study focuses on Japanese due to its complex and flexible word order.To validate the LM-based method, we test (i) parallels between LMs and human word order preference, and (ii) consistency of the results obtained using the LM-based method with previous linguistic studies.Through our experiments, we tentatively conclude that LMs display sufficient word order knowledge for usage as an analysis tool.Finally, using the LMbased method, we demonstrate the relationship between the canonical word order and topicalization, which had yet to be analyzed by largescale experiments.C Data used in Section 5.2, Section 6, and Appendix F
Tatsuki Kuribayashi, Takumi Ito, Jun Suzuki 0001, Kentaro Inui
ACL1
2020 Instance-Based Learning of Span Representations: A Case Study through Named Entity Recognition
abstract
Interpretable rationales for model predictions play a critical role in practical applications.In this study, we develop models possessing interpretable inference process for structured prediction.Specifically, we present a method of instance-based learning that learns similarities between spans.At inference time, each span is assigned a class label based on its similar spans in the training set, where it is easy to understand how much each training instance contributes to the predictions.Through empirical analysis on named entity recognition, we demonstrate that our method enables to build models that have high interpretability without sacrificing performance.
Hiroki Ouchi, Jun Suzuki 0001, Sosuke Kobayashi, Sho Yokoi, Tatsuki Kuribayashi, Ryuto Konno, Kentaro Inui
ACL5
2020 Modeling Event Salience in Narratives via Barthes' Cardinal Functions
abstract
Events in a narrative differ in salience: some are more important to the story than others.Estimating event salience is useful for tasks such as story generation, and as a tool for text analysis in narratology and folkloristics.To compute event salience without any annotations, we adopt Barthes' definition of event salience and propose several unsupervised methods that require only a pre-trained language model.Evaluating the proposed methods on folktales with event salience annotation, we show that the proposed methods outperform baseline methods and find fine-tuning a language model on narrative texts is a key factor in improving the proposed methods.
Takaki Otake, Sho Yokoi, Naoya Inoue, Tatsuki Kuribayashi, Kentaro Inui
COLING5
2020 Attention is Not Only a Weight: Analyzing Transformers with Vector Norms
abstract
Attention is a key component of Transformers, which have recently achieved considerable success in natural language processing. Hence, attention is being extensively studied to investigate various linguistic capabilities of Transformers, focusing on analyzing the parallels between attention weights and specific linguistic phenomena. This paper shows that attention weights alone are only one of the two factors that determine the output of attention and proposes a norm-based analysis that incorporates the second factor, the norm of the transformed input vectors. The findings of our norm-based analyses of BERT and a Transformer-based neural machine translation system include the following: (i) contrary to previous studies, BERT pays poor attention to special tokens, and (ii) reasonable word alignment can be extracted from attention mechanisms of Transformer. These findings provide insights into the inner workings of Transformers.
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro Inui
EMNLP (1)2
2019 An Empirical Study of Span Representations in Argumentation Structure Parsing
abstract
Tatsuki Kuribayashi, Hiroki Ouchi, Naoya Inoue, Paul Reisert, Toshinori Miyoshi, Jun Suzuki, Kentaro Inui. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Tatsuki Kuribayashi, Hiroki Ouchi, Naoya Inoue, Paul Reisert, Toshinori Miyoshi, Jun Suzuki 0001, Kentaro Inui
ACL (1)1
2019 Diamonds in the Rough: Generating Fluent Sentences from Early-Stage Drafts for Academic Writing Assistance
abstract
The writing process consists of several stages such as drafting, revising, editing, and proofreading. Studies on writing assistance, such as grammatical error correction (GEC), have mainly focused on sentence editing and proofreading, where surface-level issues such as typographical errors, spelling errors, or grammatical errors should be corrected. We broaden this focus to include the earlier revising stage, where sentences require adjustment to the information included or major rewriting and propose Sentence-level Revision (SentRev) as a new writing assistance task. Well-performing systems in this task can help inexperienced authors by producing fluent, complete sentences given their rough, incomplete drafts. We build a new freely available crowdsourced evaluation dataset consisting of incomplete sentences authored by non-native writers paired with their final versions extracted from published academic papers for developing and evaluating SentRev models. We also establish baseline performance on SentRev using our newly built evaluation dataset.
Takumi Ito, Tatsuki Kuribayashi, Hayato Kobayashi, Ana Brassard, Masato Hagiwara, Jun Suzuki 0001, Kentaro Inui
INLG2