Ryo Nagata

dblp:03/2470 · DBLP profile ↗
← Back
30ranked-venue papers
22as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 19 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Information extraction and text analysis · 56% Language models and text generation · 22% Representation and self-supervised learning · 22%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Computing education · 86% Computational social science and digital humanities · 14%
Human-computer interaction and pervasive computing
2 papers
Learning and educational technologies · 88% Human-robot interaction · 12%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
lexical semantics
1.832025
A New Formulation of Zipf's Meaning-Frequency Law through Contextual Diversity · ACL (1) 2025
Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport · ACL (1) 2025
Reinforcing English Countability Prediction with One Countability per Discourse Property · ACL 2006
Natural language and speech › Information extraction and text analysis › lexical semantics
semantic change detection
1.522025
Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport · ACL (1) 2025
Variance Matters: Detecting Semantic Differences without Corpus/Word Alignment · EMNLP 2023
Machine learning › Representation and self-supervised learning › word representation
contextualized word representation
0.912025
A New Formulation of Zipf's Meaning-Frequency Law through Contextual Diversity · ACL (1) 2025
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.712023
Variance Matters: Detecting Semantic Differences without Corpus/Word Alignment · EMNLP 2023
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.512021
Exploring Methods for Generating Feedback Comments for Writing Learning · EMNLP (1) 2021
Natural language and speech › Language models and text generation
text generation
0.512021
Exploring Methods for Generating Feedback Comments for Writing Learning · EMNLP (1) 2021
Natural language and speech › Language models and text generation › text generation › text response generation
feedback generation
0.412019
Toward a Task of Feedback Comment Generation for Writing Learning · EMNLP/IJCNLP (1) 2019
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.212016
Phrase Structure Annotation and Parsing for Learner English · ACL (1) 2016
Natural language and speech › Language models and text generation › text generation
grammatical error correction
0.212014
Correcting Preposition Errors in Learner English Using Error Case Frames and Feedback Messages · ACL (1) 2014
Computational social science and digital humanities
historical linguistics
0.212013
Reconstructing an Indo-European Family Tree from Non-native English Texts · ACL (1) 2013
Natural language and speech › Information extraction and text analysis › data annotation
corpus annotation
0.112011
Creating a manually error-tagged and shallow-parsed learner corpus · ACL 2011
Learning and educational technologies › robot-assisted learning
robot-assisted language learning
0.112011
The chanty bear: a new application for hri research · HRI 2011
Natural language and speech › Information extraction and text analysis › error detection
grammatical error detection
0.112006
A Feedback-Augmented Method for Detecting Errors in the Writing of Learners of English · ACL 2006
Natural language and speech › Information extraction and text analysis
word sense disambiguation
0.112006
Reinforcing English Countability Prediction with One Countability per Discourse Property · ACL 2006
Human-robot interaction
social robot
0.012011
The chanty bear: a new application for hri research · HRI 2011

Methods — techniques the papers use, named apart from their topics

retrieve-and-edit · 1.0pointer-generator · 1.0neural retrieval · 1.0unbalanced optimal transport · 0.9masked language model · 0.9contextualized word vectors · 0.9contextualized word embeddings · 0.9autoregressive language model · 0.9variance measurement · 0.7mean word vector norm · 0.7neural retrieval-based generation · 0.4phrase structure annotation · 0.2learner-native corpus comparison · 0.2error case frames · 0.2phylogenetic reconstruction · 0.2shallow parsing · 0.1jazz chants rhythmic teaching · 0.1error tagging · 0.1
YearPublicationVenuePosition
2025 Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport
abstract
Lexical semantic change detection aims to identify shifts in word meanings over time. While existing methods using embeddings from a diachronic corpus pair estimate the degree of change for target words, they offer limited insight into changes at the level of individual usage instances. To address this, we apply Unbalanced Optimal Transport (UOT) to sets of contextualized word embeddings, capturing semantic change through the excess and deficit in the alignment between usage instances. In particular, we propose Sense Usage Shift (SUS), a measure that quantifies changes in the usage frequency of a word sense at each usage instance. By leveraging SUS, we demonstrate that several challenges in semantic change detection can be addressed in a unified manner, including quantifying instance-level semantic change and word-level tasks such as measuring the magnitude of semantic change and the broadening or narrowing of meaning.
Ryo Kishino, Hiroaki Yamagiwa, Ryo Nagata, Sho Yokoi, Hidetoshi Shimodaira
ACL (1)3
2025 A New Formulation of Zipf's Meaning-Frequency Law through Contextual Diversity
abstract
This paper proposes formulating Zipf’s meaning-frequency law, the power law between word frequency and the number of meanings, as a relationship between word frequency and contextual diversity. The proposed formulation quantifies meaning counts as contextual diversity, which is based on the directions of contextualized word vectors obtained from a Language Model (LM). This formulation gives a new interpretation to the law and also enables us to examine it for a wider variety of words and corpora than previous studies have explored. In addition, this paper shows that the law becomes unobservable when the size of the LM used is small and that autoregressive LMs require much more parameters than masked LMs to be able to observe the law.
Ryo Nagata, Kumiko Tanaka-Ishii
ACL (1)1
2024 A Computational Approach to Quantifying Grammaticization of English Deverbal Prepositions
abstract
This paper explores grammaticization of deverbal prepositions by a computational approach based on corpus data. Deverbal prepositions are words or phrases that are derived from a verb and that behave as a preposition such as “regarding” and “according to”. Linguistic studies have revealed important aspects of grammaticization of deverbal prepositions. This paper augments them by methods for measuring the degree of grammaticization of deverbal prepositions based on non-contextualized or contextualized word vectors. Experiments show that the methods correlate well with human judgements (as high as 0.69 in Spearman’s rank correlation coefficient). Using the best-performing method, this paper further shows that the methods support previous findings in linguistics including (i) Deverbal prepositions are marginal in terms of prepositionality; and (ii) The process where verbs are grammaticized into prepositions is gradual. As a pilot study, it also conducts a diachronic analysis of grammaticization of deverbal preposition.
Ryo Nagata, Yoshifumi Kawasaki, Naoki Otani, Hiroya Takamura
LREC/COLING1
2024 n-gram F-score for Evaluating Grammatical Error Correction
abstract
M 2 and its variants are the most widely used automatic evaluation metrics for grammatical error correction (GEC), which calculate an F -score using a phrase-based alignment between sentences.However, it is not straightforward at all to align learner sentences containing errors to their correct sentences.In addition, alignment calculations are computationally expensive.We propose GREEN, an alignment-free F -score for GEC evaluation.GREEN treats a sentence as a multiset of ngrams and extracts edits between sentences by set operations instead of computing an alignment.Our experiments confirm that GREEN performs better than existing methods for the corpus-level metrics and comparably for the sentence-level metrics even without computing an alignment.GREEN is available at https://github.com/shotakoyama/green.
Shota Koyama, Ryo Nagata, Hiroya Takamura, Naoaki Okazaki
INLG2
2023 Variance Matters: Detecting Semantic Differences without Corpus/Word Alignment
abstract
In this paper, we propose methods 1 for discovering semantic differences in words appearing in two corpora.The key idea is to measure the coverage of meanings of a word in a corpus through the norm of its mean word vector, which is equivalent to examining a kind of variance of the word vector distribution.The proposed methods do not require alignments between words and/or corpora for comparison that previous methods do.All they require are to compute variance (or norms of mean word vectors) for each word type.Nevertheless, they rival the best-performing system in the SemEval-2020 Task 1.In addition, they are (i) robust for the skew in corpus sizes; (ii) capable of detecting semantic differences in infrequent words; and (iii) effective in pinpointing word instances that have a meaning missing in one of the two corpora under comparison.We show these advantages for historical corpora and also for native/non-native English corpora.
Ryo Nagata, Hiroya Takamura, Naoki Otani, Yoshifumi Kawasaki
EMNLP1
2022 Revisiting Statistical Laws of Semantic Shift in Romance Cognates
abstract
This article revisits statistical relationships across Romance cognates between lexical semantic shift and six intra-linguistic variables, such as frequency and polysemy. Cognates are words that are derived from a common etymon, in this case, a Latin ancestor. Despite their shared etymology, some cognate pairs have experienced semantic shift. The degree of semantic shift is quantified using cosine distance between the cognates’ corresponding word embeddings. In the previous literature, frequency and polysemy have been reported to be correlated with semantic shift; however, the understanding of their effects needs revision because of various methodological defects. In the present study, we perform regression analysis under improved experimental conditions, and demonstrate a genuine negative effect of frequency and positive effect of polysemy on semantic shift. Furthermore, we reveal that morphologically complex etyma are more resistant to semantic shift and that the cognates that have been in use over a longer timespan are prone to greater shift in meaning. These findings add to our understanding of the historical process of semantic change.
Yoshifumi Kawasaki, Maëlys Salingre, Marzena Karpinska, Hiroya Takamura, Ryo Nagata
COLING5
2021 Exploring Methods for Generating Feedback Comments for Writing Learning
abstract
The task of generating explanatory notes for language learners is known as feedback comment generation.Although various generation techniques are available, little is known about which methods are appropriate for this task.Nagata (2019) demonstrates the effectiveness of neural-retrieval-based methods in generating feedback comments for preposition use.Retrieval-based methods have limitations in that they can only output feedback comments existing in a given training data.Furthermore, feedback comments can be made on other grammatical and writing items than preposition use, which is still unaddressed.To shed light on these points, we investigate a wider range of methods for generating many feedback comments in this study.Our close analysis of the type of task leads us to investigate three different architectures for comment generation: (i) a neural-retrieval-based method as a baseline, (ii) a pointer-generator-based generation method as a neural seq2seq method, (iii) a retrieve-and-edit method, a hybrid of (i) and (ii).Intuitively, the pointer-generator should outperform neural-retrieval, and retrieve-andedit should perform best.However, in our experiments, this expectation is completely overturned.We closely analyze the results to reveal the major causes of these counter-intuitive results and report on our findings from the experiments.1
Kazuaki Hanawa, Ryo Nagata, Kentaro Inui
EMNLP (1)2
2021 Shared Task on Feedback Comment Generation for Language Learners
abstract
In this paper, we propose a generation challenge called Feedback comment generation for language learners.It is a task where given a text and a span, a system generates, for the span, an explanatory note that helps the writer (language learner) improve their writing skills.The motivations for this challenge are: (i) practically, it will be beneficial for both language learners and teachers if a computerassisted language learning system can provide feedback comments just as human teachers do; (ii) theoretically, feedback comment generation for language learners has a mixed aspect of other generation tasks together with its unique features and it will be interesting to explore what kind of generation method is effective against what kind of writing rule.To this end, we have created a dataset and developed baseline systems to estimate baseline performance.With these preparations, we propose a generation challenge of feedback comment generation.
Ryo Nagata, Masato Hagiwara, Kazuaki Hanawa, Masato Mita, Artem Chernodub, Olena Nahorna
INLG1
2020 Taking the Correction Difficulty into Account in Grammatical Error Correction Evaluation
abstract
This paper presents performance measures for grammatical error correction which take into account the difficulty of error correction.To the best of our knowledge, no conventional measure has such functionality despite the fact that some errors are easy to correct and others are not.The main purpose of this work is to provide a way of determining the difficulty of error correction and to motivate researchers in the domain to attack such difficult errors.The performance measures are based on the simple idea that the more systems successfully correct an error, the easier it is considered to be.This paper presents a set of algorithms to implement this idea.It evaluates the performance measures quantitatively and qualitatively on a wide variety of corpora and systems, revealing that they agree with our intuition of correction difficulty.A scorer and difficulty weight data based on the algorithms have been made available on the web.
Takumi Gotou, Ryo Nagata, Masato Mita, Kazuaki Hanawa
COLING2
2020 Creating Corpora for Research in Feedback Comment Generation
abstract
In this paper, we report on datasets that we created for research in feedback comment generation — a task of automatically generating feedback comments such as a hint or an explanatory note for writing learning. There has been almost no such corpus open to the public and accordingly there has been a very limited amount of work on this task. In this paper, we first discuss the principle and guidelines for feedback comment annotation. Then, we describe two corpora that we have manually annotated with feedback comments (approximately 50,000 general comments and 6,700 on preposition use). A part of the annotation results is now available on the web, which will facilitate research in feedback comment generation
Ryo Nagata, Kentaro Inui, Shin'ichiro Ishikawa
LREC1
2019 *Paris is Rain. or It is raining in Paris?: Detecting Overgeneralization of Be-verb in Learner English
Ryo Nagata, Koki Washio, Hokuto Ototake
CICLing (1)1
2019 Toward a Task of Feedback Comment Generation for Writing Learning
abstract
In this paper, we introduce a novel task called feedback comment generation — a task of automatically generating feedback comments such as a hint or an explanatory note for writing learning for non-native learners of English. There has been almost no work on this task nor corpus annotated with feedback comments. We have taken the first step by creating learner corpora consisting of approximately 1,900 essays where all preposition errors are manually annotated with feedback comments. We have tested three baseline methods on the dataset, showing that a simple neural retrieval-based method sets a baseline performance with an F-measure of 0.34 to 0.41. Finally, we have looked into the results to explore what modifications we need to make to achieve better performance. We also have explored problems unaddressed in this work
Ryo Nagata
EMNLP/IJCNLP (1)1
2018 Exploring the Influence of Spelling Errors on Lexical Variation Measures
abstract
This paper explores the influence of spelling errors on lexical variation measures. Lexical richness measures such as Type-Token Ration (TTR) and Yule’s K are often used for learner English analysis and assessment. When applied to learner English, however, they can be unreliable because of the spelling errors appearing in it. Namely, they are, directly or indirectly, based on the counts of distinct word types, and spelling errors undesirably increase the number of distinct words. This paper introduces and examines the hypothesis that lexical richness measures become unstable in learner English because of spelling errors. Specifically, it tests the hypothesis on English learner corpora of three groups (middle school, high school, and college students). To be precise, it estimates the difference in TTR and Yule’s K caused by spelling errors, by calculating their values before and after spelling errors are manually corrected. Furthermore, it examines the results theoretically and empirically to deepen the understanding of the influence of spelling errors on them.
Ryo Nagata, Taisei Sato, Hiroya Takamura
COLING1
2017 Analyzing Semantic Change in Japanese Loanwords
abstract
We analyze semantic changes in loanwords from English that are used in Japanese (Japanese loanwords). Specifically, we create word embeddings of English and Japanese and map the Japanese embeddings into the English space so that we can calculate the similarity of each Japanese word and each English word. We then attempt to find loanwords that are semantically different from their original, see if known meaning changes are correctly captured, and show the possibility of using our methodology in language education.
Hiroya Takamura, Ryo Nagata, Yoshifumi Kawasaki
EACL (1)2
2017 Adaptive Spelling Error Correction Models for Learner English
abstract
Spelling errors are a characteristic of learner English and degrade the performances of natural language processing systems targeting English learners. This paper describes a method specially designed for automatically correcting spelling errors in learner English that reduces the effects from noise (e.g., grammatical and spelling errors) by adaptively creating spelling error correction models from raw learner corpora. An evaluation shows that the proposed method outperforms previous edit-distance-based and language-model-based methods. We also report results of an investigation into what types of spelling errors English learners tend to make, using the spelling error models created by the proposed method as a tool for our analysis.
Ryo Nagata, Hiroya Takamura, Graham Neubig
KES1
2016 Phrase Structure Annotation and Parsing for Learner English
Ryo Nagata, Keisuke Sakaguchi
ACL (1)1
2016 Discriminative Analysis of Linguistic Features for Typological Study
Hiroya Takamura, Ryo Nagata, Yoshifumi Kawasaki
LREC2
2014 Correcting Preposition Errors in Learner English Using Error Case Frames and Feedback Messages
abstract
This paper presents a novel framework called error case frames for correcting preposition errors. They are case frames specially designed for describing and cor-recting preposition errors. Their most dis-tinct advantage is that they can correct er-rors with feedback messages explaining why the preposition is erroneous. This pa-per proposes a method for automatically generating them by comparing learner and native corpora. Experiments show (i) au-tomatically generated error case frames achieve a performance comparable to con-ventional methods; (ii) error case frames are intuitively interpretable and manually modifiable to improve them; (iii) feedback messages provided by error case frames are effective in language learning assis-tance. Considering these advantages and the fact that it has been difficult to provide feedback messages by automatically gen-erated rules, error case frames will likely be one of the major approaches for prepo-sition error correction. 1
Ryo Nagata, Mikko Vilenius, Edward W. D. Whittaker
ACL (1)1
2014 Language Family Relationship Preserved in Non-native English
Ryo Nagata
COLING1
2013 Reconstructing an Indo-European Family Tree from Non-native English Texts
Ryo Nagata, Edward W. D. Whittaker
ACL (1)1
2012 A method for detecting tense errors in learner English
abstract
Although tense errors are one major source of grammatical errors in learner English, there has been almost no work on their detection. Tense error detection seems extremely difficult considering that its determination greatly relies on intention and context. Despite the difficulties, this paper shows that tense error can be efficiently detected by exploiting a linguistic property of English verbs called stativity. Our proposed method predicts the stativity of the verbs and then detects tense errors based on the prediction. Experiments show that it achieves an F-measure of 0.571 and outperforms methods implemented for comparison.
Ryo Nagata, Vera Sheinman
KES1
2011 Creating a manually error-tagged and shallow-parsed learner corpus
Ryo Nagata, Edward W. D. Whittaker, Vera Sheinman
ACL1
2011 The chanty bear: a new application for hri research
abstract
This paper presents yet another English-teaching robot, while putting emphasis on the merits which are offered by second language education to human robot interaction (HRI) research. The chanty bear, our prototype robot based on a rhythmic teaching method of English called Jazz Chants is introduced.
Kotaro Funakoshi, Tomoya Mizumoto, Ryo Nagata, Mikio Nakano
HRI3
2011 Exploiting Learners' Tendencies for Detecting English Determiner Errors
Ryo Nagata, Atsuo Kawai
KES (2)1
2009 Edu-mining for Book Recommendation for Pupils
Ryo Nagata, Keigo Takeda, Koji Suda, Jun'ichi Kakegawa, Koichiro Morihiro
EDM1
2009 A Topic-Independent Method for Automatically Scoring Essay Content Rivaling Topic-Dependent Methods
abstract
This paper proposes a topic-independent method for automatically scoring essay content. Unlike conventional topic-dependent methods, it predicts the human score of a given essay without training essays written to the same topic as the target essay. To achieve this, this paper introduces a new measure called MIDF that measures how important and relevant a word is in a given essay. The proposed method predicts the score relying on the distribution of MIDF. Surprisingly, experiments show that the proposed method achieves an accuracy of 0.848 and performs as well as or even better than conventional topic-dependent methods.
Ryo Nagata, Jun'ichi Kakegawa, Yukiko Yabuta
ICALT1
2007 Edu-mining for finding keywords to improve message-production skills
abstract
This paper describes an edu-mining technique for finding keywords to improve pupils' message- production skills. It automatically finds keywords from blog items pupils create. It then adaptively suggests some of the keywords to pupils when they create new blog items so that they can revise their items using the keywords. Experiments show that it finds and suggests keywords surprisingly well despite its simple algorithm. They also show that pupils in an elementary school in Japan successfully improve the quality of blog items using the suggested keywords.
Ryo Nagata, Koji Suda, Jun'ichi Kakegawa, Koichiro Morihiro, Kazuhiko Showji
ICALT1
2006 A Feedback-Augmented Method for Detecting Errors in the Writing of Learners of English
abstract
This paper proposes a method for detecting errors in article usage and singular plural usage based on the mass count distinction. First, it learns decision lists from training data generated automatically to distinguish mass and count nouns. Then, in order to improve its performance, it is augmented by feedback that is obtained from the writing of learners. Finally, it detects errors by applying rules to the mass count distinction. Experiments show that it achieves a recall of 0.71 and a precision of 0.72 and outperforms other methods used for comparison when augmented by feedback.
Ryo Nagata, Atsuo Kawai, Koichiro Morihiro, Naoki Isu
ACL1
2006 Reinforcing English Countability Prediction with One Countability per Discourse Property
Ryo Nagata, Atsuo Kawai, Koichiro Morihiro, Naoki Isu
ACL1
2005 Detecting Article Errors Based on the Mass Count Distinction
Ryo Nagata, Takahiro Wakana, Fumito Masui, Atsuo Kawai, Naoki Isu
IJCNLP1