VLDB 2026 Research / reviewers in the wild / expert
Mamoru Komachi
dblp:88/2433
· DBLP profile ↗
60ranked-venue papers
2as first author
30since 2021 · last 2026
0000-0003-1166-1739ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 59 · 2 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Assessing the Capabilities of LLMs in Humor: A Multi-dimensional Analysis of Oogiri Generation and EvaluationabstractComputational humor is a frontier for creating advanced and engaging natural language processing (NLP) applications, such as sophisticated dialogue systems. While previous studies have benchmarked the humor capabilities of Large Language Models (LLMs), they have often relied on single-dimensional evaluations, such as judging whether something is simply ``funny.'' This paper argues that a multifaceted understanding of humor is necessary and addresses this gap by systematically evaluating LLMs through the lens of Oogiri, a form of Japanese improvisational comedy games. To achieve this, we expanded upon existing Oogiri datasets with data from new sources and then augmented the collection with Oogiri responses generated by LLMs. We then manually annotated this expanded collection with 5-point absolute ratings across six dimensions: Novelty, Clarity, Relevance, Intelligence, Empathy, and Overall Funniness. Using this dataset, we assessed the capabilities of state-of-the-art LLMs on two core tasks: their ability to generate creative Oogiri responses and their ability to evaluate the funniness of responses using a six-dimensional evaluation. Our results show that while LLMs can generate responses at a level between low- and mid-tier human performance, they exhibit a notable lack of Empathy. This deficit in Empathy helps explain their failure to replicate human humor assessment. Correlation analyses of human and model evaluation data further reveal a fundamental divergence in evaluation criteria: LLMs prioritize Novelty, whereas humans prioritize Empathy. We release our annotated corpus to the community to pave the way for the development of more emotionally intelligent and sophisticated conversational agents. Ritsu Sakabe, Hwichan Kim, Tosho Hirasawa, Mamoru Komachi |
AAAI | 4 |
| 2026 | Can Video LLMs See Through Illusions? Video-Illusion QA Benchmark Dataset
Souto Ohira, Tosho Hirasawa, Mamoru Komachi |
LREC | 3 |
| 2026 | Evaluation of Document-Level Text Simplification in Japanese
Iori Yamashita, Hikari Tanaka, Hajime Kiyama, Kexin Bian, Zhousi Chen, Mamoru Komachi |
LREC | 6 |
| 2025 | Targeted Syntactic Evaluation for Grammatical Error CorrectionabstractLanguage learners encounter a wide range of grammar items across the beginner, intermediate, and advanced levels.To develop grammatical error correction (GEC) models effectively, it is crucial to identify which grammar items are easier or more challenging for models to correct.However, conventional benchmarks based on learner-produced texts are insufficient for conducting detailed evaluations of GEC model performance across a wide range of grammar items due to biases in their distribution.To address this issue, we propose a new evaluation paradigm that assesses GEC models using minimal pairs of ungrammatical and grammatical sentences for each grammar item.As the first benchmark within this paradigm, we introduce the CEFR-based Targeted Syntactic Evaluation Dataset for Grammatical Error Correction (CTSEG), which complements existing English benchmarks by enabling fine-grained analyses previously unattainable with conventional datasets.Using CTSEG, we evaluate three mainstream types of English GEC models: sequence-to-sequence models, sequence tagging models, and prompt-based models.The results indicate that while current models perform well on beginner-level grammar items, their performance deteriorates substantially for intermediate and advanced items.Dataset Sents.Refs.Error Tags CEFR Level FCE (Yannakoudakis et al., 2011) 2,695 1 71 B1-B2 KJ (Nagata et al., 2011) 3,199 1 22 A1-A2?CoNLL-2014 (Ng et al., 2014) 1,312 2 28 C1 AESW (Daudaravicius et al., 2016) 143,804 1 N/A C1-C2 (+Native) JFLEG (Napoles et al., 2017) 747 4 N/A A1-C2?BEA-2019 (Bryant et al., 2019) 4,384 5 25 A1-C2 (+Native) GMEG (Napoles et al., 2019) 2,960 4 N/A B1-B2 (+Native) CWEB (Flachs et al., 2020) Aomi Koyama, Masato Mita, Su-Youn Yoon, Yasufumi Takama, Mamoru Komachi |
ACL (1) | 5 |
| 2025 | Investigating the Impact of Japanese Names and Japanese Prompts on Social Bias in Hiring Decisions Using LLMs
Saneyuki Okabe, Taisei Enomoto, Mamoru Komachi, Atsushi Keyaki |
IEEE Big Data | 3 |
| 2025 | Analyzing Continuous Semantic Shifts with Diachronic Word Similarity MatricesabstractThe meanings and relationships of words shift over time. This phenomenon is referred to as semantic shift. Research focused on understanding how semantic shifts occur over multiple time periods is essential for gaining a detailed understanding of semantic shifts. However, detecting change points only between adjacent time periods is insufficient for analyzing detailed semantic shifts, and using BERT-based methods to examine word sense proportions incurs a high computational cost. To address those issues, we propose a simple yet intuitive framework for how semantic shifts occur over multiple time periods by utilizing similarity matrices based on word embeddings. We calculate diachronic word similarity matrices using fast and lightweight word embeddings across arbitrary time periods, making it deeper to analyze continuous semantic shifts. Additionally, by clustering the resulting similarity matrices, we can categorize words that exhibit similar behavior of semantic shift in an unsupervised manner. Hajime Kiyama, Taichi Aida, Mamoru Komachi, Toshinobu Ogiso, Hiroya Takamura, Daichi Mochihashi |
COLING | 3 |
| 2024 | WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in WikipediaabstractWikipedia can be edited by anyone and thus contains various quality sentences. Therefore, Wikipedia includes some poor-quality edits, which are often marked up by other editors. While editors' reviews enhance the credibility of Wikipedia, it is hard to check all edited text. Assisting in this process is very important, but a large and comprehensive dataset for studying it does not currently exist. Here, we propose WikiSQE, the first large-scale dataset for sentence quality estimation in Wikipedia. Each sentence is extracted from the entire revision history of English Wikipedia, and the target quality labels were carefully investigated and selected. WikiSQE has about 3.4 M sentences with 153 quality labels. In the experiment with automatic classification using competitive machine learning models, sentences that had problems with citation, syntax/semantics, or propositions were found to be more difficult to detect. In addition, by performing human annotation, we found that the model we developed performed better than the crowdsourced workers. WikiSQE is expected to be a valuable resource for other tasks in NLP. Kenichiro Ando, Satoshi Sekine, Mamoru Komachi |
AAAI | 3 |
| 2024 | A Document-Level Text Simplification Dataset for JapaneseabstractDocument-level text simplification, a task that combines single-document summarization and intra-sentence simplification, has garnered significant attention. However, studies have primarily focused on languages such as English and German, leaving Japanese and similar languages underexplored because of a scarcity of linguistic resources. In this study, we devised JADOS, the first Japanese document-level text simplification dataset based on newspaper articles and Wikipedia. Our dataset focuses on simplification, to enhance readability by reducing the number of sentences and tokens in a document. We conducted investigations using our dataset. Firstly, we analyzed the characteristics of Japanese simplification by comparing it across different domains and with English counterparts. Moreover, we experimentally evaluated the performances of text summarization methods, transformer-based text simplification models, and large language models. In terms of D-SARI scores, the transformer-based models performed best across all domains. Finally, we manually evaluated several model outputs and target articles, demonstrating the need for document-level text simplification models in Japanese. Yoshinari Nagai, Teruaki Oka, Mamoru Komachi |
LREC/COLING | 3 |
| 2024 | Token-length Bias in Minimal-pair Paradigm DatasetsabstractMinimal-pair paradigm datasets have been used as benchmarks to evaluate the linguistic knowledge of models and provide an unsupervised method of acceptability judgment. The model performances are evaluated based on the percentage of minimal pairs in the MPP dataset where the model assigns a higher sentence log-likelihood to an acceptable sentence than to an unacceptable sentence. Each minimal pair in MPP datasets is controlled to align the number of words per sentence because the sentence length affects the sentence log-likelihood. However, aligning the number of words may be insufficient because recent language models tokenize sentences with subwords. Tokenization may cause a token length difference in minimal pairs, introducing token-length bias that skews the evaluation results. This study demonstrates that MPP datasets suffer from token-length bias and fail to evaluate the linguistic knowledge of a language model correctly. The results proved that sentences with a shorter token length would likely be assigned a higher log-likelihood regardless of their acceptability, which becomes problematic when comparing models with different tokenizers. To address this issue, we propose a debiased minimal pair generation method, allowing MPP datasets to measure language ability correctly and provide comparable results for all models. Naoya Ueda, Masato Mita, Teruaki Oka, Mamoru Komachi |
LREC/COLING | 4 |
| 2024 | A Survey for LLM Tuning Methods:Classifying Approaches Based on Model Internal Accessibility
Kyotaro Nakajima, Hwichan Kim, Tosho Hirasawa, Taisei Enomoto, Zhousi Chen, Mamoru Komachi |
PACLIC | 6 |
| 2024 | DejaVu: Disambiguation evaluation dataset for English-JApanese machine translation on VisUal information
Ayako Sato, Tosho Hirasawa, Hwichan Kim, Zhousi Chen, Teruaki Oka, Masato Mita, Mamoru Komachi |
PACLIC | 7 |
| 2024 | Revisiting Meta-evaluation for Grammatical Error CorrectionabstractAbstract Metrics are the foundation for automatic evaluation in grammatical error correction (GEC), with their evaluation of the metrics (meta-evaluation) relying on their correlation with human judgments. However, conventional meta-evaluations in English GEC encounter several challenges, including biases caused by inconsistencies in evaluation granularity and an outdated setup using classical systems. These problems can lead to misinterpretation of metrics and potentially hinder the applicability of GEC techniques. To address these issues, this paper proposes SEEDA, a new dataset for GEC meta-evaluation. SEEDA consists of corrections with human ratings along two different granularities: edit-based and sentence-based, covering 12 state-of-the-art systems including large language models, and two human corrections with different focuses. The results of improved correlations by aligning the granularity in the sentence-level meta-evaluation suggest that edit-based metrics may have been underestimated in existing studies. Furthermore, correlations of most metrics decrease when changing from classical to neural systems, indicating that traditional metrics are relatively poor at evaluating fluently corrected sentences with many edits. Masamune Kobayashi, Masato Mita, Mamoru Komachi |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | Simultaneous Domain Adaptation of Tokenization and Machine Translation
Taisei Enomoto, Tosho Hirasawa, Hwichan Kim, Teruaki Oka, Mamoru Komachi |
PACLIC | 5 |
| 2023 | Construction of Evaluation Dataset for Japanese Lexical Semantic Change Detection
Zhidong Ling, Taichi Aida, Teruaki Oka, Mamoru Komachi |
PACLIC | 4 |
| 2023 | Discontinuous Combinatory Constituency ParsingabstractAbstract We extend a pair of continuous combinator-based constituency parsers (one binary and one multi-branching) into a discontinuous pair. Our parsers iteratively compose constituent vectors from word embeddings without any grammar constraints. Their empirical complexities are subquadratic. Our extension includes 1) a swap action for the orientation-based binary model and 2) biaffine attention for the chunker-based multi-branching model. In tests conducted with the Discontinuous Penn Treebank and TIGER Treebank, we achieved state-of-the-art discontinuous accuracy with a significant speed advantage. Zhousi Chen, Mamoru Komachi |
Trans. Assoc. Comput. Linguistics | 2 |
| 2023 | Dataset Enhancement and Multilingual Transfer for Named Entity Recognition in the Indonesian LanguageabstractNamed entity recognition in the Indonesian language has significantly developed in recent years. However, it still lacks standardized publicly available corpora; a small dataset is available but suffers from inconsistent annotations. Therefore, we re-annotated the dataset to improve its consistency and benefit the community. Our re-annotation led to better training results from an effective baseline model consisting of bidirectional long short-term memory and conditional random fields. To fully utilize the limited available data, we utilized better contextualization and transferred external knowledge by exploiting monolingual and multilingual pre-trained language models, such as IndoBERT and XLM-RoBERTa. In addition to the general improvement from the language models, we observed that the monolingual model is more sensitive, while the multilingual ones show advantages in rich morphological knowledge. We also applied cross-lingual transfer learning to utilize high-resource corpora in other languages. We adopted English, Spanish, Dutch, and German as the source languages for the target Indonesian language and found that Dutch plays a special role in the data transfer method due to morphological similarity attributable to historical reasons. Siti Oryza Khairunnisa, Zhousi Chen, Mamoru Komachi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | North Korean Neural Machine Translation through South Korean ResourcesabstractSouth and North Korea both use the Korean language. However, Korean natural language processing (NLP) research has mostly focused on South Korean language. Therefore, existing NLP systems in the Korean language, such as neural machine translation (NMT) systems, cannot properly process North Korean inputs. Training a model using North Korean data is the most straightforward approach to solving this problem, but the data to train NMT models are insufficient. To solve this problem, we constructed a parallel corpus to develop a North Korean NMT model using a comparable corpus. We manually aligned parallel sentences to create evaluation data and automatically aligned the remaining sentences to create training data. We trained a North Korean NMT model using our North Korean parallel data and improved North Korean translation quality using South Korean resources such as parallel data and a pre-trained model. In addition, we propose Korean-specific pre-processing methods, character tokenization, and phoneme decomposition to use the South Korean resources more efficiently. We demonstrate that the phoneme decomposition consistently improves the North Korean translation accuracy compared to other pre-processing methods. Hwichan Kim, Tosho Hirasawa, Sangwhan Moon, Naoaki Okazaki, Mamoru Komachi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2023 | Chinese Grammatical Error Correction Using Pre-trained Models and Pseudo DataabstractIn recent studies, pre-trained models and pseudo data have been key factors in improving the performance of the English grammatical error correction (GEC) task. However, few studies have examined the role of pre-trained models and pseudo data in the Chinese GEC task. Therefore, we develop Chinese GEC models based on three pre-trained models: Chinese BERT, Chinese T5, and Chinese BART, and then incorporate these models with pseudo data to determine the best configuration for the Chinese GEC task. On the natural language processing and Chinese computing (NLPCC) 2018 GEC shared task test set, all our single models outperform the ensemble models developed by the top team of the shared task. Chinese BART achieves an F score of 37.15, which is a state-of-the-art result. We then combine our Chinese GEC models with three kinds of pseudo data: Lang-8 (MaskGEC), Wiki (MaskGEC), and Wiki (Backtranslation). We find that most models can benefit from pseudo data, and BART+Lang-8 (MaskGEC) is the ideal setting in terms of accuracy and training efficiency. The experimental results demonstrate the effectiveness of the pre-trained models and pseudo data on the Chinese GEC task and provide an easily reproducible and adaptable baseline for future works. Finally, we annotate the error types of the development data; the results show that word-level errors dominate all error types, and word selection errors must be addressed even when using pre-trained models and pseudo data. Our codes are available at https://github.com/wang136906578/BERT-encoder-ChineseGEC . Michiki Kurosawa, Satoru Katsumata, Masato Mita, Mamoru Komachi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2022 | Infinite SCAN: An Infinite Model of Diachronic Semantic ChangeabstractIn this study, we propose a Bayesian model that can jointly estimate the number of senses of words and their changes through time.The model combines a dynamic topic model on Gaussian Markov random fields (Frermann and Lapata, 2016) with a logistic stick-breaking process that realizes the Dirichlet process.In the experiments, we evaluated the proposed model in terms of interpretability, accuracy in estimating the number of senses, and tracking their changes using both artificial data and real data.We quantitatively verified that the model behaves as expected through evaluation using artificial data.Using the CCOHA corpus, we showed that our model outperforms the baseline model and investigated the semantic changes of several well-known target words. Seiichi Inoue, Mamoru Komachi, Toshinobu Ogiso, Hiroya Takamura, Daichi Mochihashi |
EMNLP | 2 |
| 2022 | Learning How to Translate North Korean through South KoreanabstractSouth and North Korea both use the Korean language. However, Korean NLP research has focused on South Korean only, and existing NLP systems of the Korean language, such as neural machine translation (NMT) models, cannot properly handle North Korean inputs. Training a model using North Korean data is the most straightforward approach to solving this problem, but there is insufficient data to train NMT models. In this study, we create data for North Korean NMT models using a comparable corpus. First, we manually create evaluation data for automatic alignment and machine translation, and then, investigate automatic alignment methods suitable for North Korean. Finally, we show that a model trained by North Korean bilingual data without human annotation significantly boosts North Korean translation accuracy compared to existing South Korean models in zero-shot settings. Hwichan Kim, Sangwhan Moon, Naoaki Okazaki, Mamoru Komachi |
LREC | 4 |
| 2022 | Construction of a Quality Estimation Dataset for Automatic Evaluation of Japanese Grammatical Error CorrectionabstractIn grammatical error correction (GEC), automatic evaluation is considered as an important factor for research and development of GEC systems. Previous studies on automatic evaluation have shown that quality estimation models built from datasets with manual evaluation can achieve high performance in automatic evaluation of English GEC. However, quality estimation models have not yet been studied in Japanese, because there are no datasets for constructing quality estimation models. In this study, therefore, we created a quality estimation dataset with manual evaluation to build an automatic evaluation model for Japanese GEC. By building a quality estimation model using this dataset and conducting a meta-evaluation, we verified the usefulness of the quality estimation model for Japanese GEC. Daisuke Suzuki, Yujin Takahashi, Ikumi Yamashita, Taichi Aida, Tosho Hirasawa, Michitaka Nakatsuji, Masato Mita, Mamoru Komachi |
LREC | 8 |
| 2022 | ProQE: Proficiency-wise Quality Estimation dataset for Grammatical Error CorrectionabstractThis study investigates how supervised quality estimation (QE) models of grammatical error correction (GEC) are affected by the learners’ proficiency with the data. QE models for GEC evaluations in prior work have obtained a high correlation with manual evaluations. However, when functioning in a real-world context, the data used for the reported results have limitations because prior works were biased toward data by learners with relatively high proficiency levels. To address this issue, we created a QE dataset that includes multiple proficiency levels and explored the necessity of performing proficiency-wise evaluation for QE of GEC. Our experiments demonstrated that differences in evaluation dataset proficiency affect the performance of QE models, and proficiency-wise evaluation helps create more robust models. Yujin Takahashi, Masahiro Kaneko, Masato Mita, Mamoru Komachi |
LREC | 4 |
| 2022 | Japanese Named Entity Recognition from Automatic Speech Recognition Using Pre-trained Models
Seiichiro Kondo, Naoya Ueda, Teruaki Oka, Masakazu Sugiyama, Asahi Hentona, Mamoru Komachi |
PACLIC | 6 |
| 2022 | Region-attentive multimodal neural machine translationabstractWe propose a multimodal neural machine translation (MNMT) method with semantic image regions called region-attentive multimodal neural machine translation (RA-NMT). Existing studies on MNMT have mainly focused on employing global visual features or equally sized grid local visual features extracted by convolutional neural networks (CNNs) to improve translation performance. However, they neglect the effect of semantic information captured inside the visual features. This study utilizes semantic image regions extracted by object detection for MNMT and integrates visual and textual features using two modality-dependent attention mechanisms. The proposed method was implemented and verified on two neural architectures of neural machine translation (NMT): recurrent neural network (RNN) and self-attention network (SAN). Experimental results on different language pairs of Multi30k dataset show that our proposed method improves over baselines and outperforms most of the state-of-the-art MNMT methods. Further analysis demonstrates that the proposed method can achieve better translation performance because of its better visual feature use. Mamoru Komachi, Tomoyuki Kajiwara, Chenhui Chu |
Neurocomputing | 2 |
| 2022 | Word-Region Alignment-Guided Multimodal Neural Machine TranslationabstractWe propose word-region alignment-guided multimodal neural machine translation (MNMT), a novel model for MNMT that links the semantic correlation between textual and visual modalities using word-region alignment (WRA). Existing studies on MNMT have mainly focused on the effect of integrating visual and textual modalities. However, they do not leverage the semantic relevance between the two modalities. We advance the semantic correlation between textual and visual modalities in MNMT by incorporating WRA as a bridge. This proposal has been implemented on two mainstream architectures of neural machine translation (NMT): the recurrent neural network (RNN) and the transformer. Experiments on two public benchmarks, English–German and English–French translation tasks using the Multi30k dataset and English–Japanese translation tasks using the Flickr30kEnt-JP dataset prove that our model has a significant improvement with respect to the competitive baselines across different evaluation metrics and outperforms most of the existing MNMT models. For example, 1.0 BLEU scores are improved for the English–German task and 1.1 BLEU scores are improved for the English–French task on the Multi30k test2016 set; and 0.7 BLEU scores are improved for the English–Japanese task on the Flickr30kEnt-JP test set. Further analysis demonstrates that our model can achieve better translation performance by integrating WRA, leading to better visual information use. Mamoru Komachi, Tomoyuki Kajiwara, Chenhui Chu |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | From Masked Language Modeling to Translation: Non-English Auxiliary Tasks Improve Zero-shot Spoken Language UnderstandingabstractRob van der Goot, Ibrahim Sharaf, Aizhan Imankulova, Ahmet Üstün, Marija Stepanović, Alan Ramponi, Siti Oryza Khairunnisa, Mamoru Komachi, Barbara Plank. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Rob van der Goot, Ibrahim Sharaf, Aizhan Imankulova, Ahmet Üstün, Marija Stepanovic, Alan Ramponi, Siti Oryza Khairunnisa, Mamoru Komachi, Barbara Plank |
NAACL-HLT | 8 |
| 2021 | A Comprehensive Analysis of PMI-based Models for Measuring Semantic Differences
Taichi Aida, Mamoru Komachi, Toshinobu Ogiso, Hiroya Takamura, Daichi Mochihashi |
PACLIC | 2 |
| 2021 | Can Monolingual Pre-trained Encoder-Decoder Improve NMT for Distant Language Pairs?
Hwichan Kim, Mamoru Komachi |
PACLIC | 2 |
| 2021 | Analyzing Semantic Changes in Japanese Words Using BERT
Kazuma Kobayashi, Taichi Aida, Mamoru Komachi |
PACLIC | 3 |
| 2021 | Using Sub-character Level Information for Neural Machine Translation of Logographic LanguagesabstractLogographic and alphabetic languages (e.g., Chinese vs. English) have different writing systems linguistically. Languages belonging to the same writing system usually exhibit more sharing information, which can be used to facilitate natural language processing tasks such as neural machine translation (NMT). This article takes advantage of the logographic characters in Chinese and Japanese by decomposing them into smaller units, thus more optimally utilizing the information these characters share in the training of NMT systems in both encoding and decoding processes. Experiments show that the proposed method can robustly improve the NMT performance of both “logographic” language pairs (JA–ZH) and “logographic + alphabetic” (JA–EN and ZH–EN) language pairs in both supervised and unsupervised NMT scenarios. Moreover, as the decomposed sequences are usually very long, extra position features for the transformer encoder can help with the modeling of these long sequences. The results also indicate that, theoretically, linguistic features can be manipulated to obtain higher share token rates and further improve the performance of natural language processing systems. Longtu Zhang, Mamoru Komachi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2020 | Generating Diverse Corrections with Local Beam Search for Grammatical Error CorrectionabstractIn this study, we propose a beam search method to obtain diverse outputs in a local sequence transduction task where most of the tokens in the source and target sentences overlap, such as in grammatical error correction (GEC).In GEC, it is advisable to rewrite only the local sequences that must be rewritten while leaving the correct sequences unchanged.However, existing methods of acquiring various outputs focus on revising all tokens of a sentence.Therefore, existing methods may either generate ungrammatical sentences because they force the entire sentence to be changed or produce non-diversified sentences by weakening the constraints to avoid generating ungrammatical sentences.Considering these issues, we propose a method that does not rewrite all the tokens in a text, but only rewrites those parts that need to be diversely corrected.Our beam search method adjusts the search token in the beam according to the probability that the prediction is copied from the source sentence.The experimental results show that our proposed method generates more diverse corrections than existing methods without losing accuracy in the GEC task. Kengo Hotate, Masahiro Kaneko, Mamoru Komachi |
COLING | 3 |
| 2020 | Cross-lingual Transfer Learning for Grammatical Error CorrectionabstractIn this study, we explore cross-lingual transfer learning in grammatical error correction (GEC) tasks.Many languages lack the resources required to train GEC models.Cross-lingual transfer learning from high-resource languages (the source models) is effective for training models of low-resource languages (the target models) for various tasks.However, in GEC tasks, the possibility of transferring grammatical knowledge (e.g., grammatical functions) across languages is not evident.Therefore, we investigate cross-lingual transfer learning methods for GEC.Our results demonstrate that transfer learning from other languages can improve the accuracy of GEC.We also demonstrate that proximity to source languages has a significant impact on the accuracy of correcting certain types of errors. Ikumi Yamashita, Satoru Katsumata, Masahiro Kaneko, Aizhan Imankulova, Mamoru Komachi |
COLING | 5 |
| 2020 | SOME: Reference-less Sub-Metrics Optimized for Manual Evaluations of Grammatical Error CorrectionabstractWe propose a reference-less metric trained on manual evaluations of system outputs for grammatical error correction.Previous studies have shown that reference-less metrics are promising; however, existing metrics are not optimized for manual evaluation of the system output because there is no dataset of system output with manual evaluation.This study manually evaluates the output of grammatical error correction systems to optimize the metrics.Experimental results show that the proposed metric improves the correlation with manual evaluation in both systemand sentence-level meta-evaluation.Our dataset and metric will be made publicly available. Ryoma Yoshimura, Masahiro Kaneko, Tomoyuki Kajiwara, Mamoru Komachi |
COLING | 4 |
| 2020 | Double Attention-based Multimodal Neural Machine Translation with Semantic Image RegionsabstractExisting studies on multimodal neural machine translation (MNMT) have mainly focused on the effect of combining visual and textual modalities to improve translations. However, it has been suggested that the visual modality is only marginally beneficial. Conventional visual attention mechanisms have been used to select the visual features from equally-sized grids generated by convolutional neural networks (CNNs), and may have had modest effects on aligning the visual concepts associated with textual objects, because the grid visual features do not capture semantic information. In contrast, we propose the application of semantic image regions for MNMT by integrating visual and textual features using two individual attention mechanisms (double attention). We conducted experiments on the Multi30k dataset and achieved an improvement of 0.5 and 0.9 BLEU points for English-German and English-French translation tasks, compared with the MNMT with grid visual features. We also demonstrated concrete improvements on translation performance benefited from semantic image regions. Mamoru Komachi, Tomoyuki Kajiwara, Chenhui Chu |
EAMT | 2 |
| 2020 | Automated Essay Scoring System for Nonnative Japanese LearnersabstractIn this study, we created an automated essay scoring (AES) system for nonnative Japanese learners using an essay dataset with annotations for a holistic score and multiple trait scores, including content, organization, and language scores. In particular, we developed AES systems using two different approaches: a feature-based approach and a neural-network-based approach. In the former approach, we used Japanese-specific linguistic features, including character-type features such as “kanji” and “hiragana.” In the latter approach, we used two models: a long short-term memory (LSTM) model (Hochreiter and Schmidhuber, 1997) and a bidirectional encoder representations from transformers (BERT) model (Devlin et al., 2019), which achieved the highest accuracy in various natural language processing tasks in 2018. Overall, the BERT model achieved the best root mean squared error and quadratic weighted kappa scores. In addition, we analyzed the robustness of the outputs of the BERT model. We have released and shared this system to facilitate further research on AES for Japanese as a second language learners. Reo Hirao, Mio Arai, Hiroki Shimanaka, Satoru Katsumata, Mamoru Komachi |
LREC | 5 |
| 2020 | Construction of an Evaluation Corpus for Grammatical Error Correction for Learners of Japanese as a Second LanguageabstractThe NAIST Lang-8 Learner Corpora (Lang-8 corpus) is one of the largest second-language learner corpora. The Lang-8 corpus is suitable as a training dataset for machine translation-based grammatical error correction systems. However, it is not suitable as an evaluation dataset because the corrected sentences sometimes include inappropriate sentences. Therefore, we created and released an evaluation corpus for correcting grammatical errors made by learners of Japanese as a Second Language (JSL). As our corpus has less noise and its annotation scheme reflects the characteristics of the dataset, it is ideal as an evaluation corpus for correcting grammatical errors in sentences written by JSL learners. In addition, we applied neural machine translation (NMT) and statistical machine translation (SMT) techniques to correct the grammar of the JSL learners’ sentences and evaluated their results using our corpus. We also compared the performance of the NMT system with that of the SMT system. Aomi Koyama, Tomoshige Kiyuna, Kenji Kobayashi, Mio Arai, Mamoru Komachi |
LREC | 5 |
| 2020 | Neural Machine Translation from Historical Japanese to Contemporary Japanese Using Diachronically Domain-Adapted Word Embeddings
Masashi Takaku, Tosho Hirasawa, Mamoru Komachi, Kanako Komiya |
PACLIC | 3 |
| 2020 | Filtered Pseudo-parallel Corpus Improves Low-resource Neural Machine TranslationabstractLarge-scale parallel corpora are essential for training high-quality machine translation systems; however, such corpora are not freely available for many language translation pairs. Previously, training data has been augmented by pseudo-parallel corpora obtained by using machine translation models to translate monolingual corpora into the source language. However, in low-resource language pairs, in which only low-accurate machine translation systems can be used, translation quality degrades when a pseudo-parallel corpus is naively used. To improve machine translation performance with low-resource language pairs, we propose a method to effectively expand the training data via filtering the pseudo-parallel corpus using quality estimation based on sentence-level round-trip translation. For experiments with three language pairs that utilized small, medium, and large size parallel corpora, BLEU scores significantly improved for low-resource language pairs. Additionally, the effects of iterative bootstrapping on translation performance quality is investigated; resultingly, it is confirmed that bootstrapping can further improve the translation performance. Aizhan Imankulova, Takayuki Sato, Mamoru Komachi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2019 | Debiasing Word Embeddings Improves Multimodal Machine Translation
Tosho Hirasawa, Mamoru Komachi |
MTSummit (1) | 2 |
| 2018 | Construction of a Japanese Word Similarity Dataset
Yuya Sakaizawa, Mamoru Komachi |
LREC | 2 |
| 2018 | Long Short-Term Memory for Japanese Word Segmentation
Yoshiaki Kitagawa, Mamoru Komachi |
PACLIC | 2 |
| 2018 | The Rule of Three: Abstractive Text Summarization in Three Bullet Points
Tomonori Kodaira, Mamoru Komachi |
PACLIC | 2 |
| 2018 | Japanese Sentiment Classification using a Tree-Structured Long Short-Term Memory with Attention
Ryosuke Miyazaki, Mamoru Komachi |
PACLIC | 2 |
| 2017 | MIPA: Mutual Information Based Paraphrase Acquisition via Bilingual PivotingabstractWe present a pointwise mutual information (PMI)-based approach to formalize paraphrasability and propose a variant of PMI, called MIPA, for the paraphrase acquisition. Our paraphrase acquisition method first acquires lexical paraphrase pairs by bilingual pivoting and then reranks them by PMI and distributional similarity. The complementary nature of information from bilingual corpora and from monolingual corpora makes the proposed method robust. Experimental results show that the proposed method substantially outperforms bilingual pivoting and distributional similarity themselves in terms of metrics such as MRR, MAP, coverage, and Spearman’s correlation. Tomoyuki Kajiwara, Mamoru Komachi, Daichi Mochihashi |
IJCNLP(1) | 2 |
| 2017 | Grammatical Error Detection Using Error- and Grammaticality-Specific Word EmbeddingsabstractIn this study, we improve grammatical error detection by learning word embeddings that consider grammaticality and error patterns. Most existing algorithms for learning word embeddings usually model only the syntactic context of words so that classifiers treat erroneous and correct words as similar inputs. We address the problem of contextual information by considering learner errors. Specifically, we propose two models: one model that employs grammatical error patterns and another model that considers grammaticality of the target word. We determine grammaticality of n-gram sequence from the annotated error tags and extract grammatical error patterns for word embeddings from large-scale learner corpora. Experimental results show that a bidirectional long-short term memory model initialized by our word embeddings achieved the state-of-the-art accuracy by a large margin in an English grammatical error detection task on the First Certificate in English dataset. Masahiro Kaneko, Yuya Sakaizawa, Mamoru Komachi |
IJCNLP(1) | 3 |
| 2016 | Building a Monolingual Parallel Corpus for Text Simplification Using Sentence Similarity Based on Alignment between Word EmbeddingsabstractMethods for text simplification using the framework of statistical machine translation have been extensively studied in recent years. However, building the monolingual parallel corpus necessary for training the model requires costly human annotation. Monolingual parallel corpora for text simplification have therefore been built only for a limited number of languages, such as English and Portuguese. To obviate the need for human annotation, we propose an unsupervised method that automatically builds the monolingual parallel corpus for text simplification using sentence similarity based on word embeddings. For any sentence pair comprising a complex sentence and its simple counterpart, we employ a many-to-one method of aligning each word in the complex sentence with the most similar word in the simple sentence and compute sentence similarity by averaging these word similarities. The experimental results demonstrate the excellent performance of the proposed method in a monolingual parallel corpus construction task for English text simplification. The results also demonstrated the superior accuracy in text simplification that use the framework of statistical machine translation trained using the corpus built by the proposed method to that using the existing corpora. Tomoyuki Kajiwara, Mamoru Komachi |
COLING | 2 |
| 2016 | Analysis of English Spelling Errors in a Word-Typing Game
Ryuichi Tachibana, Mamoru Komachi |
LREC | 2 |
| 2015 | Who caught a cold ? - Identifying the subject of a symptomabstractShin Kanouchi, Mamoru Komachi, Naoaki Okazaki, Eiji Aramaki, Hiroshi Ishikawa. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Shin Kanouchi, Mamoru Komachi, Naoaki Okazaki, Eiji Aramaki, Hiroshi Ishikawa 0004 |
ACL (1) | 2 |
| 2015 | Japanese Sentiment Classification with Stacked Denoising Auto-Encoder using Distributed Word Representation
Peinan Zhang, Mamoru Komachi |
PACLIC | 2 |
| 2014 | Extracting a Chinese Learner Corpus from the Web: Grammatical Error Correction for Learning Chinese as a Foreign Language with Statistical Machine Translation
Yinchen Zhao, Mamoru Komachi, Hiroshi Ishikawa 0004 |
ICCE | 2 |
| 2013 | Towards Automatic Error Type Classification of Japanese Language Learners' Writings
Hiromi Oyama, Mamoru Komachi, Yuji Matsumoto 0001 |
PACLIC | 2 |
| 2012 | Joint English Spelling Error Correction and POS Tagging for Language Learners Writing
Keisuke Sakaguchi, Tomoya Mizumoto, Mamoru Komachi, Yuji Matsumoto 0001 |
COLING | 3 |
| 2012 | UniDic for Early Middle Japanese: a Dictionary for Morphological Analysis of Classical Japanese
Toshinobu Ogiso, Mamoru Komachi, Yasuharu Den, Yuji Matsumoto 0001 |
LREC | 2 |
| 2011 | Using the Mutual k-Nearest Neighbor Graphs for Semi-supervised Classification on Natural Language Data
Kohei Ozaki, Masashi Shimbo, Mamoru Komachi, Yuji Matsumoto 0001 |
CoNLL | 3 |
| 2011 | Japanese Predicate Argument Structure Analysis Exploiting Argument Position and Type
Yuta Hayashibe, Mamoru Komachi, Yuji Matsumoto 0001 |
IJCNLP | 2 |
| 2011 | Mining Revision Log of Language Learning SNS for Automated Japanese Error Correction of Second Language Learners
Tomoya Mizumoto, Mamoru Komachi, Masaaki Nagata, Yuji Matsumoto 0001 |
IJCNLP | 2 |
| 2011 | Automatic Labeling of Voiced Consonants for Morphological Analysis of Modern Japanese Literature
Teruaki Oka, Mamoru Komachi, Toshinobu Ogiso, Yuji Matsumoto 0001 |
IJCNLP | 2 |
| 2011 | Japanese Abbreviation Expansion with Query and Clickthrough Logs
Kei Uchiumi, Mamoru Komachi, Keigo Machinaga, Toshiyuki Maezawa, Toshinori Satou, Yoshinori Kobayashi |
IJCNLP | 2 |
| 2008 | Graph-based Analysis of Semantic Drift in Espresso-like Bootstrapping Algorithms
Mamoru Komachi, Taku Kudo, Masashi Shimbo, Yuji Matsumoto 0001 |
EMNLP | 1 |
| 2008 | Minimally Supervised Learning of Semantic Knowledge from Query Logs
Mamoru Komachi, Hisami Suzuki |
IJCNLP | 1 |