EDBT 2026 Demo / reviewers in the wild / expert
Masato Mita
dblp:213/1183
· DBLP profile ↗
19ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0001-6210-3716ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 3 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Language Acquisition Device in Large Language ModelsabstractLarge Language Models (LLMs) remain substantially less data-efficient than humans.Prepretraining (PPT) on synthetic languages has been proposed to close this gap, with prior work emphasizing highly expressive formal languages such as k-Shuffle Dyck.Inspired by the Language Acquisition Device (LAD) hypothesis, which posits that innate constraints preemptively restrict the learner's hypothesis space to natural-language-like structure, we propose LAD-inspired PPT: pre-pretraining on MP-STRUCT, a formal language whose strings encode hierarchical composition, feature-based dependencies, and long-distance displacement via MERGE, AGREE, and MOVE.A brief 500step PPT with MP-STRUCT matches strong formal-language baselines in token efficiency while additionally imparting a human-like resistance to structurally implausible languages.Analyzing simplified variants, we find that MP-STRUCT CORE outperforms k-Shuffle Dyck despite not being definable in C-RASP (a formal bound on transformer expressivity), challenging the prior hypothesis that effective PPT languages must be both hierarchically expressive and circuit-theoretically learnable.We show that functional landmarks, which reduce dependency resolution ambiguity, are a key driver, suggesting that effective PPT design depends not only on expressivity but also on the accessibility of dependency resolution. Masato Mita, Taiga Someya, Ryo Yoshida, Yohei Oseki |
ACL (1) | 1 |
| 2025 | Targeted Syntactic Evaluation for Grammatical Error CorrectionabstractLanguage learners encounter a wide range of grammar items across the beginner, intermediate, and advanced levels.To develop grammatical error correction (GEC) models effectively, it is crucial to identify which grammar items are easier or more challenging for models to correct.However, conventional benchmarks based on learner-produced texts are insufficient for conducting detailed evaluations of GEC model performance across a wide range of grammar items due to biases in their distribution.To address this issue, we propose a new evaluation paradigm that assesses GEC models using minimal pairs of ungrammatical and grammatical sentences for each grammar item.As the first benchmark within this paradigm, we introduce the CEFR-based Targeted Syntactic Evaluation Dataset for Grammatical Error Correction (CTSEG), which complements existing English benchmarks by enabling fine-grained analyses previously unattainable with conventional datasets.Using CTSEG, we evaluate three mainstream types of English GEC models: sequence-to-sequence models, sequence tagging models, and prompt-based models.The results indicate that while current models perform well on beginner-level grammar items, their performance deteriorates substantially for intermediate and advanced items.Dataset Sents.Refs.Error Tags CEFR Level FCE (Yannakoudakis et al., 2011) 2,695 1 71 B1-B2 KJ (Nagata et al., 2011) 3,199 1 22 A1-A2?CoNLL-2014 (Ng et al., 2014) 1,312 2 28 C1 AESW (Daudaravicius et al., 2016) 143,804 1 N/A C1-C2 (+Native) JFLEG (Napoles et al., 2017) 747 4 N/A A1-C2?BEA-2019 (Bryant et al., 2019) 4,384 5 25 A1-C2 (+Native) GMEG (Napoles et al., 2019) 2,960 4 N/A B1-B2 (+Native) CWEB (Flachs et al., 2020) Aomi Koyama, Masato Mita, Su-Youn Yoon, Yasufumi Takama, Mamoru Komachi |
ACL (1) | 2 |
| 2025 | Developmentally-plausible Working Memory Shapes a Critical Period for Language AcquisitionabstractLarge language models possess general linguistic abilities but acquire language less efficiently than humans. This study proposes a method for integrating the developmental characteristics of working memory during the critical period, a stage when human language acquisition is particularly efficient, into the training process of language models. The proposed method introduces a mechanism that initially constrains working memory during the early stages of training and gradually relaxes this constraint in an exponential manner as learning progresses. Targeted syntactic evaluation shows that the proposed method outperforms conventional methods without memory constraints or with static memory constraints. These findings not only provide new directions for designing data-efficient language models but also offer indirect evidence supporting the role of the developmental characteristics of working memory as the underlying mechanism of the critical period in language acquisition. Masato Mita, Ryo Yoshida, Yohei Oseki |
ACL (1) | 1 |
| 2025 | AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine AdvertisingabstractPeinan Zhang, Yusuke Sakai, Masato Mita, Hiroki Ouchi, Taro Watanabe. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Peinan Zhang, Yusuke Sakai 0010, Masato Mita, Hiroki Ouchi, Taro Watanabe |
NAACL (Long Papers) | 3 |
| 2024 | Striking Gold in Advertising: Standardization and Exploration of Ad Text GenerationabstractIn response to the limitations of manual ad creation, significant research has been conducted in the field of automatic ad text generation (ATG).However, the lack of comprehensive benchmarks and well-defined problem sets has made comparing different methods challenging.To tackle these challenges, we standardize the task of ATG and propose a first benchmark dataset, CAMERA , carefully designed and enabling the utilization of multi-modal information and facilitating industry-wise evaluations.Our extensive experiments with a variety of nine baselines, from classical methods to state-of-the-art models including large language models (LLMs), show the current state and the remaining challenges.We also explore how existing metrics in ATG and an LLMbased evaluator align with human evaluations.ORIX Card Loan Keyword … Diagnosis of instant loan Cards! 3 recommended companies to borrow … Landing page (LP) Ad text 1. [Official] Top 3 Popular Card Loans 2. Easily diagnose recommended card loans 3. Diagnose Cards Availbale for Same-Day Borrowing ! 4. Get Financing in as Fast as 30 Mnutes Online ! Masato Mita, Soichiro Murakami, Akihiko Kato, Peinan Zhang |
ACL (1) | 1 |
| 2024 | CAMERA³: An Evaluation Dataset for Controllable Ad Text Generation in JapaneseabstractAd text generation is the task of creating compelling text from an advertising asset that describes products or services, such as a landing page. In advertising, diversity plays an important role in enhancing the effectiveness of an ad text, mitigating a phenomenon called “ad fatigue,” where users become disengaged due to repetitive exposure to the same advertisement. Despite numerous efforts in ad text generation, the aspect of diversifying ad texts has received limited attention, particularly in non-English languages like Japanese. To address this, we present CAMERA³, an evaluation dataset for controllable text generation in the advertising domain in Japanese. Our dataset includes 3,980 ad texts written by expert annotators, taking into account various aspects of ad appeals. We make CAMERA³ publicly available, allowing researchers to examine the capabilities of recent NLG models in controllable text generation in a real-world scenario. Go Inoue, Akihiko Kato, Masato Mita, Ukyo Honda, Peinan Zhang |
LREC/COLING | 3 |
| 2024 | Token-length Bias in Minimal-pair Paradigm DatasetsabstractMinimal-pair paradigm datasets have been used as benchmarks to evaluate the linguistic knowledge of models and provide an unsupervised method of acceptability judgment. The model performances are evaluated based on the percentage of minimal pairs in the MPP dataset where the model assigns a higher sentence log-likelihood to an acceptable sentence than to an unacceptable sentence. Each minimal pair in MPP datasets is controlled to align the number of words per sentence because the sentence length affects the sentence log-likelihood. However, aligning the number of words may be insufficient because recent language models tokenize sentences with subwords. Tokenization may cause a token length difference in minimal pairs, introducing token-length bias that skews the evaluation results. This study demonstrates that MPP datasets suffer from token-length bias and fail to evaluate the linguistic knowledge of a language model correctly. The results proved that sentences with a shorter token length would likely be assigned a higher log-likelihood regardless of their acceptability, which becomes problematic when comparing models with different tokenizers. To address this issue, we propose a debiased minimal pair generation method, allowing MPP datasets to measure language ability correctly and provide comparable results for all models. Naoya Ueda, Masato Mita, Teruaki Oka, Mamoru Komachi |
LREC/COLING | 2 |
| 2024 | DejaVu: Disambiguation evaluation dataset for English-JApanese machine translation on VisUal information
Ayako Sato, Tosho Hirasawa, Hwichan Kim, Zhousi Chen, Teruaki Oka, Masato Mita, Mamoru Komachi |
PACLIC | 6 |
| 2024 | Not Eliminate but Aggregate: Post-Hoc Control over Mixture-of-Experts to Address Shortcut Shifts in Natural Language UnderstandingabstractAbstract Recent models for natural language understanding are inclined to exploit simple patterns in datasets, commonly known as shortcuts. These shortcuts hinge on spurious correlations between labels and latent features existing in the training data. At inference time, shortcut-dependent models are likely to generate erroneous predictions under distribution shifts, particularly when some latent features are no longer correlated with the labels. To avoid this, previous studies have trained models to eliminate the reliance on shortcuts. In this study, we explore a different direction: pessimistically aggregating the predictions of a mixture-of-experts, assuming each expert captures relatively different latent features. The experimental results demonstrate that our post-hoc control over the experts significantly enhances the model’s robustness to the distribution shift in shortcuts. Additionally, we show that our approach has some practical advantages. We also analyze our model and provide results to support the assumption.1 Ukyo Honda, Tatsushi Oka, Peinan Zhang, Masato Mita |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | Revisiting Meta-evaluation for Grammatical Error CorrectionabstractAbstract Metrics are the foundation for automatic evaluation in grammatical error correction (GEC), with their evaluation of the metrics (meta-evaluation) relying on their correlation with human judgments. However, conventional meta-evaluations in English GEC encounter several challenges, including biases caused by inconsistencies in evaluation granularity and an outdated setup using classical systems. These problems can lead to misinterpretation of metrics and potentially hinder the applicability of GEC techniques. To address these issues, this paper proposes SEEDA, a new dataset for GEC meta-evaluation. SEEDA consists of corrections with human ratings along two different granularities: edit-based and sentence-based, covering 12 state-of-the-art systems including large language models, and two human corrections with different focuses. The results of improved correlations by aligning the granularity in the sentence-level meta-evaluation suggest that edit-based metrics may have been underestimated in existing studies. Furthermore, correlations of most metrics decrease when changing from classical to neural systems, indicating that traditional metrics are relatively poor at evaluating fluently corrected sentences with many edits. Masamune Kobayashi, Masato Mita, Mamoru Komachi |
Trans. Assoc. Comput. Linguistics | 2 |
| 2023 | Chinese Grammatical Error Correction Using Pre-trained Models and Pseudo DataabstractIn recent studies, pre-trained models and pseudo data have been key factors in improving the performance of the English grammatical error correction (GEC) task. However, few studies have examined the role of pre-trained models and pseudo data in the Chinese GEC task. Therefore, we develop Chinese GEC models based on three pre-trained models: Chinese BERT, Chinese T5, and Chinese BART, and then incorporate these models with pseudo data to determine the best configuration for the Chinese GEC task. On the natural language processing and Chinese computing (NLPCC) 2018 GEC shared task test set, all our single models outperform the ensemble models developed by the top team of the shared task. Chinese BART achieves an F score of 37.15, which is a state-of-the-art result. We then combine our Chinese GEC models with three kinds of pseudo data: Lang-8 (MaskGEC), Wiki (MaskGEC), and Wiki (Backtranslation). We find that most models can benefit from pseudo data, and BART+Lang-8 (MaskGEC) is the ideal setting in terms of accuracy and training efficiency. The experimental results demonstrate the effectiveness of the pre-trained models and pseudo data on the Chinese GEC task and provide an easily reproducible and adaptable baseline for future works. Finally, we annotate the error types of the development data; the results show that word-level errors dominate all error types, and word selection errors must be addressed even when using pre-trained models and pseudo data. Our codes are available at https://github.com/wang136906578/BERT-encoder-ChineseGEC . Michiki Kurosawa, Satoru Katsumata, Masato Mita, Mamoru Komachi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2022 | Construction of a Quality Estimation Dataset for Automatic Evaluation of Japanese Grammatical Error CorrectionabstractIn grammatical error correction (GEC), automatic evaluation is considered as an important factor for research and development of GEC systems. Previous studies on automatic evaluation have shown that quality estimation models built from datasets with manual evaluation can achieve high performance in automatic evaluation of English GEC. However, quality estimation models have not yet been studied in Japanese, because there are no datasets for constructing quality estimation models. In this study, therefore, we created a quality estimation dataset with manual evaluation to build an automatic evaluation model for Japanese GEC. By building a quality estimation model using this dataset and conducting a meta-evaluation, we verified the usefulness of the quality estimation model for Japanese GEC. Daisuke Suzuki, Yujin Takahashi, Ikumi Yamashita, Taichi Aida, Tosho Hirasawa, Michitaka Nakatsuji, Masato Mita, Mamoru Komachi |
LREC | 7 |
| 2022 | ProQE: Proficiency-wise Quality Estimation dataset for Grammatical Error CorrectionabstractThis study investigates how supervised quality estimation (QE) models of grammatical error correction (GEC) are affected by the learners’ proficiency with the data. QE models for GEC evaluations in prior work have obtained a high correlation with manual evaluations. However, when functioning in a real-world context, the data used for the reported results have limitations because prior works were biased toward data by learners with relatively high proficiency levels. To address this issue, we created a QE dataset that includes multiple proficiency levels and explored the necessity of performing proficiency-wise evaluation for QE of GEC. Our experiments demonstrated that differences in evaluation dataset proficiency affect the performance of QE models, and proficiency-wise evaluation helps create more robust models. Yujin Takahashi, Masahiro Kaneko, Masato Mita, Mamoru Komachi |
LREC | 3 |
| 2021 | Shared Task on Feedback Comment Generation for Language LearnersabstractIn this paper, we propose a generation challenge called Feedback comment generation for language learners.It is a task where given a text and a span, a system generates, for the span, an explanatory note that helps the writer (language learner) improve their writing skills.The motivations for this challenge are: (i) practically, it will be beneficial for both language learners and teachers if a computerassisted language learning system can provide feedback comments just as human teachers do; (ii) theoretically, feedback comment generation for language learners has a mixed aspect of other generation tasks together with its unique features and it will be interesting to explore what kind of generation method is effective against what kind of writing rule.To this end, we have created a dataset and developed baseline systems to estimate baseline performance.With these preparations, we propose a generation challenge of feedback comment generation. Ryo Nagata, Masato Hagiwara, Kazuaki Hanawa, Masato Mita, Artem Chernodub, Olena Nahorna |
INLG | 4 |
| 2020 | Encoder-Decoder Models Can Benefit from Pre-trained Masked Language Models in Grammatical Error CorrectionabstractThis paper investigates how to effectively incorporate a pre-trained masked language model (MLM), such as BERT, into an encoderdecoder (EncDec) model for grammatical error correction (GEC).The answer to this question is not as straightforward as one might expect because the previous common methods for incorporating a MLM into an EncDec model have potential drawbacks when applied to GEC.For example, the distribution of the inputs to a GEC model can be considerably different (erroneous, clumsy, etc.) from that of the corpora used for pre-training MLMs; however, this issue is not addressed in the previous methods.Our experiments show that our proposed method, where we first fine-tune a MLM with a given GEC corpus and then use the output of the finetuned MLM as additional features in the GEC model, maximizes the benefit of the MLM.The best-performing model achieves state-ofthe-art performances on the BEA-2019 and CoNLL-2014 benchmarks.Our code is publicly available at: https://github.com/ kanekomasahiro/bert-gec. Masahiro Kaneko, Masato Mita, Shun Kiyono, Jun Suzuki 0001, Kentaro Inui |
ACL | 2 |
| 2020 | PheMT: A Phenomenon-wise Dataset for Machine Translation Robustness on User-Generated ContentsabstractNeural Machine Translation (NMT) has shown drastic improvement in its quality when translating clean input, such as text from the news domain.However, existing studies suggest that NMT still struggles with certain kinds of input with considerable noise, such as User-Generated Contents (UGC) on the Internet.To make better use of NMT for cross-cultural communication, one of the most promising directions is to develop a model that correctly handles these expressions.Though its importance has been recognized, it is still not clear as to what creates the great gap in performance between the translation of clean input and that of UGC.To answer the question, we present a new dataset, PheMT, for evaluating the robustness of MT systems against specific linguistic phenomena in Japanese-English translation.Our experiments with the created dataset revealed that not only our in-house models but even widely used off-the-shelf systems are greatly disturbed by the presence of certain phenomena. Ryo Fujii, Masato Mita, Kaori Abe, Kazuaki Hanawa, Makoto Morishita, Jun Suzuki 0001, Kentaro Inui |
COLING | 2 |
| 2020 | Taking the Correction Difficulty into Account in Grammatical Error Correction EvaluationabstractThis paper presents performance measures for grammatical error correction which take into account the difficulty of error correction.To the best of our knowledge, no conventional measure has such functionality despite the fact that some errors are easy to correct and others are not.The main purpose of this work is to provide a way of determining the difficulty of error correction and to motivate researchers in the domain to attack such difficult errors.The performance measures are based on the simple idea that the more systems successfully correct an error, the easier it is considered to be.This paper presents a set of algorithms to implement this idea.It evaluates the performance measures quantitatively and qualitatively on a wide variety of corpora and systems, revealing that they agree with our intuition of correction difficulty.A scorer and difficulty weight data based on the algorithms have been made available on the web. Takumi Gotou, Ryo Nagata, Masato Mita, Kazuaki Hanawa |
COLING | 3 |
| 2020 | GitHub Typo Corpus: A Large-Scale Multilingual Dataset of Misspellings and Grammatical ErrorsabstractThe lack of large-scale datasets has been a major hindrance to the development of NLP tasks such as spelling correction and grammatical error correction (GEC). As a complementary new resource for these tasks, we present the GitHub Typo Corpus, a large-scale, multilingual dataset of misspellings and grammatical errors along with their corrections harvested from GitHub, a large and popular platform for hosting and sharing git repositories. The dataset, which we have made publicly available, contains more than 350k edits and 65M characters in more than 15 languages, making it the largest dataset of misspellings to date. We also describe our process for filtering true typo edits based on learned classifiers on a small annotated subset, and demonstrate that typo edits can be identified with F1 0.9 using a very simple classifier with only three features. The detailed analyses of the dataset show that existing spelling correctors merely achieve an F-measure of approx. 0.5, suggesting that the dataset serves as a new, rich source of spelling errors that complement existing datasets. Masato Hagiwara, Masato Mita |
LREC | 2 |
| 2019 | An Empirical Study of Incorporating Pseudo Data into Grammatical Error CorrectionabstractShun Kiyono, Jun Suzuki, Masato Mita, Tomoya Mizumoto, Kentaro Inui. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shun Kiyono, Jun Suzuki 0001, Masato Mita, Tomoya Mizumoto, Kentaro Inui |
EMNLP/IJCNLP (1) | 3 |