VLDB 2026 Research / reviewers in the wild / expert
Tomoyuki Kajiwara
dblp:140/3305
· DBLP profile ↗
33ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0002-3233-4879ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 6 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Japanese Dataset for Aspect-based Sentiment Polarity Classification and Emotion Intensity Estimation
Kentaro Hanafusa, Kota Manabe, Yuki Maeda, Daisuke Maekawa, Tomoyuki Kajiwara, Hideaki Hayashi, Yuta Nakashima, Hajime Nagahara |
LREC | 5 |
| 2026 | Parallel Corpus Filtering Based on Semantic Similarity and Surface Dissimilarity for Japanese Text Simplification with LLMs
Daisuke Maekawa, Tomoyuki Kajiwara, Takashi Ninomiya |
LREC | 2 |
| 2026 | HOTATE: A Japanese Dialogue Corpus Annotated with Responses of Private Thoughts and Public Statements
Yuko Toda, Daisuke Maekawa, Kota Manabe, Eito Yoneyama, Kanade Nonomura, Yuki Fujiwara, Tomoyuki Kajiwara |
LREC | 7 |
| 2026 | A Dialogue System for Second Language Learning with Dynamic Adaptation to In-Dialogue Difficulty ChangesabstractWe propose a dialogue system for second language learning that dynamically adjusts the difficulty of its utterances by adapting to changes in the learner’s utterance difficulty across dialogue topics. Recent studies have advanced second language learning dialogue systems based on large language models (LLMs) grounded in the Input Hypothesis. However, although learners’ language proficiency varies based on their familiarity with a given topic, conventional fixed-difficulty dialogue systems do not adequately account for this variability. In this study, we propose a dialogue system for second language learning that minimizes the difficulty gap between the system and a user model. The user model dynamically varies utterance difficulty based on the topic during dialogue. Our LLM-based dialogue system controls utterance difficulty by prompt engineering and Direct Preference Optimization (DPO). Experimental results in both Japanese and English demonstrate that the proposed method is effective in tracking utterance difficulty for second language learners. Taku Morioka, Junya Takayama, Tomoyuki Kajiwara |
SIGDIAL | 3 |
| 2026 | MultiLexNorm++: A Unified Benchmark and a Generative Model for Lexical Normalization for Asian LanguagesabstractSocial media data has been of interest to Natural Language Processing (NLP) practitioners for over a decade, because of its richness in information, but also challenges for automatic processing. Since language use is more informal, spontaneous, and adheres to many different sociolects, the performance of NLP models often deteriorates. One solution to this problem is to transform data to a standard variant before processing it, which is also called lexical normalization. There has been a wide variety of benchmarks and models proposed for this task. The MultiLexNorm benchmark proposed to unify these efforts, but it consists almost solely of languages from the Indo-European language family in the Latin script. Hence, we propose an extension to MultiLexNorm, which covers five Asian languages from different language families in four different scripts. We show that the previous state-of-the-art model performs worse on the new languages and propose a new architecture based on Large Language Models (LLMs), which shows more robust performance. Finally, we analyze remaining errors, revealing future directions for this task. Weerayut Buaphet, Thanh-Nhi Nguyen, Risa Kondo, Tomoyuki Kajiwara, Yumin Kim, Jimin Lee 0001, Hwanhee Lee, Holy Lovenia, Peerat Limkonchotiwat, Sarana Nutanong, Rob van der Goot |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2024 | Transfer Fine-tuning for Quality Estimation of Text SimplificationabstractTo efficiently train quality estimation of text simplification on a small-scale labeled corpus, we train sentence difficulty estimation prior to fine-tuning the pre-trained language models. Our proposed method improves the quality estimation of text simplification in the framework of transfer fine-tuning, in which pre-trained language models can improve the performance of the target task by additional training on the relevant task prior to fine-tuning. Since the labeled corpus for quality estimation of text simplification is small (600 sentence pairs), an efficient training method is desired. Therefore, we propose a training method for pseudo quality estimation that does not require labels for quality estimation. As a relevant task for quality estimation of text simplification, we train the estimation of sentence difficulty. This is a binary classification task that identifies which sentence is simpler using an existing parallel corpus for text simplification. Experimental results on quality estimation of English text simplification showed that not only the quality estimation performance on simplicity that was trained, but also the quality estimation performance on fluency and meaning preservation could be improved in some cases. Yuki Hironaka, Tomoyuki Kajiwara, Takashi Ninomiya |
LREC/COLING | 2 |
| 2024 | Utilizing Longer Context than Speech Bubbles in Automated Manga TranslationabstractThis paper focuses on improving the performance of machine translation for manga (Japanese-style comics). In manga machine translation, text consists of a sequence of speech bubbles and each speech bubble is translated individually. However, each speech bubble itself does not contain sufficient information for translation. Therefore, previous work has proposed methods to use contextual information, such as the previous speech bubble, speech bubbles within the same scene, and corresponding scene images. In this research, we propose two new approaches to capture broader contextual information. Our first approach involves scene-based translation that considers the previous scene. The second approach considers broader context information, including details about the work, author, and manga genre. Through our experiments, we confirm that each of our methods improves translation quality, with the combination of both methods achieving the highest quality. Additionally, detailed analysis reveals the effect of zero-anaphora resolution in translation, such as supplying missing subjects not mentioned within a scene, highlighting the usefulness of longer contextual information in manga machine translation. Hiroto Kaino, Soichiro Sugihara, Tomoyuki Kajiwara, Takashi Ninomiya, Joshua B. Tanner, Shonosuke Ishiwatari |
LREC/COLING | 3 |
| 2024 | Controllable Paraphrase Generation for Semantic and Lexical SimilaritiesabstractWe developed a controllable paraphrase generation model for semantic and lexical similarities using a simple and intuitive mechanism: attaching tags to specify these values at the head of the input sentence. Lexically diverse paraphrases have been long coveted for data augmentation. However, their generation is not straightforward because diversifying surfaces easily degrades semantic similarity. Furthermore, our experiments revealed two critical features in data augmentation by paraphrasing: appropriate similarities of paraphrases are highly downstream task-dependent, and mixing paraphrases of various similarities negatively affects the downstream tasks. These features indicated that the controllability in paraphrase generation is crucial for successful data augmentation. We tackled these challenges by fine-tuning a pre-trained sequence-to-sequence model employing tags that indicate the semantic and lexical similarities of synthetic paraphrases selected carefully based on the similarities. The resultant model could paraphrase an input sentence according to the tags specified. Extensive experiments on data augmentation for contrastive learning and pre-fine-tuning of pretrained masked language models confirmed the effectiveness of the proposed model. We release our paraphrase generation model and a corpus of 87 million diverse paraphrases. (https://github.com/Ogamon958/ConPGS) Yuya Ogasa, Tomoyuki Kajiwara, Yuki Arase |
LREC/COLING | 2 |
| 2024 | Automatic Decomposition of Text Editing Examples into Primitive Edit Operations: Toward Analytic Evaluation of Editing SystemsabstractThis paper presents our work on a task of automatic decomposition of text editing examples into primitive edit operations. Toward a detailed analysis of the behavior of text editing systems, identification of fine-grained edit operations performed by the systems is essential. Given a pair of source and edited sentences, the goal of our task is to generate a non-redundant sequence of primitive edit operations, i.e., the semantically minimal edit operations preserving grammaticality, that iteratively converts the source sentence to the edited sentence. First, we formalize this task, explaining its significant features and specifying the constraints that primitive edit operations should satisfy. Then, we propose a method to automate this task, which consists of two steps: generation of an edit operation lattice and selection of an optimal path. To obtain a wide range of edit operation candidates in the first step, we combine a phrase aligner and a large language model. Experimental results show that our method perfectly decomposes 44% and 64% of editing examples in the text simplification and machine translation post-editing datasets, respectively. Detailed analyses also provide insights into the difficulties of this task, suggesting directions for improvement. Daichi Yamaguchi, Rei Miyata, Atsushi Fujita, Tomoyuki Kajiwara, Satoshi Sato |
LREC/COLING | 4 |
| 2023 | Self-Ensemble of N-best Generation Hypotheses by Lexically Constrained DecodingabstractWe propose a method that ensembles N -best hypotheses to improve natural language generation.Previous studies have achieved notable improvements in generation quality by explicitly reranking N -best candidates.These studies assume that there exists a hypothesis of higher quality.We expand the assumption to be more practical as there exist partly higher quality hypotheses in the N -best yet they may be imperfect as the entire sentences.By merging these high-quality fragments, we can obtain a higher-quality output than the single-best sentence.Specifically, we first obtain N -best hypotheses and conduct token-level quality estimation.We then apply tokens that should or should not be present in the final output as lexical constraints in decoding.Empirical experiments on paraphrase generation, summarisation, and constrained text generation confirm that our method outperforms the strong N -best reranking methods. Ryota Miyano, Tomoyuki Kajiwara, Yuki Arase |
EMNLP | 2 |
| 2022 | Adversarial Training on Disentangling Meaning and Language Representations for Unsupervised Quality EstimationabstractWe propose a method to distill language-agnostic meaning embeddings from multilingual sentence encoders for unsupervised quality estimation of machine translation. Our method facilitates that the meaning embeddings focus on semantics by adversarial training that attempts to eliminate language-specific information. Experimental results on unsupervised quality estimation reveal that our method achieved higher correlations with human evaluations. Yuto Kuroda, Tomoyuki Kajiwara, Yuki Arase, Takashi Ninomiya |
COLING | 2 |
| 2022 | CEFR-Based Sentence Difficulty Annotation and AssessmentabstractControllable text simplification is a crucial assistive technique for language learning and teaching.One of the primary factors hindering its advancement is the lack of a corpus annotated with sentence difficulty levels based on language ability descriptions.To address this problem, we created the CEFR-based Sentence Profile (CEFR-SP) corpus, containing 17k English sentences annotated with the levels based on the Common European Framework of Reference for Languages assigned by English-education professionals.In addition, we propose a sentence-level assessment model to handle unbalanced level distribution because the most basic and highly proficient sentences are naturally scarce.In the experiments in this study, our method achieved a macro-F1 score of 84.5% in the level assessment, thus outperforming strong baselines employed in readability assessment. Yuki Arase, Satoru Uchida, Tomoyuki Kajiwara |
EMNLP | 3 |
| 2022 | JADE: Corpus for Japanese Definition ModellingabstractThis study investigated and released the JADE, a corpus for Japanese definition modelling, which is a technique that automatically generates definitions of a given target word and phrase. It is a crucial technique for practical applications that assist language learning and education, as well as for those supporting reading documents in unfamiliar domains. Although corpora for development of definition modelling techniques have been actively created, their languages are mostly limited to English. In this study, a corpus for Japanese, named JADE, was created following the previous study that mines an online encyclopedia. The JADE provides about 630k sets of targets, their definitions, and usage examples as contexts for about 41k unique targets, which is sufficiently large to train neural models. The targets are both words and phrases, and the coverage of domains and topics is diverse. The performance of a pre-trained sequence-to-sequence model and the state-of-the-art definition modelling method was also benchmarked on JADE for future development of the technique in Japanese. The JADE corpus has been released and available online. Tomoyuki Kajiwara, Yuki Arase |
LREC | 2 |
| 2022 | A Japanese Dataset for Subjective and Objective Sentiment Polarity Classification in Micro Blog DomainabstractWe annotate 35,000 SNS posts with both the writer’s subjective sentiment polarity labels and the reader’s objective ones to construct a Japanese sentiment analysis dataset. Our dataset includes intensity labels (none, weak, medium, and strong) for each of the eight basic emotions by Plutchik (joy, sadness, anticipation, surprise, anger, fear, disgust, and trust) as well as sentiment polarity labels (strong positive, positive, neutral, negative, and strong negative). Previous studies on emotion analysis have studied the analysis of basic emotions and sentiment polarity independently. In other words, there are few corpora that are annotated with both basic emotions and sentiment polarity. Our dataset is the first large-scale corpus to annotate both of these emotion labels, and from both the writer’s and reader’s perspectives. In this paper, we analyze the relationship between basic emotion intensity and sentiment polarity on our dataset and report the results of benchmarking sentiment polarity classification. Haruya Suzuki, Yuto Miyauchi, Kazuki Akiyama, Tomoyuki Kajiwara, Takashi Ninomiya, Noriko Takemura, Yuta Nakashima, Hajime Nagahara |
LREC | 4 |
| 2022 | A Benchmark Dataset for Multi-Level Complexity-Controllable Machine TranslationabstractThis paper presents a new benchmark test dataset for multi-level complexity-controllable machine translation (MLCC-MT), which is MT controlling the complexity of the output at more than two levels. In previous research, MLCC-MT models have been evaluated on a test dataset automatically constructed from the Newsela corpus, which is a document-level comparable corpus with document-level complexity. The existing test dataset has the following three problems: (i) A source language sentence and its target language sentence are not necessarily an exact translation pair because they are automatically detected. (ii) A target language sentence and its simplified target language sentence are not necessarily exactly parallel because they are automatically aligned. (iii) A sentence-level complexity is not necessarily appropriate because it is transferred from an article-level complexity attached to the Newsela corpus. Therefore, we create a benchmark test dataset for Japanese-to-English MLCC-MT from the Newsela corpus by introducing an automatic filtering of data with inappropriate sentence-level complexity, manual check for parallel target language sentences with different complexity levels, and manual translation. Moreover, we implement two MLCC-NMT frameworks with a Transformer architecture and report their performance on our test dataset as baselines for future research. Our test dataset and codes are released. Kazuki Tani, Ryoya Yuasa, Kazuki Takikawa, Akihiro Tamura, Tomoyuki Kajiwara, Takashi Ninomiya, Tsuneo Kato |
LREC | 5 |
| 2022 | Region-attentive multimodal neural machine translationabstractWe propose a multimodal neural machine translation (MNMT) method with semantic image regions called region-attentive multimodal neural machine translation (RA-NMT). Existing studies on MNMT have mainly focused on employing global visual features or equally sized grid local visual features extracted by convolutional neural networks (CNNs) to improve translation performance. However, they neglect the effect of semantic information captured inside the visual features. This study utilizes semantic image regions extracted by object detection for MNMT and integrates visual and textual features using two modality-dependent attention mechanisms. The proposed method was implemented and verified on two neural architectures of neural machine translation (NMT): recurrent neural network (RNN) and self-attention network (SAN). Experimental results on different language pairs of Multi30k dataset show that our proposed method improves over baselines and outperforms most of the state-of-the-art MNMT methods. Further analysis demonstrates that the proposed method can achieve better translation performance because of its better visual feature use. Mamoru Komachi, Tomoyuki Kajiwara, Chenhui Chu |
Neurocomputing | 3 |
| 2022 | Word-Region Alignment-Guided Multimodal Neural Machine TranslationabstractWe propose word-region alignment-guided multimodal neural machine translation (MNMT), a novel model for MNMT that links the semantic correlation between textual and visual modalities using word-region alignment (WRA). Existing studies on MNMT have mainly focused on the effect of integrating visual and textual modalities. However, they do not leverage the semantic relevance between the two modalities. We advance the semantic correlation between textual and visual modalities in MNMT by incorporating WRA as a bridge. This proposal has been implemented on two mainstream architectures of neural machine translation (NMT): the recurrent neural network (RNN) and the transformer. Experiments on two public benchmarks, English–German and English–French translation tasks using the Multi30k dataset and English–Japanese translation tasks using the Flickr30kEnt-JP dataset prove that our model has a significant improvement with respect to the competitive baselines across different evaluation metrics and outperforms most of the existing MNMT models. For example, 1.0 BLEU scores are improved for the English–German task and 1.1 BLEU scores are improved for the English–French task on the Multi30k test2016 set; and 0.7 BLEU scores are improved for the English–Japanese task on the Flickr30kEnt-JP test set. Further analysis demonstrates that our model can achieve better translation performance by integrating WRA, leading to better visual information use. Mamoru Komachi, Tomoyuki Kajiwara, Chenhui Chu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Definition Modelling for Appropriate SpecificityabstractDefinition generation techniques aim to generate a definition of a target word or phrase given a context. In previous studies, researchers have faced various issues such as the out-of-vocabulary problem and over/under-specificity problems. Over-specific definitions present narrow word meanings, whereas under-specific definitions present general and context-insensitive meanings. Herein, we propose a method for definition generation with appropriate specificity. The proposed method addresses the aforementioned problems by leveraging a pre-trained encoder-decoder model, namely Text-to-Text Transfer Transformer, and introducing a re-ranking mechanism to model specificity in definitions. Experimental results on standard evaluation datasets indicate that our method significantly outperforms the previous state-of-the-art method. Moreover, manual evaluation confirms that our method effectively addresses the over/under-specificity problems. Tomoyuki Kajiwara, Yuki Arase |
EMNLP (1) | 2 |
| 2021 | Language-agnostic Representation from Multilingual Sentence Encoders for Cross-lingual Similarity EstimationabstractWe propose a method to distil languageagnostic meaning embedding using a multilingual sentence encoder.By removing languagespecific information from the original embedding, we retrieve an embedding that fully represents the meaning of the sentence.The proposed method relies only on parallel corpora without any human annotations.Our meaning embedding allows for efficient cross-lingual sentence similarity estimation using a simple cosine similarity calculation.Experimental results of both the quality estimation of machine translation and cross-lingual semantic textual similarity tasks reveal that our method consistently outperforms the strong baselines using the original multilingual embeddings.The method also consistently improves the performance of any pre-trained multilingual sentence encoder, even in low-resource language pairs, where only tens of thousands of parallel sentence pairs are available.1 Nattapong Tiyajamorn, Tomoyuki Kajiwara, Yuki Arase, Makoto Onizuka |
EMNLP (1) | 2 |
| 2021 | WRIME: A New Dataset for Emotional Intensity Estimation with Subjective and Objective AnnotationsabstractTomoyuki Kajiwara, Chenhui Chu, Noriko Takemura, Yuta Nakashima, Hajime Nagahara. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tomoyuki Kajiwara, Chenhui Chu, Noriko Takemura, Yuta Nakashima, Hajime Nagahara |
NAACL-HLT | 1 |
| 2020 | Monolingual Transfer Learning via Bilingual Translators for Style-Sensitive Paraphrase Generation
Tomoyuki Kajiwara, Biwa Miura, Yuki Arase |
AAAI | 1 |
| 2020 | Text Classification with Negative SupervisionabstractAdvanced pre-trained models for text representation have achieved state-of-the-art performance on various text classification tasks.However, the discrepancy between the semantic similarity of texts and labelling standards affects classifiers, i.e. leading to lower performance in cases where classifiers should assign different labels to semantically similar texts.To address this problem, we propose a simple multitask learning model that uses negative supervision.Specifically, our model encourages texts with different labels to have distinct representations.Comprehensive experiments show that our model outperforms the stateof-the-art pre-trained model on both singleand multi-label classifications, sentence and document classifications, and classifications in three different languages. Sora Ohashi, Junya Takayama, Tomoyuki Kajiwara, Chenhui Chu, Yuki Arase |
ACL | 3 |
| 2020 | Tiny Word Embeddings Using Globally Informed ReconstructionabstractWe reduce the model size of pre-trained word embeddings by a factor of 200 while preserving its quality.Previous studies in this direction created a smaller word embedding model by reconstructing pre-trained word representations from those of subwords, which allows to store only a smaller number of subword embeddings in the memory.However, previous studies that train the reconstruction models using only target words cannot reduce the model size extremely while preserving its quality.Inspired by the observation of words with similar meanings having similar embeddings, our reconstruction training learns the global relationships among words, which can be employed in various models for word embedding reconstruction.Experimental results on word similarity benchmarks show that the proposed method improves the performance of the all subword-based reconstruction models. Sora Ohashi, Mao Isogawa, Tomoyuki Kajiwara, Yuki Arase |
COLING | 3 |
| 2020 | SOME: Reference-less Sub-Metrics Optimized for Manual Evaluations of Grammatical Error CorrectionabstractWe propose a reference-less metric trained on manual evaluations of system outputs for grammatical error correction.Previous studies have shown that reference-less metrics are promising; however, existing metrics are not optimized for manual evaluation of the system output because there is no dataset of system output with manual evaluation.This study manually evaluates the output of grammatical error correction systems to optimize the metrics.Experimental results show that the proposed metric improves the correlation with manual evaluation in both systemand sentence-level meta-evaluation.Our dataset and metric will be made publicly available. Ryoma Yoshimura, Masahiro Kaneko, Tomoyuki Kajiwara, Mamoru Komachi |
COLING | 3 |
| 2020 | Double Attention-based Multimodal Neural Machine Translation with Semantic Image RegionsabstractExisting studies on multimodal neural machine translation (MNMT) have mainly focused on the effect of combining visual and textual modalities to improve translations. However, it has been suggested that the visual modality is only marginally beneficial. Conventional visual attention mechanisms have been used to select the visual features from equally-sized grids generated by convolutional neural networks (CNNs), and may have had modest effects on aligning the visual concepts associated with textual objects, because the grid visual features do not capture semantic information. In contrast, we propose the application of semantic image regions for MNMT by integrating visual and textual features using two individual attention mechanisms (double attention). We conducted experiments on the Multi30k dataset and achieved an improvement of 0.5 and 0.9 BLEU points for English-German and English-French translation tasks, compared with the MNMT with grid visual features. We also demonstrated concrete improvements on translation performance benefited from semantic image regions. Mamoru Komachi, Tomoyuki Kajiwara, Chenhui Chu |
EAMT | 3 |
| 2020 | Annotation of Adverse Drug Reactions in Patients' WeblogsabstractAdverse drug reactions are a severe problem that significantly degrade quality of life, or even threaten the life of patients. Patient-generated texts available on the web have been gaining attention as a promising source of information in this regard. While previous studies annotated such patient-generated content, they only reported on limited information, such as whether a text described an adverse drug reaction or not. Further, they only annotated short texts of a few sentences crawled from online forums and social networking services. The dataset we present in this paper is unique for the richness of annotated information, including detailed descriptions of drug reactions with full context. We crawled patient’s weblog articles shared on an online patient-networking platform and annotated the effects of drugs therein reported. We identified spans describing drug reactions and assigned labels for related drug names, standard codes for the symptoms of the reactions, and types of effects. As a first dataset, we annotated 677 drug reactions with these detailed labels based on 169 weblog articles by Japanese lung cancer patients. Our annotation dataset is made publicly available at our web site (https://yukiar.github.io/adr-jp/) for further research on the detection of adverse drug reactions and more broadly, on patient-generated text processing. Yuki Arase, Tomoyuki Kajiwara, Chenhui Chu |
LREC | 2 |
| 2020 | Word Complexity Estimation for Japanese Lexical SimplificationabstractWe introduce three language resources for Japanese lexical simplification: 1) a large-scale word complexity lexicon, 2) the first synonym lexicon for converting complex words to simpler ones, and 3) the first toolkit for developing and benchmarking Japanese lexical simplification system. Our word complexity lexicon is expanded to a broader vocabulary using a classifier trained on a small, high-quality word complexity lexicon created by Japanese language teachers. Based on this word complexity estimator, we extracted simplified word pairs from a large-scale synonym lexicon and constructed a simplified synonym lexicon useful for lexical simplification. In addition, we developed a Python library that implements automatic evaluation and key methods in each subtask to ease the construction of a lexical simplification pipeline. Experimental results show that the proposed method based on our lexicon achieves the highest performance of Japanese lexical simplification. The current lexical simplification is mainly studied in English, which is rich in language resources such as lexicons and toolkits. The language resources constructed in this study will help advance the lexical simplification system in Japanese. Daiki Nishihara, Tomoyuki Kajiwara |
LREC | 2 |
| 2020 | SAPPHIRE: Simple Aligner for Phrasal Paraphrase with Hierarchical RepresentationabstractWe present SAPPHIRE, a Simple Aligner for Phrasal Paraphrase with HIerarchical REpresentation. Monolingual phrase alignment is a fundamental problem in natural language understanding and also a crucial technique in various applications such as natural language inference and semantic textual similarity assessment. Previous methods for monolingual phrase alignment are language-resource intensive; they require large-scale synonym/paraphrase lexica and high-quality parsers. Different from them, SAPPHIRE depends only on a monolingual corpus to train word embeddings. Therefore, it is easily transferable to specific domains and different languages. Specifically, SAPPHIRE first obtains word alignments using pre-trained word embeddings and then expands them to phrase alignments by bilingual phrase extraction methods. To estimate the likelihood of phrase alignments, SAPPHIRE uses phrase embeddings that are hierarchically composed of word embeddings. Finally, SAPPHIRE searches for a set of consistent phrase alignments on a lattice of phrase alignment candidates. It achieves search-efficiency by constraining the lattice so that all the paths go through a phrase alignment pair with the highest alignment score. Experimental results using the standard dataset for phrase alignment evaluation show that SAPPHIRE outperforms the previous method and establishes the state-of-the-art performance. Masato Yoshinaka, Tomoyuki Kajiwara, Yuki Arase |
LREC | 2 |
| 2019 | Negative Lexically Constrained Decoding for Paraphrase GenerationabstractParaphrase generation can be regarded as monolingual translation.Unlike bilingual machine translation, paraphrase generation rewrites only a limited portion of an input sentence.Hence, previous methods based on machine translation often perform conservatively to fail to make necessary rewrites.To solve this problem, we propose a neural model for paraphrase generation that first identifies words in the source sentence that should be paraphrased.Then, these words are paraphrased by the negative lexically constrained decoding that avoids outputting these words as they are.Experiments on text simplification and formality transfer show that our model improves the quality of paraphrasing by making necessary rewrites to an input sentence. Tomoyuki Kajiwara |
ACL (1) | 1 |
| 2018 | Contextualized Word Representations for Multi-Sense Embedding
Kazuki Ashihara, Tomoyuki Kajiwara, Yuki Arase, Satoru Uchida |
PACLIC | 2 |
| 2017 | MIPA: Mutual Information Based Paraphrase Acquisition via Bilingual PivotingabstractWe present a pointwise mutual information (PMI)-based approach to formalize paraphrasability and propose a variant of PMI, called MIPA, for the paraphrase acquisition. Our paraphrase acquisition method first acquires lexical paraphrase pairs by bilingual pivoting and then reranks them by PMI and distributional similarity. The complementary nature of information from bilingual corpora and from monolingual corpora makes the proposed method robust. Experimental results show that the proposed method substantially outperforms bilingual pivoting and distributional similarity themselves in terms of metrics such as MRR, MAP, coverage, and Spearman’s correlation. Tomoyuki Kajiwara, Mamoru Komachi, Daichi Mochihashi |
IJCNLP(1) | 1 |
| 2016 | Building a Monolingual Parallel Corpus for Text Simplification Using Sentence Similarity Based on Alignment between Word EmbeddingsabstractMethods for text simplification using the framework of statistical machine translation have been extensively studied in recent years. However, building the monolingual parallel corpus necessary for training the model requires costly human annotation. Monolingual parallel corpora for text simplification have therefore been built only for a limited number of languages, such as English and Portuguese. To obviate the need for human annotation, we propose an unsupervised method that automatically builds the monolingual parallel corpus for text simplification using sentence similarity based on word embeddings. For any sentence pair comprising a complex sentence and its simple counterpart, we employ a many-to-one method of aligning each word in the complex sentence with the most similar word in the simple sentence and compute sentence similarity by averaging these word similarities. The experimental results demonstrate the excellent performance of the proposed method in a monolingual parallel corpus construction task for English text simplification. The results also demonstrated the superior accuracy in text simplification that use the framework of statistical machine translation trained using the corpus built by the proposed method to that using the existing corpora. Tomoyuki Kajiwara, Mamoru Komachi |
COLING | 1 |
| 2014 | Noun Paraphrasing Based on a Variety of Contexts
Tomoyuki Kajiwara, Kazuhide Yamamoto |
PACLIC | 1 |