VLDB 2026 Research / reviewers in the wild / expert
Naoaki Okazaki
dblp:49/4018
· DBLP profile ↗
115ranked-venue papers
11as first author
36since 2021 · last 2026
0000-0001-7635-6175ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 104 · 9 first-author · 36 since 2021Databases, data management, data science and information retrieval · 8Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Discriminability of Vision-Language Models
Masayasu Muraoka, Naoaki Okazaki |
LREC | 2 |
| 2026 | Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
Masanari Oi, Masahiro Kaneko, Naoaki Okazaki, Nakamasa Inoue |
LREC | 3 |
| 2025 | HMoE: Heterogeneous Mixture of Experts for Language ModelingabstractAn Wang, Xingwu Sun, Ruobing Xie, Shuaipeng Li, Jiaqi Zhu, Zhen Yang, Pinxue Zhao, Weidong Han, Zhanhui Kang, Di Wang, Naoaki Okazaki, Cheng-zhong Xu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xingwu Sun, Ruobing Xie, Shuaipeng Li, Jiaqi Zhu 0004, Pinxue Zhao, Weidong Han 0006, Zhanhui Kang, Di Wang 0052, Naoaki Okazaki, Cheng-Zhong Xu 0001 |
EMNLP | 11 |
| 2024 | OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated ExamplesabstractLarge Language Models (LLMs) have achieved human-level fluency in text generation, making it difficult to distinguish between human-written and LLM-generated texts. This poses a growing risk of misuse of LLMs and demands the development of detectors to identify LLM-generated texts. However, existing detectors lack robustness against attacks: they degrade detection accuracy by simply paraphrasing LLM-generated texts. Furthermore, a malicious user might attempt to deliberately evade the detectors based on detection results, but this has not been assumed in previous studies. In this paper, we propose OUTFOX, a framework that improves the robustness of LLM-generated-text detectors by allowing both the detector and the attacker to consider each other's output. In this framework, the attacker uses the detector's prediction labels as examples for in-context learning and adversarially generates essays that are harder to detect, while the detector uses the adversarially generated essays as examples for in-context learning to learn to detect essays from a strong attacker. Experiments in the domain of student essays show that the proposed detector improves the detection performance on the attacker-generated texts by up to +41.3 points F1-score. Furthermore, the proposed detector shows a state-of-the-art detection performance: up to 96.9 points F1-score, beating existing detectors on non-attacked texts. Finally, the proposed attacker drastically degrades the performance of detectors by up to -57.0 points F1-score, massively outperforming the baseline paraphrasing method for evading detection. Ryuto Koike, Masahiro Kaneko, Naoaki Okazaki |
AAAI | 3 |
| 2024 | Evaluating Gender Bias of Pre-trained Language Models in Natural Language Inference by Considering All LabelsabstractDiscriminatory gender biases have been found in Pre-trained Language Models (PLMs) for multiple languages. In Natural Language Inference (NLI), existing bias evaluation methods have focused on the prediction results of one specific label out of three labels, such as neutral. However, such evaluation methods can be inaccurate since unique biased inferences are associated with unique prediction labels. Addressing this limitation, we propose a bias evaluation method for PLMs, called NLI-CoAL, which considers all the three labels of NLI task. First, we create three evaluation data groups that represent different types of biases. Then, we define a bias measure based on the corresponding label output of each data group. In the experiments, we introduce a meta-evaluation technique for NLI bias measures and use it to confirm that our bias measure can distinguish biased, incorrect inferences from non-biased incorrect inferences better than the baseline, resulting in a more accurate bias evaluation. We create the datasets in English, Japanese, and Chinese, and successfully validate the compatibility of our bias measure across multiple languages. Lastly, we observe the bias tendencies in PLMs of different languages. To our knowledge, we are the first to construct evaluation datasets and measure PLMs’ bias from NLI in Japanese and Chinese. Panatchakorn Anantaprayoon, Masahiro Kaneko, Naoaki Okazaki |
LREC/COLING | 3 |
| 2024 | Two Counterexamples to Tokenization and the Noiseless ChannelabstractIn Tokenization and the Noiseless Channel (Zouhar et al., 2023), Rényi efficiency is suggested as an intrinsic mechanism for evaluating a tokenizer: for NLP tasks, the tokenizer which leads to the highest Rényi efficiency of the unigram distribution should be chosen. The Rényi efficiency is thus treated as a predictor of downstream performance (e.g., predicting BLEU for a machine translation task), without the expensive step of training multiple models with different tokenizers. Although useful, the predictive power of this metric is not perfect, and the authors note there are additional qualities of a good tokenization scheme that Rényi efficiency alone cannot capture. We describe two variants of BPE tokenization which can arbitrarily increase Rényi efficiency while decreasing the downstream model performance. These counterexamples expose cases where Rényi efficiency fails as an intrinsic tokenization metric and thus give insight for building more accurate predictors. Marco Cognetta, Vilém Zouhar, Sangwhan Moon, Naoaki Okazaki |
LREC/COLING | 4 |
| 2024 | Controlled Generation with Prompt Insertion for Natural Language Explanations in Grammatical Error CorrectionabstractIn Grammatical Error Correction (GEC), it is crucial to ensure the user’s comprehension of a reason for correction. Existing studies present tokens, examples, and hints for corrections, but do not directly explain the reasons in natural language. Although methods that use Large Language Models (LLMs) to provide direct explanations in natural language have been proposed for various tasks, no such method exists for GEC. Generating explanations for GEC corrections involves aligning input and output tokens, identifying correction points, and presenting corresponding explanations consistently. However, it is not straightforward to specify a complex format to generate explanations, because explicit control of generation is difficult with prompts. This study introduces a method called controlled generation with Prompt Insertion (PI) so that LLMs can explain the reasons for corrections in natural language. In PI, LLMs first correct the input text, and then we automatically extract the correction points based on the rules. The extracted correction points are sequentially inserted into the LLM’s explanation output as prompts, guiding the LLMs to generate explanations for the correction points. We also create an Explainable GEC (XGEC) dataset of correction reasons by annotating NUCLE, CoNLL2013, and CoNLL2014. Although generations from GPT-3.5 and ChatGPT using original prompts miss some correction points, the generation control using PI can explicitly guide to describe explanations for all correction points, contributing to improved performance in generating correction reasons. Masahiro Kaneko, Naoaki Okazaki |
LREC/COLING | 2 |
| 2024 | Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual TransferabstractDocument-level Relation Extraction (DocRE) is the task of extracting all semantic relationships from a document. While studies have been conducted on English DocRE, limited attention has been given to DocRE in non-English languages. This work delves into effectively utilizing existing English resources to promote DocRE studies in non-English languages, with Japanese as the representative case. As an initial attempt, we construct a dataset by transferring an English dataset to Japanese. However, models trained on such a dataset are observed to suffer from low recalls. We investigate the error cases and attribute the failure to different surface structures and semantics of documents translated from English and those written by native speakers. We thus switch to explore if the transferred dataset can assist human annotation on Japanese documents. In our proposal, annotators edit relation predictions from a model trained on the transferred dataset. Quantitative analysis shows that relation recommendations suggested by the model help reduce approximately 50% of the human edit steps compared with the previous approach. Experiments quantify the performance of existing DocRE models on our collected dataset, portraying the challenges of Japanese and cross-lingual DocRE. Youmi Ma, Naoaki Okazaki |
LREC/COLING | 3 |
| 2024 | SAIE Framework: Support Alone Isn't Enough - Advancing LLM Training with Adversarial RemarksabstractLarge Language Models (LLMs) can justify or critique their predictions through discussions with other models or humans, thereby enriching their intrinsic understanding of instances. While proactive discussions in the inference phase have been shown to boost performance, such interactions have not been extensively explored during the training phase. We hypothesize that incorporating interactive discussions into the training process can enhance the models’ understanding and improve their reasoning and verbal expression abilities during inference. This work introduces the SAIE framework, facilitating supportive and adversarial discussions between learner and partner models. The learner model receives responses from the partner, and its parameters are then updated based on this discussion. This dynamic adjustment process continues throughout the training phase, responding to the evolving outputs of the learner model. Our empirical evaluation across various tasks, including math word problems, commonsense reasoning, and multi-domain knowledge, demonstrates that models fine-tuned with the SAIE framework outperform those trained with conventional fine-tuning approaches. Furthermore, our method enhances the models’ reasoning capabilities, improving both individual and multi-agent inference performance. Our code is available at https://github.com/loem-ms/saie. Mengsay Loem, Masahiro Kaneko, Naoaki Okazaki |
ECAI | 3 |
| 2024 | Distributional Properties of Subword RegularizationabstractSubword regularization, used widely in NLP, improves model performance by reducing the dependency on exact tokenizations, augmenting the training corpus, and exposing the model to more unique contexts during training.BPE and MaxMatch, two popular subword tokenization schemes, have stochastic dropout regularization variants.However, there has not been an analysis of the distributions formed by them.We show that these stochastic variants are heavily biased towards a small set of tokenizations per word.If the benefits of subword regularization are as mentioned, we hypothesize that biasedness artificially limits the effectiveness of these schemes.Thus, we propose an algorithm to uniformly sample tokenizations that we use as a drop-in replacement for the stochastic aspects of existing tokenizers, and find that it improves machine translation quality. Marco Cognetta, Vilém Zouhar, Naoaki Okazaki |
EMNLP | 3 |
| 2024 | n-gram F-score for Evaluating Grammatical Error CorrectionabstractM 2 and its variants are the most widely used automatic evaluation metrics for grammatical error correction (GEC), which calculate an F -score using a phrase-based alignment between sentences.However, it is not straightforward at all to align learner sentences containing errors to their correct sentences.In addition, alignment calculations are computationally expensive.We propose GREEN, an alignment-free F -score for GEC evaluation.GREEN treats a sentence as a multiset of ngrams and extracts edits between sentences by set operations instead of computing an alignment.Our experiments confirm that GREEN performs better than existing methods for the corpus-level metrics and comparably for the sentence-level metrics even without computing an alignment.GREEN is available at https://github.com/shotakoyama/green. Shota Koyama, Ryo Nagata, Hiroya Takamura, Naoaki Okazaki |
INLG | 4 |
| 2023 | Parameter-Efficient Korean Character-Level Language ModelingabstractCharacter-level language modeling has been shown empirically to perform well on highly agglutinative or morphologically rich languages while using only a small fraction of the parameters required by (sub)word models.Korean fits nicely into this framework, except that, like other CJK languages, it has a very large character vocabulary of 11,172 unique syllables.However, unlike Japanese Kanji and Chinese Hanzi, each Korean syllable can be uniquely factored into a small set of subcharacters, called jamo.We explore a "three-hot" scheme, where we exploit the decomposability of Korean characters to model at the syllable level but using only jamo-level representations.We find that our three-hot embedding and decoding scheme alleviates the two major issues with prior syllableand jamo-level models.Namely, it requires fewer than 1% of the embedding parameters of a syllable model, and it does not require tripling the sequence length, as with jamo models.In addition, it addresses a theoretical flaw in a prior three-hot modeling scheme.Our experiments show that, even when reducing the number of embedding parameters by > 99.6% (from 11.4M to just 36k), our model suffers no loss in translation quality compared to the baseline syllable model. Marco Cognetta, Sangwhan Moon, Lawrence Wolf-Sonkin, Naoaki Okazaki |
EACL | 4 |
| 2023 | Comparing Intrinsic Gender Bias Evaluation Measures without using Human Annotated ExamplesabstractNumerous types of social biases have been identified in pre-trained language models (PLMs), and various intrinsic bias evaluation measures have been proposed for quantifying those social biases.Prior works have relied on human annotated examples to compare existing intrinsic bias evaluation measures.However, this approach is not easily adaptable to different languages nor amenable to large scale evaluations due to the costs and difficulties when recruiting human annotators.To overcome this limitation, we propose a method to compare intrinsic gender bias evaluation measures without relying on human-annotated examples.Specifically, we create multiple bias-controlled versions of PLMs using varying amounts of male vs. female gendered sentences, mined automatically from an unannotated corpus using genderrelated word lists.Next, each bias-controlled PLM is evaluated using an intrinsic bias evaluation measure, and the rank correlation between the computed bias scores and the gender proportions used to fine-tune the PLMs is computed.Experiments on multiple corpora and PLMs repeatedly show that the correlations reported by our proposed method that does not require human annotated examples are comparable to those computed using human annotated examples in prior work. Masahiro Kaneko, Danushka Bollegala, Naoaki Okazaki |
EACL | 3 |
| 2023 | DREEAM: Guiding Attention with Evidence for Improving Document-Level Relation ExtractionabstractDocument-level relation extraction (DocRE)is the task of identifying all relations between each entity pair in a document.Evidence, defined as sentences containing clues for the relationship between an entity pair, has been shown to help DocRE systems focus on relevant texts, thus improving relation extraction.However, evidence retrieval (ER) in DocRE faces two major issues: high memory consumption and limited availability of annotations.This work aims at addressing these issues to improve the usage of ER in DocRE.First, we propose DREEAM, a memory-efficient approach that adopts evidence information as the supervisory signal, thereby guiding the attention modules of the DocRE system to assign high weights to evidence.Second, we propose a self-training strategy for DREEAM to learn ER from automatically-generated evidence on massive data without evidence annotations.Experimental results reveal that our approach exhibits state-of-the-art performance on the Do-cRED benchmark for both DocRE and ER.To the best of our knowledge, DREEAM is the first approach to employ ER self-training 1 . Youmi Ma, Naoaki Okazaki |
EACL | 3 |
| 2023 | Semantic Specialization for Knowledge-based Word Sense DisambiguationabstractA promising approach for knowledge-based Word Sense Disambiguation (WSD) is to select the sense whose contextualized embeddings computed for its definition sentence are closest to those computed for a target word in a given sentence.This approach relies on the similarity of the sense and context embeddings computed by a pre-trained language model.We propose a semantic specialization for WSD where contextualized embeddings are adapted to the WSD task using solely lexical knowledge.The key idea is, for a given sense, to bring semantically related senses and contexts closer and send different/unrelated senses farther away.We realize this idea as the joint optimization of the Attract-Repel objective for sense pairs and the self-training objective for context-sense pairs while controlling deviations from the original embeddings.The proposed method outperformed previous studies that adapt contextualized embeddings.It achieved state-of-the-art performance on knowledge-based WSD when combined with the reranking heuristic that uses the sense inventory.We found that the similarity characteristics of specialized embeddings conform to the key idea.We also found that the (dis)similarity of embeddings between the related/different/unrelated senses correlates well with the performance of WSD. Sakae Mizuki, Naoaki Okazaki |
EACL | 2 |
| 2023 | Reducing Sequence Length by Predicting Edit Spans with Large Language ModelsabstractLarge Language Models (LLMs) have demonstrated remarkable performance in various tasks and gained significant attention.LLMs are also used for local sequence transduction tasks, including grammatical error correction (GEC) and formality style transfer, where most tokens in a source text are kept unchanged.However, the models that generate all target tokens in such tasks have a tendency to simply copy the input text as is, without making needed changes, because the difference between input and output texts is minimal in the training data.This is also inefficient because the computational cost grows quadratically with the target sequence length with Transformer.This paper proposes predicting edit spans for the source text for local sequence transduction tasks.Representing an edit span with a position of the source text and corrected tokens, we can reduce the length of the target sequence and the computational cost for inference.We apply instruction tuning for LLMs on the supervision data of edit spans.Experiments show that the proposed method achieves comparable performance to the baseline in four tasks, paraphrasing, formality style transfer, GEC, and text simplification, despite reducing the length of the target text by as small as 21%.Furthermore, we report that the task-specific fine-tuning with the proposed method achieved state-of-the-art performance in the four tasks. Masahiro Kaneko, Naoaki Okazaki |
EMNLP | 2 |
| 2023 | Causal Reasoning through Two Cognition Layers for Improving Generalization in Visual Question AnsweringabstractGeneralization in Visual Question Answering (VQA) requires models to answer questions about images with contexts beyond the training distribution.Existing attempts primarily refine unimodal aspects, overlooking enhancements in multimodal aspects.Besides, diverse interpretations of the input lead to various modes of answer generation, highlighting the role of causal reasoning between interpreting and answering steps in VQA.Through this lens, we propose Cognitive pathways VQA (CopVQA) improving the multimodal predictions by emphasizing causal reasoning factors.CopVQA first operates a pool of pathways that capture diverse causal reasoning flows through interpreting and answering stages.Mirroring human cognition, we decompose the responsibility of each stage into distinct experts and a cognition-enabled component (CC).The two CCs strategically execute one expert for each stage at a time.Finally, we prioritize answer predictions governed by pathways involving both CCs while disregarding answers produced by either CC, thereby emphasizing causal reasoning and supporting generalization.Our experiments on real-life and medical data consistently verify that CopVQA improves VQA performance and generalization across baselines and domains.Notably, CopVQA achieves a new state-of-the-art (SOTA) on the PathVQA dataset and comparable accuracy to the current SOTA on VQA-CPv2, VQAv2, and VQA-RAD, with one-fourth of the model size. Naoaki Okazaki |
EMNLP | 2 |
| 2023 | North Korean Neural Machine Translation through South Korean ResourcesabstractSouth and North Korea both use the Korean language. However, Korean natural language processing (NLP) research has mostly focused on South Korean language. Therefore, existing NLP systems in the Korean language, such as neural machine translation (NMT) systems, cannot properly process North Korean inputs. Training a model using North Korean data is the most straightforward approach to solving this problem, but the data to train NMT models are insufficient. To solve this problem, we constructed a parallel corpus to develop a North Korean NMT model using a comparable corpus. We manually aligned parallel sentences to create evaluation data and automatically aligned the remaining sentences to create training data. We trained a North Korean NMT model using our North Korean parallel data and improved North Korean translation quality using South Korean resources such as parallel data and a pre-trained model. In addition, we propose Korean-specific pre-processing methods, character tokenization, and phoneme decomposition to use the South Korean resources more efficiently. We demonstrate that the phoneme decomposition consistently improves the North Korean translation accuracy compared to other pre-processing methods. Hwichan Kim, Tosho Hirasawa, Sangwhan Moon, Naoaki Okazaki, Mamoru Komachi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2022 | Interpretability for Language Learners Using Example-Based Grammatical Error CorrectionabstractGrammatical Error Correction (GEC) should focus not only on correction accuracy but also on the interpretability of the results for language learners.However, existing neuralbased GEC models mostly focus on improving accuracy, while their interpretability has not been explored.Example-based methods are promising for improving interpretability, which use similar retrieved examples to generate corrections.Furthermore, examples are beneficial in language learning, helping learners to understand the basis for grammatically incorrect/correct texts and improve their confidence in writing.Therefore, we hypothesized that incorporating an example-based method into GEC could improve interpretability and support language learners.In this study, we introduce an Example-Based GEC (EB-GEC) that presents examples to language learners as a basis for correction result.The examples consist of pairs of correct and incorrect sentences similar to a given input and its predicted correction.Experiments demonstrate that the examples presented by EB-GEC help language learners decide whether to accept or refuse suggestions from the GEC output.Furthermore, the experiments show that retrieved examples also improve the accuracy of corrections. Masahiro Kaneko, Sho Takase, Ayana Niwa, Naoaki Okazaki |
ACL (1) | 4 |
| 2022 | Semi-Supervised Formality Style Transfer with Consistency TrainingabstractFormality style transfer (FST) is a task that involves paraphrasing an informal sentence into a formal one without altering its meaning.To address the data-scarcity problem of existing parallel datasets, previous studies tend to adopt a cycle-reconstruction scheme to utilize additional unlabeled data, where the FST model mainly benefits from target-side unlabeled sentences.In this work, we propose a simple yet effective semi-supervised framework to better utilize source-side unlabeled sentences based on consistency training.Specifically, our approach augments pseudo-parallel data obtained from a source-side informal sentence by enforcing the model to generate similar outputs for its perturbed version.Moreover, we empirically examined the effects of various data perturbation methods and propose effective data filtering strategies to improve our framework.Experimental results on the GYAFC benchmark demonstrate that our approach can achieve state-of-the-art results, even with less than 40% of the parallel data 1 . Ao Liu 0008, Naoaki Okazaki |
ACL (1) | 3 |
| 2022 | Debiasing Isn't Enough! - on the Effectiveness of Debiasing MLMs and Their Social Biases in Downstream TasksabstractWe study the relationship between task-agnostic intrinsic and task-specific extrinsic social bias evaluation measures for MLMs, and find that there exists only a weak correlation between these two types of evaluation measures. Moreover, we find that MLMs debiased using different methods still re-learn social biases during fine-tuning on downstream tasks. We identify the social biases in both training instances as well as their assigned labels as reasons for the discrepancy between intrinsic and extrinsic bias evaluation measurements. Overall, our findings highlight the limitations of existing MLM bias evaluation measures and raise concerns on the deployment of MLMs in downstream applications using those measures. Masahiro Kaneko, Danushka Bollegala, Naoaki Okazaki |
COLING | 3 |
| 2022 | IMPARA: Impact-Based Metric for GEC Using Parallel DataabstractAutomatic evaluation of grammatical error correction (GEC) is essential in developing useful GEC systems. Existing methods for automatic evaluation require multiple reference sentences or manual scores. However, such resources are expensive, thereby hindering automatic evaluation for various domains and correction styles. This paper proposes an Impact-based Metric for GEC using PARAllel data, IMPARA, which utilizes correction impacts computed by parallel data comprising pairs of grammatical/ungrammatical sentences. As parallel data is cheaper than manually assessing evaluation scores, IMPARA can reduce the cost of data creation for automatic evaluation. Correlations between IMPARA and human scores indicate that IMPARA is comparable or better than existing evaluation methods. Furthermore, we find that IMPARA can perform evaluations that fit different domains and correction styles trained on various parallel data. Koki Maeda, Masahiro Kaneko, Naoaki Okazaki |
COLING | 3 |
| 2022 | PLOG: Table-to-Logic Pretraining for Logical Table-to-Text GenerationabstractLogical table-to-text generation is a task that involves generating logically faithful sentences from tables, which requires models to derive logical-level facts from table records via logical inference.It raises a new challenge on the logical-level content planning of table-to-text models.However, directly learning the logical inference knowledge from table-text pairs is very difficult for neural models because of the ambiguity of natural language and the scarcity of parallel data.Hence even large-scale pretrained language models present low logical fidelity on logical table-to-text.In this work, we propose a Pretrained Logical Form Generator (PLOG) framework to improve generation fidelity.Specifically, PLOG is first pretrained on a table-to-logical-form generation (table-to-logic) task, then finetuned on downstream table-to-text tasks.The logical forms are formally defined with unambiguous semantics.Hence we can collect a large amount of accurate logical forms from tables without human annotation.In addition, PLOG can learn logical inference from table-logic pairs much more reliably than from table-text pairs.To evaluate our model, we further collect a controlled logical table-to-text dataset CONTLOG based on an existing dataset.On two benchmarks, LOGICNLG and CONTLOG, PLOG outperforms strong baselines by a large margin on logical fidelity, demonstrating the effectiveness of table-to-logic pretraining. Ao Liu 0008, Haoyu Dong 0001, Naoaki Okazaki, Shi Han, Dongmei Zhang 0001 |
EMNLP | 3 |
| 2022 | Learning How to Translate North Korean through South KoreanabstractSouth and North Korea both use the Korean language. However, Korean NLP research has focused on South Korean only, and existing NLP systems of the Korean language, such as neural machine translation (NMT) models, cannot properly handle North Korean inputs. Training a model using North Korean data is the most straightforward approach to solving this problem, but there is insufficient data to train NMT models. In this study, we create data for North Korean NMT models using a comparable corpus. First, we manually create evaluation data for automatic alignment and machine translation, and then, investigate automatic alignment methods suitable for North Korean. Finally, we show that a model trained by North Korean bilingual data without human annotation significantly boosts North Korean translation accuracy compared to existing South Korean models in zero-shot settings. Hwichan Kim, Sangwhan Moon, Naoaki Okazaki, Mamoru Komachi |
LREC | 3 |
| 2022 | OpenKorPOS: Democratizing Korean Tokenization with Voting-Based Open Corpus AnnotationabstractKorean is a language with complex morphology that uses spaces at larger-than-word boundaries, unlike other East-Asian languages. While morpheme-based text generation can provide significant semantic advantages compared to commonly used character-level approaches, Korean morphological analyzers only provide a sequence of morpheme-level tokens, losing information in the tokenization process. Two crucial issues are the loss of spacing information and subcharacter level morpheme normalization, both of which make the tokenization result challenging to reconstruct the original input string, deterring the application to generative tasks. As this problem originates from the conventional scheme used when creating a POS tagging corpus, we propose an improvement to the existing scheme, which makes it friendlier to generative tasks. On top of that, we suggest a fully-automatic annotation of a corpus by leveraging public analyzers. We vote the surface and POS from the outcome and fill the sequence with the selected morphemes, yielding tokenization with a decent quality that incorporates space information. Our scheme is verified via an evaluation done on an external corpus, and subsequently, it is adapted to Korean Wikipedia to construct an open, permissive resource. We compare morphological analyzer performance trained on our corpus with existing methods, then perform an extrinsic evaluation on a downstream task. Sangwhan Moon, Won-Ik Cho, Hye Joo Han, Naoaki Okazaki, Nam Soo Kim |
LREC | 4 |
| 2022 | Multi-Task Learning for Cross-Lingual Abstractive SummarizationabstractWe present a multi-task learning framework for cross-lingual abstractive summarization to augment training data. Recent studies constructed pseudo cross-lingual abstractive summarization data to train their neural encoder-decoders. Meanwhile, we introduce existing genuine data such as translation pairs and monolingual abstractive summarization data into training. Our proposed method, Transum, attaches a special token to the beginning of the input sentence to indicate the target task. The special token enables us to incorporate the genuine data into the training data easily. The experimental results show that Transum achieves better performance than the model trained with only pseudo cross-lingual summarization data. In addition, we achieve the top ROUGE score on Chinese-English and Arabic-English abstractive summarization. Moreover, Transum also has a positive effect on machine translation. Experimental results indicate that Transum improves the performance from the strong baseline, Transformer, in Chinese-English, Arabic-English, and English-Japanese translation datasets. Sho Takase, Naoaki Okazaki |
LREC | 2 |
| 2022 | Gender Bias in Masked Language Models for Multiple LanguagesabstractMasahiro Kaneko, Aizhan Imankulova, Danushka Bollegala, Naoaki Okazaki. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Masahiro Kaneko, Aizhan Imankulova, Danushka Bollegala, Naoaki Okazaki |
NAACL-HLT | 4 |
| 2022 | Improving Automatic Evaluation of Acceptability Based on Language Models with a Coarse Sentence Representation
Vijay Daultani, Naoaki Okazaki |
PACLIC | 2 |
| 2022 | Annotating Entity and Causal Relationships on Japanese Vehicle Recall Information
Hsuan-Yu Kuo, Youmi Ma, Naoaki Okazaki |
PACLIC | 3 |
| 2022 | Recurrent Neural Hidden Markov Model for High-order TransitionabstractWe propose a method to pay attention to high-order relations among latent states to improve the conventional HMMs that focus only on the latest latent state, since they assume Markov property. To address the high-order relations, we apply an RNN to each sequence of latent states, because the RNN can represent the information of an arbitrary-length sequence with their cell: a fixed-size vector. However, the simplest way, which provides all latent sequences explicitly for the RNN, is intractable due to the combinatorial explosion of the search space of latent states. Thus, we modify the RNN to represent the history of latent states from the beginning of the sequence to the current state with a fixed number of RNN cells whose number is equal to the number of possible states. We conduct experiments on unsupervised POS tagging and synthetic datasets. Experimental results show that the proposed method achieves better performance than previous methods. In addition, the results on the synthetic dataset indicate that the proposed method can capture the high-order relations. Tatsuya Hiraoka, Sho Takase, Kei Uchiumi, Atsushi Keyaki, Naoaki Okazaki |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2021 | Transformer-based Lexically Constrained Headline GenerationabstractThis paper explores a variant of automatic headline generation methods, where a generated headline is required to include a given phrase such as a company or a product name.Previous methods using Transformer-based models generate a headline including a given phrase by providing the encoder with additional information corresponding to the given phrase.However, these methods cannot always include the phrase in the generated headline.Inspired by previous RNN-based methods generating token sequences in backward and forward directions from the given phrase, we propose a simple Transformerbased method that guarantees to include the given phrase in the high-quality generated headline.We also consider a new headline generation strategy that takes advantage of the controllable generation order of Transformer.Our experiments with the Japanese News Corpus demonstrate that our methods, which are guaranteed to include the phrase in the generated headline, achieve ROUGE scores comparable to previous Transformer-based methods.We also show that our generation strategy performs better than previous strategies. Kosuke Yamada, Yuta Hitomi, Hideaki Tamori, Ryohei Sasano, Naoaki Okazaki, Kentaro Inui, Koichi Takeda 0003 |
EMNLP (1) | 5 |
| 2021 | Predicting Antonyms in Context using BERTabstractWe address the task of antonym prediction in a context, which is a fill-in-the-blanks problem.This task setting is unique and practical because it requires contrastiveness to the other word and naturalness as a text in filling a blank.We propose methods for fine-tuning pre-trained masked language models (BERT) for contextaware antonym prediction.The experimental results show that these methods have positive impacts on the prediction of antonyms within a context.Moreover, human evaluation reveals that more than 85% of the predictions using the proposed method are acceptable as antonyms. Ayana Niwa, Keisuke Nishiguchi, Naoaki Okazaki |
INLG | 3 |
| 2021 | Incorporating Semantic Textual Similarity and Lexical Matching for Information Retrieval
Hiroki Iida, Naoaki Okazaki |
PACLIC | 2 |
| 2021 | Various Errors Improve Neural Grammatical Error Correction
Shota Koyama, Hiroya Takamura, Naoaki Okazaki |
PACLIC | 3 |
| 2021 | Multimodal Pretraining Unmasked: A Meta-Analysis and a Unified Framework of Vision-and-Language BERTsabstractAbstract Large-scale pretraining and task-specific fine- tuning is now the standard methodology for many tasks in computer vision and natural language processing. Recently, a multitude of methods have been proposed for pretraining vision and language BERTs to tackle challenges at the intersection of these two key areas of AI. These models can be categorized into either single-stream or dual-stream encoders. We study the differences between these two categories, and show how they can be unified under a single theoretical framework. We then conduct controlled experiments to discern the empirical differences between five vision and language BERTs. Our experiments show that training data and hyperparameters are responsible for most of the differences between the reported results, but they also reveal that the embedding layer plays a crucial role in these massive models. Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, Desmond Elliott |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Construction of a Corpus of Rhetorical Devices in Slogans and Structural Analysis of AntithesesabstractAn advertising slogan is a sentence that expresses a product or a work of art in a straightforward manner and is used for advertising and publicity. Moving the consumer's mind and attracting their interest can significantly influence sales. Although rhetorical techniques in a slogan are known to improve the effectiveness of advertising, not much attention has been devoted to analyze or automatically generate sentences with the techniques. Therefore, we constructed a large corpus of slogans and revealed the linguistic characteristics of the basic statistics and rhetorical devices. Another point of focus was antitheses, of which the usage rates are relatively high and which have a specific sentence structure and lexical constraints. The generation of a slogan that contains an antithesis necessitates the structure of sentences, known as templates, to be extracted and also requires knowledge of word pairs with semantic contrast. Thus, the next step involved analysis of the structure to extract the sentence structure and lexical knowledge about the antithesis. Despite its simple architecture, the proposed method exceeds the prediction accuracy and efficiency of a comparable method. Lexical knowledge that is not available in existing dictionaries was also extracted. Ayana Niwa, Naoaki Okazaki, Kohei Wakimoto, Keisuke Nishiguchi, Masataka Mouri |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2020 | It's Easier to Translate out of English than into it: Measuring Neural Translation Difficulty by Cross-Mutual InformationabstractThe performance of neural machine translation systems is commonly evaluated in terms of BLEU. However, due to its reliance on target language properties and generation, the BLEU metric does not allow an assessment of which translation directions are more difficult to model. In this paper, we propose cross-mutual information (XMI): an asymmetric information-theoretic metric of machine translation difficulty that exploits the probabilistic nature of most neural machine translation models. XMI allows us to better evaluate the difficulty of translating text into the target language while controlling for the difficulty of the target-side generation component independent of the translation task. We then present the first systematic and controlled study of cross-lingual translation difficulties using modern neural translation systems. Code for replicating our experiments is available online at https://github.com/e-bug/nmt-difficulty. Emanuele Bugliarello, Sabrina J. Mielke, Antonios Anastasopoulos, Ryan Cotterell, Naoaki Okazaki |
ACL | 5 |
| 2020 | Enhancing Machine Translation with Dependency-Aware Self-AttentionabstractMost neural machine translation models only rely on pairs of parallel sentences, assuming syntactic information is automatically learned by an attention mechanism.In this work, we investigate different approaches to incorporate syntactic knowledge in the Transformer model and also propose a novel, parameter-free, dependency-aware self-attention mechanism that improves its translation quality, especially for long sentences and in low-resource scenarios.We show the efficacy of each approach on WMT English↔German and English→Turkish, and WAT English→Japanese translation tasks. Emanuele Bugliarello, Naoaki Okazaki |
ACL | 2 |
| 2020 | Improving Truthfulness of Headline GenerationabstractMost studies on abstractive summarization report ROUGE scores between system and reference summaries.However, we have a concern about the truthfulness of generated summaries: whether all facts of a generated summary are mentioned in the source text.This paper explores improving the truthfulness in headline generation on two popular datasets.Analyzing headlines generated by the stateof-the-art encoder-decoder model, we show that the model sometimes generates untruthful headlines.We conjecture that one of the reasons lies in untruthful supervision data used for training the model.In order to quantify the truthfulness of article-headline pairs, we consider the textual entailment of whether an article entails its headline.After confirming quite a few untruthful instances in the datasets, this study hypothesizes that removing untruthful instances from the supervision data may remedy the problem of the untruthful behaviors of the model.Building a binary classifier that predicts an entailment relation between an article and its headline, we filter out untruthful instances from the supervision data.Experimental results demonstrate that the headline generation model trained on filtered supervision data shows no clear difference in ROUGE scores but remarkable improvements in automatic and manual evaluations of the generated headlines. Kazuki Matsumaru, Sho Takase, Naoaki Okazaki |
ACL | 3 |
| 2020 | Image Caption Generation for News ArticlesabstractIn this paper, we address the task of news-image captioning, which generates a description of an image given the image and its article body as input.This task is more challenging than the conventional image captioning, because it requires a joint understanding of image and text.We present a Transformer model that integrates text and image modalities and attends to textual features from visual features in generating a caption.Experiments based on automatic evaluation metrics and human evaluation show that an article text provides primary information to reproduce news-image captions written by journalists.The results also demonstrate that the proposed model outperforms the state-of-the-art model.In addition, we also confirm that visual features contribute to improving the quality of news-image captions. Zhishen Yang, Naoaki Okazaki |
COLING | 2 |
| 2020 | PatchBERT: Just-in-Time, Out-of-Vocabulary PatchingabstractLarge scale pre-trained language models have shown groundbreaking performance improvements for transfer learning in the domain of natural language processing.In our paper, we study a pre-trained multilingual BERT model and analyze the OOV rate on downstream tasks, how it introduces information loss, and as a side-effect, obstructs the potential of the underlying model.We then propose multiple approaches for mitigation and demonstrate that it improves performance with the same parameter count when combined with finetuning. Sangwhan Moon, Naoaki Okazaki |
EMNLP (1) | 2 |
| 2020 | Jamo Pair Encoding: Subcharacter Representation-based Extreme Korean Vocabulary Compression for Efficient Subword TokenizationabstractIn the context of multilingual language model pre-training, vocabulary size for languages with a broad set of potential characters is an unsolved problem. We propose two algorithms applicable in any unsupervised multilingual pre-training task, increasing the elasticity of budget required for building the vocabulary in Byte-Pair Encoding inspired tokenizers, significantly reducing the cost of supporting Korean in a multilingual model. Sangwhan Moon, Naoaki Okazaki |
LREC | 2 |
| 2020 | Evaluation Dataset for Zero Pronoun in Japanese to English TranslationabstractIn natural language, we often omit some words that are easily understandable from the context. In particular, pronouns of subject, object, and possessive cases are often omitted in Japanese; these are known as zero pronouns. In translation from Japanese to other languages, we need to find a correct antecedent for each zero pronoun to generate a correct and coherent translation. However, it is difficult for conventional automatic evaluation metrics (e.g., BLEU) to focus on the success of zero pronoun resolution. Therefore, we present a hand-crafted dataset to evaluate whether translation models can resolve the zero pronoun problems in Japanese to English translations. We manually and statistically validate that our dataset can effectively evaluate the correctness of the antecedents selected in translations. Through the translation experiments using our dataset, we reveal shortcomings of an existing context-aware neural machine translation model. Sho Shimazu, Sho Takase, Toshiaki Nakazawa, Naoaki Okazaki |
LREC | 4 |
| 2019 | Learning to Select, Track, and Generate for Data-to-TextabstractHayate Iso, Yui Uehara, Tatsuya Ishigaki, Hiroshi Noji, Eiji Aramaki, Ichiro Kobayashi, Yusuke Miyao, Naoaki Okazaki, Hiroya Takamura. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Hayate Iso, Yui Uehara, Tatsuya Ishigaki, Hiroshi Noji, Eiji Aramaki, Ichiro Kobayashi 0001, Yusuke Miyao, Naoaki Okazaki, Hiroya Takamura |
ACL (1) | 8 |
| 2019 | A Large-Scale Multi-Length Headline Corpus for Analyzing Length-Constrained Headline Generation Model EvaluationabstractYuta Hitomi, Yuya Taguchi, Hideaki Tamori, Ko Kikuta, Jiro Nishitoba, Naoaki Okazaki, Kentaro Inui, Manabu Okumura. Proceedings of the 12th International Conference on Natural Language Generation. 2019. Yuta Hitomi, Yuya Taguchi, Hideaki Tamori, Ko Kikuta, Jiro Nishitoba, Naoaki Okazaki, Kentaro Inui, Manabu Okumura |
INLG | 6 |
| 2019 | Neural Question Generation using Interrogative PhrasesabstractQuestion Generation (QG) is the task of generating questions from a given passage.One of the key requirements of QG is to generate a question such that it results in a target answer.Previous works used a target answer to obtain a desired question.However, we also want to specify how to ask questions and improve the quality of generated questions.In this study, we explore the use of interrogative phrases as additional sources to control QG.By providing interrogative phrases, we expect that QG can generate a more reliable sequence of words subsequent to an interrogative phrase.We present a baseline sequenceto-sequence model with the attention, copy, and coverage mechanisms, and show that the simple baseline achieves state-of-the-art performance.The experiments demonstrate that interrogative phrases contribute to improving the performance of QG.In addition, we report the superiority of using interrogative phrases in human evaluation.Finally, we show that a question answering system can provide target answers more correctly when the questions are generated with interrogative phrases. 1 Yuichi Sasazawa, Sho Takase, Naoaki Okazaki |
INLG | 3 |
| 2018 | Predicting Stances from Social Media Posts using Factorization MachinesabstractSocial media provide platforms to express, discuss, and shape opinions about events and issues in the real world. An important step to analyze the discussions on social media and to assist in healthy decision-making is stance detection. This paper presents an approach to detect the stance of a user toward a topic based on their stances toward other topics and the social media posts of the user. We apply factorization machines, a widely used method in item recommendation, to model user preferences toward topics from the social media data. The experimental results demonstrate that users’ posts are useful to model topic preferences and therefore predict stances of silent users. Akira Sasaki, Kazuaki Hanawa, Naoaki Okazaki, Kentaro Inui |
COLING | 3 |
| 2018 | Incorporating Semantic Attention in Video Description Generation
Natsuda Laokulrat, Naoaki Okazaki, Hideki Nakayama |
LREC | 2 |
| 2018 | PoB: Toward Reasoning Patterns of Beauty in Image DataabstractAiming to develop of computational grammar system for visual information, we design a 4-tier framework that consists of four levels of 'visual grammar of images.' As a first step of realization, we propose a new dataset, named the PoB dataset, in which each image is annotated with multiple labels of armature patterns that compose the pictorial scene. The PoB dataset includes of a 10,000-painting dataset for art and a 4,959-image dataset for photography. In this paper, we discuss the consistency analysis of our dataset and its applicability. We also demonstrate how the armature patterns in the PoB dataset are useful in assessing aesthetic quality of images, and how well a deep learning algorithm can recognize these patterns. This paper seeks to set a new direction in image understanding with a more holistic approach beyond discrete objects and in aesthetic reasoning with a more interpretative way. Diep Thi Ngoc Nguyen, Hideki Nakayama, Naoaki Okazaki, Tatsuya Sakaeda |
ACM Multimedia | 3 |
| 2018 | Multi-dialect Neural Machine Translation and Dialectometry
Kaori Abe, Yuichiroh Matsubayashi, Naoaki Okazaki, Kentaro Inui |
PACLIC | 3 |
| 2018 | Reducing Odd Generation from Neural Headline Generation
Shun Kiyono, Sho Takase, Jun Suzuki 0001, Naoaki Okazaki, Kentaro Inui, Masaaki Nagata |
PACLIC | 4 |
| 2017 | Other Topics You May Also Agree or Disagree: Modeling Inter-Topic Preferences using Tweets and Matrix FactorizationabstractWe present in this paper our approach for modeling inter-topic preferences of Twitter users: for example, those who agree with the Trans-Pacific Partnership (TPP) also agree with free trade.This kind of knowledge is useful not only for stance detection across multiple topics but also for various real-world applications including public opinion surveys, electoral predictions, electoral campaigns, and online debates.In order to extract users' preferences on Twitter, we design linguistic patterns in which people agree and disagree about specific topics (e.g., "A is completely wrong").By applying these linguistic patterns to a collection of tweets, we extract statements agreeing and disagreeing with various topics.Inspired by previous work on item recommendation, we formalize the task of modeling intertopic preferences as matrix factorization: representing users' preferences as a usertopic matrix and mapping both users and topics onto a latent feature space that abstracts the preferences.Our experimental results demonstrate both that our proposed approach is useful in predicting missing preferences of users and that the latent vector representations of topics successfully encode inter-topic preferences. Akira Sasaki, Kazuaki Hanawa, Naoaki Okazaki, Kentaro Inui |
ACL (1) | 3 |
| 2017 | Monitoring Geographical Entities with Temporal Awareness in Tweets
Koji Matsuda, Mizuki Sango, Naoaki Okazaki, Kentaro Inui |
CICLing (2) | 3 |
| 2017 | Learning Co-Substructures by Kernel Dependence MaximizationabstractModeling associations between items in a dataset is a problem that is frequently encountered in data and knowledge mining research. Most previous studies have simply applied a predefined fixed pattern for extracting the substructure of each item pair and then analyzed the associations between these substructures. Using such fixed patterns may not, however, capture the significant association. We, therefore, propose the novel machine learning task of extracting a strongly associated substructure pair (co-substructure) from each input item pair. We call this task dependent co-substructure extraction (DCSE), and formalize it as a dependence maximization problem. Then, we discuss critical issues with this task: the data sparsity problem and a huge search space. To address the data sparsity problem, we adopt the Hilbert--Schmidt independence criterion as an objective function. To improve search efficiency, we adopt the Metropolis--Hastings algorithm. We report the results of empirical evaluations, in which the proposed method is applied for acquiring and predicting narrative event pairs, an active task in the field of natural language processing. Sho Yokoi, Daichi Mochihashi, Naoaki Okazaki, Kentaro Inui |
IJCAI | 4 |
| 2017 | A Neural Language Model for Dynamically Representing the Meanings of Unknown Words and Entities in a DiscourseabstractThis study addresses the problem of identifying the meaning of unknown words or entities in a discourse with respect to the word embedding approaches used in neural language models. We proposed a method for on-the-fly construction and exploitation of word embeddings in both the input and output layers of a neural model by tracking contexts. This extends the dynamic entity representation used in Kobayashi et al. (2016) and incorporates a copy mechanism proposed independently by Gu et al. (2016) and Gulcehre et al. (2016). In addition, we construct a new task and dataset called Anonymized Language Modeling for evaluating the ability to capture word meanings while reading. Experiments conducted using our novel dataset show that the proposed variant of RNN language model outperformed the baseline model. Furthermore, the experiments also demonstrate that dynamic updates of an output layer help a model predict reappearing entities, whereas those of an input layer are effective to predict words following reappearing entities. Sosuke Kobayashi, Naoaki Okazaki, Kentaro Inui |
IJCNLP(1) | 2 |
| 2017 | A Crowdsourcing Approach for Annotating Causal Relation Instances in Wikipedia
Kazuaki Hanawa, Akira Sasaki, Naoaki Okazaki, Kentaro Inui |
PACLIC | 3 |
| 2017 | The mechanism of additive compositionabstractAdditive composition (Foltz et al. in Discourse Process 15:285–307, 1998 ; Landauer and Dumais in Psychol Rev 104(2):211, 1997 ; Mitchell and Lapata in Cognit Sci 34(8):1388–1429, 2010 ) is a widely used method for computing meanings of phrases, which takes the average of vector representations of the constituent words. In this article, we prove an upper bound for the bias of additive composition, which is the first theoretical analysis on compositional frameworks from a machine learning point of view. The bound is written in terms of collocation strength; we prove that the more exclusively two successive words tend to occur together, the more accurate one can guarantee their additive composition as an approximation to the natural phrase vector. Our proof relies on properties of natural language data that are empirically verified, and can be theoretically derived from an assumption that the data is generated from a Hierarchical Pitman–Yor Process. The theory endorses additive composition as a reasonable operation for calculating meanings of phrases, and suggests ways to improve additive compositionality, including: transforming entries of distributional word vectors by a function that meets a specific condition, constructing a novel type of vector representations to make additive composition sensitive to word order, and utilizing singular value decomposition to train word vectors. Naoaki Okazaki, Kentaro Inui |
Mach. Learn. | 2 |
| 2016 | Learning to Describe E-Commerce Images from Noisy Online Data
Takuya Yashima, Naoaki Okazaki, Kentaro Inui, Kota Yamaguchi, Takayuki Okatani |
ACCV (5) | 2 |
| 2016 | Composing Distributed Representations of Relational PatternsabstractLearning distributed representations for relation instances is a central technique in downstream NLP applications. In order to address semantic modeling of relational patterns, this paper constructs a new dataset that provides multiple similarity ratings for every pair of relational patterns on the existing dataset. In addition, we conduct a comparative study of different encoders including additive composition, RNN, LSTM, and GRU for composing distributed representations of relational patterns. We also present Gated Additive Composition, which is an enhancement of additive composition with the gating mechanism. Experiments show that the new dataset does not only enable detailed analyses of the different encoders, but also provides a gauge to predict successes of distributed representations of relational patterns in the relation classification task. Sho Takase, Naoaki Okazaki, Kentaro Inui |
ACL (1) | 2 |
| 2016 | Learning Semantically and Additively Compositional Distributional RepresentationsabstractThis paper connects a vector-based composition model to a formal semantics, the Dependency-based Compositional Semantics (DCS).We show theoretical evidence that the vector compositions in our model conform to the logic of DCS.Experimentally, we show that vector-based composition brings a strong ability to calculate similar phrases as similar vectors, achieving near state-of-the-art on a wide range of phrase similarity tasks and relation classification; meanwhile, DCS can guide building vectors for structured queries that can be directly executed.We evaluate this utility on sentence completion task and report a new state-of-the-art. Naoaki Okazaki, Kentaro Inui |
ACL (1) | 2 |
| 2016 | Modeling Context-sensitive Selectional Preference with Distributed RepresentationsabstractThis paper proposes a novel problem setting of selectional preference (SP) between a predicate and its arguments, called as context-sensitive SP (CSP). CSP models the narrative consistency between the predicate and preceding contexts of its arguments, in addition to the conventional SP based on semantic types. Furthermore, we present a novel CSP model that extends the neural SP model (Van de Cruys, 2014) to incorporate contextual information into the distributed representations of arguments. Experimental results demonstrate that the proposed CSP model successfully learns CSP and outperforms the conventional SP model in coreference cluster ranking. Naoya Inoue, Yuichiroh Matsubayashi, Masayuki Ono, Naoaki Okazaki, Kentaro Inui |
COLING | 4 |
| 2016 | Generating Video Description using Sequence-to-sequence Model with Temporal AttentionabstractAutomatic video description generation has recently been getting attention after rapid advancement in image caption generation. Automatically generating description for a video is more challenging than for an image due to its temporal dynamics of frames. Most of the work relied on Recurrent Neural Network (RNN) and recently attentional mechanisms have also been applied to make the model learn to focus on some frames of the video while generating each word in a describing sentence. In this paper, we focus on a sequence-to-sequence approach with temporal attention mechanism. We analyze and compare the results from different attention model configuration. By applying the temporal attention mechanism to the system, we can achieve a METEOR score of 0.310 on Microsoft Video Description dataset, which outperformed the state-of-the-art system so far. Natsuda Laokulrat, Sang Phan Le, Noriki Nishida, Raphael Shu, Yo Ehara, Naoaki Okazaki, Yusuke Miyao, Hideki Nakayama |
COLING | 6 |
| 2016 | Modeling Discourse Segments in Lyrics Using Repeated PatternsabstractThis study proposes a computational model of the discourse segments in lyrics to understand and to model the structure of lyrics. To test our hypothesis that discourse segmentations in lyrics strongly correlate with repeated patterns, we conduct the first large-scale corpus study on discourse segments in lyrics. Next, we propose the task to automatically identify segment boundaries in lyrics and train a logistic regression model for the task with the repeated pattern and textual features. The results of our empirical experiments illustrate the significance of capturing repeated patterns in predicting the boundaries of discourse segments in lyrics. Kento Watanabe, Yuichiroh Matsubayashi, Naho Orita, Naoaki Okazaki, Kentaro Inui, Satoru Fukayama, Tomoyasu Nakano, Jordan B. L. Smith, Masataka Goto |
COLING | 4 |
| 2016 | Neural Headline Generation on Abstract Meaning RepresentationabstractNeural network-based encoder-decoder models are among recent attractive methodologies for tackling natural language generation tasks.This paper investigates the usefulness of structural syntactic and semantic information additionally incorporated in a baseline neural attention-based model.We encode results obtained from an abstract meaning representation (AMR) parser using a modified version of Tree-LSTM.Our proposed attention-based AMR encoder-decoder model improves headline generation benchmarks compared with the baseline neural attention-based model. Sho Takase, Jun Suzuki 0001, Naoaki Okazaki, Tsutomu Hirao, Masaaki Nagata |
EMNLP | 3 |
| 2016 | Dynamic Entity Representation with Max-pooling Improves Machine ReadingabstractSosuke Kobayashi, Ran Tian, Naoaki Okazaki, Kentaro Inui. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Sosuke Kobayashi, Naoaki Okazaki, Kentaro Inui |
HLT-NAACL | 3 |
| 2016 | Recognizing Open-Vocabulary Relations between Objects in Images
Masayasu Muraoka, Sumit Maharjan, Masaki Saito, Kota Yamaguchi, Naoaki Okazaki, Takayuki Okatani, Kentaro Inui |
PACLIC | 5 |
| 2016 | Neural Joint Learning for Classifying Wikipedia Articles into Fine-grained Named Entity Types
Masatoshi Suzuki, Koji Matsuda, Satoshi Sekine, Naoaki Okazaki, Kentaro Inui |
PACLIC | 4 |
| 2016 | Toward the automatic extraction of knowledge of usable goods
Mei Uemura, Naho Orita, Naoaki Okazaki, Kentaro Inui |
PACLIC | 3 |
| 2016 | Stance Classification by Recognizing Related Events about TargetsabstractRecently, many people express their opinions using social networking services such as Twitter and Facebook. Each opinion has a stance related to something such as product, service, and politics. The task of detecting a stance is known as sentiment analysis, reputation mining, and stance detection. A popular approach for stance detection uses sentiment polarity towards a target in a text. This approach is known as targeted sentiment analysis. If a target appears in text, the detecting stance based on targeted sentiment polarity would work well. However, how can we detect stance towards an event? (e.g. "I cannot understand why man can marry only with a woman", "The problem of low birth rate becomes more severe" to the event "Allowing same-sex marriage"). To detect these stances, it is necessary to recognize a situation in which the event occurs or does not occur. To classify texts including these phenomena, we propose a classification method based on machine learning considering PRIOR-SITUATION and EFFECT. Akira Sasaki, Junta Mizuno, Naoaki Okazaki, Kentaro Inui |
WI | 3 |
| 2016 | Fine-Grained Named Entity Classification with Wikipedia Article VectorsabstractThis paper addresses the task of assigning multiple labels of fine-grained named entity (NE) types to Wikipedia articles. To address the sparseness of the input feature space, which is salient particularly in fine-grained type classification, we propose to learn article vectors (i.e. entity embeddings) from hypertext structure of Wikipedia using a Skip-gram model and incorporate them into the input feature set. To conduct large-scale practical experiments, we created a new dataset containing over 22,000 manually labeled instances. The results of our experiments show that our idea gained statistically significant improvements in classification results. Masatoshi Suzuki, Koji Matsuda, Satoshi Sekine, Naoaki Okazaki, Kentaro Inui |
WI | 4 |
| 2016 | Modeling semantic compositionality of relational patternsabstractVector representation is a common approach for expressing the meaning of a relational pattern. Most previous work obtained a vector of a relational pattern based on the distribution of its context words (e.g., arguments of the relational pattern), regarding the pattern as a single ‘word’. However, this approach suffers from the data sparseness problem, because relational patterns are productive, i.e., produced by combinations of words. To address this problem, we propose a novel method for computing the meaning of a relational pattern based on the semantic compositionality of constituent words. We extend the Skip-gram model (Mikolov et al., 2013) to handle semantic compositions of relational patterns using recursive neural networks. The experimental results show the superiority of the proposed method for modeling the meanings of relational patterns, and demonstrate the contribution of this work to the task of relation extraction. Sho Takase, Naoaki Okazaki, Kentaro Inui |
Eng. Appl. Artif. Intell. | 2 |
| 2015 | Who caught a cold ? - Identifying the subject of a symptomabstractShin Kanouchi, Mamoru Komachi, Naoaki Okazaki, Eiji Aramaki, Hiroshi Ishikawa. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Shin Kanouchi, Mamoru Komachi, Naoaki Okazaki, Eiji Aramaki, Hiroshi Ishikawa 0004 |
ACL (1) | 3 |
| 2015 | Reducing Lexical Features in Parsing by Word Embeddings
Hiroya Komatsu, Naoaki Okazaki, Kentaro Inui |
PACLIC | 3 |
| 2015 | Fast and Large-scale Unsupervised Relation Extraction
Sho Takase, Naoaki Okazaki, Kentaro Inui |
PACLIC | 2 |
| 2014 | Finding The Best Model Among Representative Compositional Models
Masayasu Muraoka, Sonse Shimaoka, Kazeto Yamamoto, Yotaro Watanabe, Naoaki Okazaki, Kentaro Inui |
PACLIC | 5 |
| 2013 | Is a 204 cm Man Tall or Small ? Acquisition of Numerical Common Sense from the Web
Katsuma Narisawa, Yotaro Watanabe, Junta Mizuno, Naoaki Okazaki, Kentaro Inui |
ACL (1) | 4 |
| 2013 | Evidence in Automatic Error Correction Improves Learners' English Skill
Jiro Umezawa, Junta Mizuno, Naoaki Okazaki, Kentaro Inui |
CICLing (2) | 3 |
| 2013 | Discriminative Learning of First-Order Weighted Abduction from Partial Discourse Explanations
Kazeto Yamamoto, Naoya Inoue, Yotaro Watanabe, Naoaki Okazaki, Kentaro Inui |
CICLing (1) | 4 |
| 2013 | Inducing Context Gazetteers from Encyclopedic Databases for Named Entity Recognition
Hancheol Cho, Naoaki Okazaki, Kentaro Inui |
PAKDD (1) | 2 |
| 2013 | Named entity recognition with multiple segment representations
Hancheol Cho, Naoaki Okazaki, Makoto Miwa, Jun'ichi Tsujii |
Inf. Process. Manag. | 2 |
| 2013 | Learning Abbreviations from Chinese and English Terms by Modeling Non-Local InformationabstractThe present article describes a robust approach for abbreviating terms. First, in order to incorporate non-local information into abbreviation generation tasks, we present both implicit and explicit solutions: the latent variable model and the label encoding with global information. Although the two approaches compete with one another, we find they are also highly complementary. We propose a combination of the two approaches, and we will show the proposed method outperforms all of the existing methods on abbreviation generation datasets. In order to reduce computational complexity of learning non-local information, we further present an online training method, which can arrive the objective optimum with accelerated training speed. We used a Chinese newswire dataset and a English biomedical dataset for experiments. Experiments revealed that the proposed abbreviation generator with non-local information achieved the best results for both the Chinese and English languages. Xu Sun 0001, Naoaki Okazaki, Jun'ichi Tsujii, Houfeng Wang |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2012 | A Latent Discriminative Model for Compositional Entailment Relation Recognition using Natural Logic
Yotaro Watanabe, Junta Mizuno, Eric Nichols, Naoaki Okazaki, Kentaro Inui |
COLING | 4 |
| 2012 | Set Expansion using Sibling Relations between Semantic Categories
Sho Takase, Naoaki Okazaki, Kentaro Inui |
PACLIC | 2 |
| 2012 | A preference learning approach to sentence ordering for multi-document summarization
Danushka Bollegala, Naoaki Okazaki, Mitsuru Ishizuka |
Inf. Sci. | 2 |
| 2012 | Leveraging Diverse Lexical Resources for Textual Entailment RecognitionabstractSince the problem of textual entailment recognition requires capturing semantic relations between diverse expressions of language, linguistic and world knowledge play an important role. In this article, we explore the effectiveness of different types of currently available resources including synonyms, antonyms, hypernym-hyponym relations, and lexical entailment relations for the task of textual entailment recognition. In order to do so, we develop an entailment relation recognition system which utilizes diverse linguistic analyses and resources to align the linguistic units in a pair of texts and identifies entailment relations based on these alignments. We use the Japanese subset of the NTCIR-9 RITE-1 dataset for evaluation and error analysis, conducting ablation testing and evaluation on hand-crafted alignment gold standard data to evaluate the contribution of individual resources. Error analysis shows that existing knowledge sources are effective for RTE, but that their coverage is limited, especially for domain-specific and other low-frequency expressions. To increase alignment coverage on such expressions, we propose a method of alignment inference that uses syntactic and semantic dependency information to identify likely alignments without relying on external resources. Evaluation adding alignment inference to a system using all available knowledge sources shows improvements in both precision and recall of entailment relation recognition. Yotaro Watanabe, Junta Mizuno, Eric Nichols, Katsuma Narisawa, Keita Nabeshima, Naoaki Okazaki, Kentaro Inui |
ACM Trans. Asian Lang. Inf. Process. | 6 |
| 2011 | Fast Newton-CG Method for Batch Learning of Conditional Random FieldsabstractWe propose a fast batch learning method for linear-chain Conditional Random Fields (CRFs) based on Newton-CG methods. Newton-CG methods are a variant of Newton method for high-dimensional problems. They only require the Hessian-vector products instead of the full Hessian matrices. To speed up Newton-CG methods for the CRF learning, we derive a novel dynamic programming procedure for the Hessian-vector products of the CRF objective function. The proposed procedure can reuse the byproducts of the time-consuming gradient computation for the Hessian-vector products to drastically reduce the total computation time of the Newton-CG methods. In experiments with tasks in natural language processing, the proposed method outperforms a conventional quasi-Newton method. Remarkably, the proposed method is competitive with online learning algorithms that are fast but unstable. Yuta Tsuboi, Yuya Unno, Hisashi Kashima, Naoaki Okazaki |
AAAI | 4 |
| 2011 | BioCreative III interactive task: an overviewabstractBACKGROUND: The BioCreative challenge evaluation is a community-wide effort for evaluating text mining and information extraction systems applied to the biological domain. The biocurator community, as an active user of biomedical literature, provides a diverse and engaged end user group for text mining tools. Earlier BioCreative challenges involved many text mining teams in developing basic capabilities relevant to biological curation, but they did not address the issues of system usage, insertion into the workflow and adoption by curators. Thus in BioCreative III (BC-III), the InterActive Task (IAT) was introduced to address the utility and usability of text mining tools for real-life biocuration tasks. To support the aims of the IAT in BC-III, involvement of both developers and end users was solicited, and the development of a user interface to address the tasks interactively was requested. RESULTS: A User Advisory Group (UAG) actively participated in the IAT design and assessment. The task focused on gene normalization (identifying gene mentions in the article and linking these genes to standard database identifiers), gene ranking based on the overall importance of each gene mentioned in the article, and gene-oriented document retrieval (identifying full text papers relevant to a selected gene). Six systems participated and all processed and displayed the same set of articles. The articles were selected based on content known to be problematic for curation, such as ambiguity of gene names, coverage of multiple genes and species, or introduction of a new gene name. Members of the UAG curated three articles for training and assessment purposes, and each member was assigned a system to review. A questionnaire related to the interface usability and task performance (as measured by precision and recall) was answered after systems were used to curate articles. Although the limited number of articles analyzed and users involved in the IAT experiment precluded rigorous quantitative analysis of the results, a qualitative analysis provided valuable insight into some of the problems encountered by users when using the systems. The overall assessment indicates that the system usability features appealed to most users, but the system performance was suboptimal (mainly due to low accuracy in gene normalization). Some of the issues included failure of species identification and gene name ambiguity in the gene normalization task leading to an extensive list of gene identifiers to review, which, in some cases, did not contain the relevant genes. The document retrieval suffered from the same shortfalls. The UAG favored achieving high performance (measured by precision and recall), but strongly recommended the addition of features that facilitate the identification of correct gene and its identifier, such as contextual information to assist in disambiguation. DISCUSSION: The IAT was an informative exercise that advanced the dialog between curators and developers and increased the appreciation of challenges faced by each group. A major conclusion was that the intended users should be actively involved in every phase of software development, and this will be strongly encouraged in future tasks. The IAT Task provides the first steps toward the definition of metrics and functional requirements that are necessary for designing a formal evaluation of interactive curation systems in the BioCreative IV challenge. Cecilia N. Arighi, Phoebe M. Roberts, Shashank Agarwal, Sanmitra Bhattacharya, Gianni Cesareni, Andrew Chatr-aryamontri, Simon Clematide, Pascale Gaudet, Michelle G. Giglio, Ian Harrow, Eva Huala, Martin Krallinger, Ulf Leser, Zhiyong Lu, Lois J. Maltais, Naoaki Okazaki, Livia Perfetto, Fabio Rinaldi 0001, Rune Sætre, David Salgado, Padmini Srinivasan, Philippe Thomas 0002, Luca Toldo, Lynette Hirschman, Cathy H. Wu |
BMC Bioinform. | 18 |
| 2011 | The gene normalization task in BioCreative IIIabstractBACKGROUND: We report the Gene Normalization (GN) challenge in BioCreative III where participating teams were asked to return a ranked list of identifiers of the genes detected in full-text articles. For training, 32 fully and 500 partially annotated articles were prepared. A total of 507 articles were selected as the test set. Due to the high annotation cost, it was not feasible to obtain gold-standard human annotations for all test articles. Instead, we developed an Expectation Maximization (EM) algorithm approach for choosing a small number of test articles for manual annotation that were most capable of differentiating team performance. Moreover, the same algorithm was subsequently used for inferring ground truth based solely on team submissions. We report team performance on both gold standard and inferred ground truth using a newly proposed metric called Threshold Average Precision (TAP-k). RESULTS: We received a total of 37 runs from 14 different teams for the task. When evaluated using the gold-standard annotations of the 50 articles, the highest TAP-k scores were 0.3297 (k=5), 0.3538 (k=10), and 0.3535 (k=20), respectively. Higher TAP-k scores of 0.4916 (k=5, 10, 20) were observed when evaluated using the inferred ground truth over the full test set. When combining team results using machine learning, the best composite system achieved TAP-k scores of 0.3707 (k=5), 0.4311 (k=10), and 0.4477 (k=20) on the gold standard, representing improvements of 12.4%, 21.8%, and 26.6% over the best team results, respectively. CONCLUSIONS: By using full text and being species non-specific, the GN task in BioCreative III has moved closer to a real literature curation task than similar tasks in the past and presents additional challenges for the text mining community, as revealed in the overall team results. By evaluating teams using the gold standard, we show that the EM algorithm allows team submissions to be differentiated while keeping the manual annotation effort feasible. Using the inferred ground truth we show measures of comparative performance between teams. Finally, by comparing team rankings on gold standard vs. inferred ground truth, we further demonstrate that the inferred ground truth is as effective as the gold standard for detecting good team performance. Zhiyong Lu, Hung-Yu Kao, Chih-Hsuan Wei, Minlie Huang, Jingchen Liu, Cheng-Ju Kuo, Chun-Nan Hsu, Richard Tzong-Han Tsai, Hong-Jie Dai, Naoaki Okazaki, Hancheol Cho, Martin Gerner, Illés Solt, Shashank Agarwal, Dina Vishnyakova, Patrick Ruch, Martin Romacker, Fabio Rinaldi 0001, Sanmitra Bhattacharya, Padmini Srinivasan, Manabu Torii, Sérgio Matos, David Campos 0001, Karin Verspoor, Kevin M. Livingston, W. John Wilbur |
BMC Bioinform. | 10 |
| 2010 | Simple and Efficient Algorithm for Approximate Dictionary Matching
Naoaki Okazaki, Jun'ichi Tsujii |
COLING | 1 |
| 2010 | Building a high-quality sense inventory for improved abbreviation disambiguationabstractMOTIVATION: The ultimate goal of abbreviation management is to disambiguate every occurrence of an abbreviation into its expanded form (concept or sense). To collect expanded forms for abbreviations, previous studies have recognized abbreviations and their expanded forms in parenthetical expressions of bio-medical texts. However, expanded forms extracted by abbreviation recognition are mixtures of concepts/senses and their term variations. Consequently, a list of expanded forms should be structured into a sense inventory, which provides possible concepts or senses for abbreviation disambiguation. RESULTS: A sense inventory is a key to robust management of abbreviations. Therefore, we present a supervised approach for clustering expanded forms. The experimental result reports 0.915 F1 score in clustering expanded forms. We then investigate the possibility of conflicts of protein and gene names with abbreviations. Finally, an experiment of abbreviation disambiguation on the sense inventory yielded 0.984 accuracy and 0.986 F1 score using the dataset obtained from MEDLINE abstracts. AVAILABILITY: The sense inventory and disambiguator of abbreviations are accessible at http://www.nactem.ac.uk/software/acromine/ and http://www.nactem.ac.uk/software/acromine_disambiguation/. Naoaki Okazaki, Sophia Ananiadou, Jun'ichi Tsujii |
Bioinform. | 1 |
| 2010 | Medie and Info-pubmed: 2010 updateabstractIn the recent decades, high-throughput screening methods were established, bringing forth major breakthroughs in the fields of molecular biology and biomedicine. Since researchers in these fields need to interpret an enormous quantity of data and the publication rates of scientific articles are exploding, demands on text mining technology are growing with each passing year. Tomoko Ohta, Takuya Matsuzaki, Naoaki Okazaki, Makoto Miwa, Rune Sætre, Sampo Pyysalo, Jun'ichi Tsujii |
BMC Bioinform. | 3 |
| 2010 | A bottom-up approach to sentence ordering for multi-document summarization
Danushka Bollegala, Naoaki Okazaki, Mitsuru Ishizuka |
Inf. Process. Manag. | 2 |
| 2009 | A Comparative Study on Generalization of Semantic Roles in FrameNet
Yuichiroh Matsubayashi, Naoaki Okazaki, Jun'ichi Tsujii |
ACL/IJCNLP | 2 |
| 2009 | Robust Approach to Abbreviating Terms: A Discriminative Latent Variable Model with Global Information
Xu Sun 0001, Naoaki Okazaki, Jun'ichi Tsujii |
ACL/IJCNLP | 2 |
| 2009 | Unsupervised Relation Extraction by Mining Wikipedia Texts Using Information from the Web
Yulan Yan, Naoaki Okazaki, Yutaka Matsuo, Zhenglu Yang, Mitsuru Ishizuka |
ACL/IJCNLP | 2 |
| 2009 | Opinion classification with tree kernel SVM using linguistic modality analysisabstractWe propose a method for classifying opinions which captures the role of linguistic modalities in the sentence. We use features than simple bag-of-words or opinion-holding predicates. The method is based on a machine learning and utilizes opinion-holding predicates and linguistic modalities as features. Two different detectors help to classify the opinions: the opinion-holding predicate detector and the modality detector. An opinion in the target is first parsed into a dependency structure, and then the opinion-holding predicates and modalities stick onto the leaf nodes of the dependency tree. The whole tree is regarded as input features of the opinion, and it becomes the input of tree kernel support vector machines. We have applied method to opinions in Japanese about television programs, and have confirmed the effectiveness of the method against conventional bag-of-words features, or against simple opinion-holding predicates features Takeshi S. Kobayakawa, Tadashi Kumano, Hideki Tanaka, Naoaki Okazaki, Jin-Dong Kim, Jun'ichi Tsujii |
CIKM | 4 |
| 2009 | Semi-Supervised Lexicon Mining from Parenthetical Expressions in Monolingual Web Pages
Xianchao Wu, Naoaki Okazaki, Jun'ichi Tsujii |
HLT-NAACL | 2 |
| 2009 | A Chinese-Japanese Lexical Machine Translation through a Pivot LanguageabstractThe bilingual lexicon is an expensive but critical resource for multilingual applications in natural language processing. This article proposes an integrated framework for building a bilingual lexicon between the Chinese and Japanese languages. Since the language pair Chinese-Japanese does not include English, which is a central language of the world, few large-scale bilingual resources between Chinese and Japanese have been constructed. One solution to alleviate this problem is to build a Chinese-Japanese bilingual lexicon through English as the pivot language. In addition to the pivotal approach, we can make use of the characteristics of Chinese and Japanese languages that use Han characters. We incorporate a translation model obtained from a small Chinese-Japanese lexicon and use the similarity of the hanzi and kanji characters by using the log-linear model. Our experimental results show that the use of the pivotal approach can improve the translation performance over the translation model built from a small Chinese-Japanese lexicon. The results also demonstrate that the similarity between the hanzi and kanji characters provides a positive effect for translating technical terms. Takashi Tsunakawa, Naoaki Okazaki, Jun'ichi Tsujii |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2008 | A Discriminative Alignment Model for Abbreviation Recognition
Naoaki Okazaki, Sophia Ananiadou, Jun'ichi Tsujii |
COLING | 1 |
| 2008 | A Discriminative Candidate Generator for String Transformations
Naoaki Okazaki, Yoshimasa Tsuruoka, Sophia Ananiadou, Jun'ichi Tsujii |
EMNLP | 1 |
| 2008 | Identifying Sections in Scientific Abstracts using Conditional Random Fields
Kenji Hirohata, Naoaki Okazaki, Sophia Ananiadou, Mitsuru Ishizuka |
IJCNLP | 2 |
| 2008 | A Discriminative Approach to Japanese Abbreviation Extraction
Naoaki Okazaki, Mitsuru Ishizuka, Jun'ichi Tsujii |
IJCNLP | 1 |
| 2008 | Connecting Text Mining and Pathways using the PathText Resource
Rune Sætre, Brian Kemper, Kanae Oda, Naoaki Okazaki, Yukiko Matsuoka, Norihiro Kikuchi, Hiroaki Kitano, Yoshimasa Tsuruoka, Sophia Ananiadou, Jun'ichi Tsujii |
LREC | 4 |
| 2008 | Building Bilingual Lexicons using Lexical Translation Probabilities via Pivot Languages
Takashi Tsunakawa, Naoaki Okazaki, Jun'ichi Tsujii |
LREC | 2 |
| 2008 | Kleio: a knowledge-enriched information retrieval system for biologyabstractKleio is an advanced information retrieval (IR) system developed at the UK National Centre for Text Mining (NaCTeM)1. The system offers textual and metadata searches across MEDLINE and provides enhanced searching functionality by leveraging terminology management technologies. Chikashi Nobata, Philip Cotter, Naoaki Okazaki, Brian Rea, Yutaka Sasaki, Yoshimasa Tsuruoka, Jun'ichi Tsujii, Sophia Ananiadou |
SIGIR | 3 |
| 2007 | Inferring Long-term User Properties Based on Users' Location History
Yutaka Matsuo, Naoaki Okazaki, Kiyoshi Izumi, Yoshiyuki Nakamura, Takuichi Nishimura, Kôiti Hasida, Hideyuki Nakashima |
IJCAI | 2 |
| 2006 | A Bottom-Up Approach to Sentence Ordering for Multi-Document SummarizationabstractOrdering information is a difficult but important task for applications generating natural-language text. We present a bottom-up approach to arranging sentences extracted for multi-document summarization. To capture the association and order of two textual segments (eg, sentences), we define four criteria, chronology, topical-closeness, precedence, and succession. These criteria are integrated into a criterion by a supervised learning approach. We repeatedly concatenate two textual segments into one segment based on the criterion until we obtain the overall segment with all sentences arranged. Our experimental results show a significant improvement over existing sentence ordering strategies. Danushka Bollegala, Naoaki Okazaki, Mitsuru Ishizuka |
ACL | 2 |
| 2006 | A Term Recognition Approach to Acronym Recognition
Naoaki Okazaki, Sophia Ananiadou |
ACL | 1 |
| 2006 | Towards a terminological resource for biomedical text mining
Goran Nenadic, Naoaki Okazaki, Sophia Ananiadou |
LREC | 2 |
| 2006 | Clustering acronyms in biomedical text for disambiguation
Naoaki Okazaki, Sophia Ananiadou |
LREC | 1 |
| 2006 | Building an abbreviation dictionary using a term recognition approachabstractMOTIVATION: Acronyms result from a highly productive type of term variation and trigger the need for an acronym dictionary to establish associations between acronyms and their expanded forms. RESULTS: We propose a novel method for recognizing acronym definitions in a text collection. Assuming a word sequence co-occurring frequently with a parenthetical expression to be a potential expanded form, our method identifies acronym definitions in a similar manner to the statistical term recognition task. Applied to the whole MEDLINE (7 811 582 abstracts), the implemented system extracted 886 755 acronym candidates and recognized 300 954 expanded forms in reasonable time. Our method outperformed base-line systems, achieving 99% precision and 82-95% recall on our evaluation corpus that roughly emulates the whole MEDLINE. AVAILABILITY AND SUPPLEMENTARY INFORMATION: The implementations and supplementary information are available at our web site: http://www.chokkan.org/research/acromine/ Naoaki Okazaki, Sophia Ananiadou |
Bioinform. | 1 |
| 2005 | A Machine Learning Approach to Sentence Ordering for Multidocument Summarization and Its Evaluation
Danushka Bollegala, Naoaki Okazaki, Mitsuru Ishizuka |
IJCNLP | 2 |
| 2005 | Improving chronological ordering of sentences extracted from multiple newspaper articlesabstractIt is necessary to determine a proper arrangement of extracted sentences to generate a well-organized summary from multiple documents. This paper describes our Multi-Document Summarization (MDS) system for TSC-3. It specifically addresses an approach to coherent sentence ordering for MDS. An impediment to the use of chronological ordering, which is widely used by conventional summarization system, is that it arranges sentences without considering the presupposed information of each sentence. We propose a method to improve chronological ordering by resolving precedent information of arranging sentences. Combining the refinement algorithm with topical segmentation and chronological ordering, we address our experiments and metrics to test the effectiveness of MDS tasks. Results demonstrate that the proposed method significantly improves chronological sentence ordering. At the end of the paper, we also report an outline/evaluation of important sentence extraction and redundant clause elimination integrated in our MDS system. Naoaki Okazaki, Yutaka Matsuo, Mitsuru Ishizuka |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2004 | Improving Chronological Sentence Ordering by Precedence Relation
Naoaki Okazaki, Yutaka Matsuo, Mitsuru Ishizuka |
COLING | 1 |
| 2004 | Coherent Arrangement of Sentences Extracted from Multiple Newspaper Articles
Naoaki Okazaki, Yutaka Matsuo, Mitsuru Ishizuka |
PRICAI | 1 |