VLDB 2026 Research / reviewers in the wild / expert
Sadao Kurohashi
dblp:42/2149
· DBLP profile ↗
178ranked-venue papers
5as first author
37since 2021 · last 2026
0000-0001-5398-8399ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 173 · 5 first-author · 35 since 2021Databases, data management, data science and information retrieval · 5Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Memorization, Emergence, and Explaining Reversal Failures: A Controlled Study of Relational Semantics in LLMsabstractYihua Zhu, Qianying Liu, Jiaxin Wang, Fei Cheng, Chaoran Liu, Akiko Aizawa, Sadao Kurohashi, Hidetoshi Shimodaira. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yihua Zhu 0002, Qianying Liu, Fei Cheng 0002, Akiko Aizawa, Sadao Kurohashi, Hidetoshi Shimodaira |
ACL (1) | 7 |
| 2026 | Building Effective Japanese Medical LLMs with an Open Recipe for Domain Adaptation through Continued Pre-training
Akiko Aizawa, Yuki Arase, Fei Cheng 0002, Teruhito Kanazawa, Daisuke Kawahara, Kazuma Kobayashi, Takashi Kodama, Sadao Kurohashi, Yusuke Oda, Tsuta Yuma, Zhishen Yang, Rio Yokota |
LREC | 11 |
| 2026 | Scaling LLM Reasoning from Minimal Labels: A Semi-Supervised Framework with a Lightweight Verifier
Keizo Kato, Chenhui Chu, Yugo Murawaki, Sadao Kurohashi |
LREC | 4 |
| 2026 | BIS Reasoning 1.0: The First Large-Scale Japanese Benchmark for Belief-Inconsistent Syllogistic ReasoningabstractWe present BIS Reasoning 1.0, the first large-scale Japanese dataset of syllogistic reasoning problems explicitly designed to evaluate belief-inconsistent reasoning in large language models (LLMs). Unlike prior resources such as NeuBAROCO and JFLD, which emphasize general or belief-aligned logic, BIS Reasoning 1.0 systematically introduces logically valid yet belief-inconsistent syllogisms to expose belief bias, the tendency to accept believable conclusions irrespective of validity. We benchmark a representative suite of cutting-edge models, including OpenAI GPT-5 variants, GPT-4o, Qwen, and prominent Japanese LLMs, under a uniform, zero-shot protocol. Reasoning-centric models achieve near-perfect accuracy on BIS Reasoning 1.0 (e.g., Qwen3-32B $\approx$99% and GPT-5-mini up to $\approx$99.7%), while GPT-4o attains around 80%. Earlier Japanese-specialized models underperform, often well below 60%, whereas the latest llm-jp-3.1-13b-instruct4 markedly improves to the mid-80% range. These results indicate that robustness to belief-inconsistent inputs is driven more by explicit reasoning optimization than by language specialization or scale alone. Our analysis further shows that even top-tier systems falter when logical validity conflicts with intuitive or factual beliefs, and that performance is sensitive to prompt design and inference-time reasoning effort. We discuss implications for safety-critical domains, including law, healthcare, and scientific literature, where strict logical fidelity must override intuitive belief to ensure reliability. Ha-Thanh Nguyen, Hideyuki Tachibana, Qianying Liu, Su Myat Noe, Koichi Takeda 0003, Sadao Kurohashi |
LREC | 7 |
| 2025 | SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language ModelsabstractZhen Wan, Chao-Han Huck Yang, Yahan Yu, Jinchuan Tian, Sheng Li, Ke Hu, Zhehuai Chen, Shinji Watanabe, Fei Cheng, Chenhui Chu, Sadao Kurohashi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chao-Han Huck Yang, Yahan Yu, Jinchuan Tian, Sheng Li 0010, Zhehuai Chen, Shinji Watanabe 0001, Fei Cheng 0002, Chenhui Chu, Sadao Kurohashi |
ACL (1) | 11 |
| 2025 | Causal Tree Extraction from Medical Case Reports: A Novel Task for Experts-like Text ComprehensionabstractExtracting causal relationships from a medical case report is essential for comprehending the case, particularly its diagnostic process.Since the diagnostic process is regarded as a bottom-up inference, causal relationships in cases naturally form a multi-layered tree structure.The existing tasks, such as medical relation extraction, are insufficient for capturing the causal relationships of an entire case, as they treat all relations equally without considering the hierarchical structure inherent in the diagnostic process.Thus, we propose a novel task, Causal Tree Extraction (CTE), which receives a case report and generates a causal tree with the primary disease as the root, providing an intuitive understanding of a case's diagnostic process.Subsequently, we construct a Japanese case report CTE dataset, J-Casemap, propose a generation-based CTE method that outperforms the baseline by 20.2 points in the human evaluation, and introduce evaluation metrics that reflect clinician preferences.Further experiments also show that J-Casemap enhances the performance of solving other medical tasks, such as question answering. Sakiko Yahata, Fei Cheng 0002, Sadao Kurohashi, Hisahiko Sato, Ryozo Nagai |
EMNLP | 4 |
| 2024 | Domain Transferable Semantic Frames for Expert Interview DialoguesabstractInterviews are an effective method to elicit critical skills to perform particular processes in various domains. In order to understand the knowledge structure of these domain-specific processes, we consider semantic role and predicate annotation based on Frame Semantics. We introduce a dataset of interview dialogues with experts in the culinary and gardening domains, each annotated with semantic frames. This dataset consists of (1) 308 interview dialogues related to the culinary domain, originally assembled by Okahisa et al. (2022), and (2) 100 interview dialogues associated with the gardening domain, which we newly acquired. The labeling specifications take into account the domain-transferability by adopting domain-agnostic labels for frame elements. In addition, we conducted domain transfer experiments from the culinary domain to the gardening domain to examine the domain transferability with our dataset. The experimental results showed the effectiveness of our domain-agnostic labeling scheme. Taishi Chika, Taro Okahisa, Takashi Kodama, Yin Jou Huang, Yugo Murawaki, Sadao Kurohashi |
LREC/COLING | 6 |
| 2024 | An Empirical Study of Synthetic Data Generation for Implicit Discourse Relation RecognitionabstractImplicit Discourse Relation Recognition (IDRR), which is the task of recognizing the semantic relation between given text spans that do not contain overt clues, is a long-standing and challenging problem. In particular, the paucity of training data for some error-prone discourse relations makes the problem even more challenging. To address this issue, we propose a method of generating synthetic data for IDRR using a large language model. The proposed method is summarized as two folds: extraction of confusing discourse relation pairs based on false negative rate and synthesis of data focused on the confusion. The key points of our proposed method are utilizing a confusion matrix and adopting two-stage prompting to obtain effective synthetic data. According to the proposed method, we generated synthetic data several times larger than training examples for some error-prone discourse relations and incorporated it into training. As a result of experiments, we achieved state-of-the-art macro-F1 performance thanks to the synthetic data without sacrificing micro-F1 performance and demonstrated its positive effects especially on recognizing some infrequent discourse relations. Kazumasa Omura, Fei Cheng 0002, Sadao Kurohashi |
LREC/COLING | 3 |
| 2024 | Identifying Source Language Expressions for Pre-editing in Machine TranslationabstractMachine translation-mediated communication can benefit from pre-editing source language texts to ensure accurate transmission of intended meaning in the target language. The primary challenge lies in identifying source language expressions that pose difficulties in translation. In this paper, we hypothesize that such expressions tend to be distinctive features of texts originally written in the source language (native language) rather than translations generated from the target language into the source language (machine translation). To identify such expressions, we train a neural classifier to distinguish native language from machine translation, and subsequently isolate the expressions that contribute to the model’s prediction of native language. Our manual evaluation revealed that our method successfully identified characteristic expressions of the native language, despite the noise and the inherent nuances of the task. We also present case studies where we edit the identified expressions to improve translation quality. Norizo Sakaguchi, Yugo Murawaki, Chenhui Chu, Sadao Kurohashi |
LREC/COLING | 4 |
| 2024 | Rapidly Developing High-quality Instruction Data and Evaluation Benchmark for Large Language Models with Minimal Human Effort: A Case Study on JapaneseabstractThe creation of instruction data and evaluation benchmarks for serving Large language models often involves enormous human annotation. This issue becomes particularly pronounced when rapidly developing such resources for a non-English language like Japanese. Instead of following the popular practice of directly translating existing English resources into Japanese (e.g., Japanese-Alpaca), we propose an efficient self-instruct method based on GPT-4. We first translate a small amount of English instructions into Japanese and post-edit them to obtain native-level quality. GPT-4 then utilizes them as demonstrations to automatically generate Japanese instruction data. We also construct an evaluation benchmark containing 80 questions across 8 categories, using GPT-4 to automatically assess the response quality of LLMs without human references. The empirical results suggest that the models fine-tuned on our GPT-4 self-instruct data significantly outperformed the Japanese-Alpaca across all three base pre-trained models. Our GPT-4 self-instruct data allowed the LLaMA 13B model to defeat GPT-3.5 (Davinci-003) with a 54.37% win-rate. The human evaluation exhibits the consistency between GPT-4’s assessments and human preference. Our high-quality instruction data and evaluation benchmark are released here. Yikun Sun, Nobuhiro Ueda, Sakiko Yahata, Fei Cheng 0002, Chenhui Chu, Sadao Kurohashi |
LREC/COLING | 7 |
| 2024 | Abstractive Multi-Video Captioning: Benchmark Dataset Construction and Extensive EvaluationabstractThis paper introduces a new task, abstractive multi-video captioning, which focuses on abstracting multiple videos with natural language. Unlike conventional video captioning tasks generating a specific caption for a video, our task generates an abstract caption of the shared content in a video group containing multiple videos. To address our task, models must learn to understand each video in detail and have strong abstraction abilities to find commonalities among videos. We construct a benchmark dataset for abstractive multi-video captioning named AbstrActs. AbstrActs contains 13.5k video groups and corresponding abstract captions. As abstractive multi-video captioning models, we explore two approaches: end-to-end and cascade. For evaluation, we proposed a new metric, CocoA, which can evaluate the model performance based on the abstractness of the generated captions. In experiments, we report the impact of the way of combining multiple video features, the overall model architecture, and the number of input videos. Rikito Takahashi, Hirokazu Kiyomaru, Chenhui Chu, Sadao Kurohashi |
LREC/COLING | 4 |
| 2024 | J-CRe3: A Japanese Conversation Dataset for Real-world Reference ResolutionabstractUnderstanding expressions that refer to the physical world is crucial for such human-assisting systems in the real world, as robots that must perform actions that are expected by users. In real-world reference resolution, a system must ground the verbal information that appears in user interactions to the visual information observed in egocentric views. To this end, we propose a multimodal reference resolution task and construct a Japanese Conversation dataset for Real-world Reference Resolution (J-CRe3). Our dataset contains egocentric video and dialogue audio of real-world conversations between two people acting as a master and an assistant robot at home. The dataset is annotated with crossmodal tags between phrases in the utterances and the object bounding boxes in the video frames. These tags include indirect reference relations, such as predicate-argument structures and bridging references as well as direct reference relations. We also constructed an experimental model and clarified the challenges in multimodal reference resolution tasks. Nobuhiro Ueda, Hideko Habe, Akishige Yuguchi, Seiya Kawano, Yasutomo Kawanishi, Sadao Kurohashi, Koichiro Yoshino |
LREC/COLING | 6 |
| 2024 | SubMerge: Merging Equivalent Subword Tokenizations for Subword Regularized Models in Neural Machine TranslationabstractSubword regularized models leverage multiple subword tokenizations of one target sentence during training. However, selecting one tokenization during inference leads to the underutilization of knowledge learned about multiple tokenizations.We propose the SubMerge algorithm to rescue the ignored Subword tokenizations through merging equivalent ones during inference.SubMerge is a nested search algorithm where the outer beam search treats the word as the minimal unit, and the inner beam search provides a list of word candidates and their probabilities, merging equivalent subword tokenizations. SubMerge estimates the probability of the next word more precisely, providing better guidance during inference.Experimental results on six low-resource to high-resource machine translation datasets show that SubMerge utilizes a greater proportion of a model’s probability weight during decoding (lower word perplexities for hypotheses). It also improves BLEU and chrF++ scores for many translation directions, most reliably for low-resource scenarios. We investigate the effect of different beam sizes, training set sizes, dropout rates, and whether it is effective on non-regularized models. Haiyue Song, Francois Meyer, Raj Dabre, Hideki Tanaka, Chenhui Chu, Sadao Kurohashi |
EAMT (1) | 6 |
| 2024 | A Comprehensive Analysis of Memorization in Large Language ModelsabstractThis paper presents a comprehensive study that investigates memorization in large language models (LLMs) from multiple perspectives.Experiments are conducted with the Pythia and LLM-jp model suites, both of which offer LLMs with over 10B parameters and full access to their pre-training corpora.Our findings include: (1) memorization is more likely to occur with larger model sizes, longer prompt lengths, and frequent texts, which aligns with findings in previous studies; (2) memorization is less likely to occur for texts not trained during the latter stages of training, even if they frequently appear in the training corpus; (3) the standard methodology for judging memorization can yield false positives, and texts that are infrequent yet flagged as memorized typically result from causes other than true memorization 1 . Hirokazu Kiyomaru, Issa Sugiura, Daisuke Kawahara, Sadao Kurohashi |
INLG | 4 |
| 2024 | EMS: Efficient and Effective Massively Multilingual Sentence Embedding LearningabstractMassively multilingual sentence representation models, e.g., LASER, SBERT-distill, and LaBSE, help significantly improve cross-lingual downstream tasks. However, the use of a large amount of data or inefficient model architectures results in heavy computation to train a new model according to our preferred languages and domains. To resolve this issue, we introduce efficient and effective massively multilingual sentence embedding (EMS), using cross-lingual token-level reconstruction (XTR) and sentence-level contrastive learning as training objectives. Compared with related studies, the proposed model can be efficiently trained using significantly fewer parallel sentences and GPU computation resources. Empirical results showed that the proposed model significantly yields better or comparable results with regard to cross-lingual sentence retrieval, zero-shot cross-lingual genre classification, and sentiment classification. Ablative analyses demonstrated the efficiency and effectiveness of each component of the proposed model. We release the codes for model training and the EMS pre-trained sentence embedding model, which supports 62 languages (https://github.com/Mao-KU/EMS). Zhuoyuan Mao, Chenhui Chu, Sadao Kurohashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | ComSearch: Equation Searching with Combinatorial Strategy for Solving Math Word Problems with Weak SupervisionabstractPrevious studies have introduced a weaklysupervised paradigm for solving math word problems requiring only the answer value annotation.While these methods search for correct value equation candidates as pseudo labels, they search among a narrow sub-space of the enormous equation space.To address this problem, we propose a novel search algorithm with combinatorial strategy ComSearch, which can compress the search space by excluding mathematically equivalent equations.The compression allows the searching algorithm to enumerate all possible equations and obtain high-quality data.We investigate the noise in the pseudo labels that hold wrong mathematical logic, which we refer to as the false-matching problem, and propose a ranking model to denoise the pseudo labels.Our approach holds a flexible framework to utilize two existing supervised math word problem solvers to train pseudo labels, and both achieve state-of-the-art performance in the weak supervision task. 1 Qianying Liu, Wenyu Guan, Jianhao Shen, Fei Cheng 0002, Sadao Kurohashi |
EACL | 5 |
| 2023 | SuperDialseg: A Large-scale Dataset for Supervised Dialogue SegmentationabstractDialogue segmentation is a crucial task for dialogue systems allowing a better understanding of conversational texts.Despite recent progress in unsupervised dialogue segmentation methods, their performances are limited by the lack of explicit supervised signals for training.Furthermore, the precise definition of segmentation points in conversations still remains as a challenging problem, increasing the difficulty of collecting manual annotations.In this paper, we provide a feasible definition of dialogue segmentation points with the help of document-grounded dialogues and release a large-scale supervised dataset called Su-perDialseg, containing 9,478 dialogues based on two prevalent document-grounded dialogue corpora, and also inherit their useful dialoguerelated annotations.Moreover, we provide a benchmark including 18 models across five categories for the dialogue segmentation task with several proper evaluation metrics.Empirical studies show that supervised learning is extremely effective in in-domain datasets and models trained on SuperDialseg can achieve good generalization ability on out-of-domain data.Additionally, we also conducted human verification on the test set and the Kappa score confirmed the quality of our automatically constructed dataset.We believe our work is an important step forward in the field of dialogue segmentation.Our codes and data can be found from: https://github.com/ Coldog2333/SuperDialseg.A2: Are you looking for family benefits?U1: Hello.I'd like to learn about your retirement program. Chengzhang Dong, Sadao Kurohashi, Akiko Aizawa |
EMNLP | 3 |
| 2023 | Video-Helpful Multimodal Machine TranslationabstractExisting multimodal machine translation (MMT) datasets consist of images and video captions or instructional video subtitles, which rarely contain linguistic ambiguity, making visual information ineffective in generating appropriate translations.Recent work has constructed an ambiguous subtitles dataset to alleviate this problem but is still limited to the problem that videos do not necessarily contribute to disambiguation.We introduce EVA (Extensive training set and Videohelpful evaluation set for Ambiguous subtitles translation), an MMT dataset containing 852k Japanese-English (Ja-En) parallel subtitle pairs, 520k Chinese-English (Zh-En) parallel subtitle pairs, and corresponding video clips collected from movies and TV episodes.In addition to the extensive training set, EVA contains a video-helpful evaluation set in which subtitles are ambiguous, and videos are guaranteed helpful for disambiguation.Furthermore, we propose SAFA, an MMT model based on the Selective Attention model with two novel methods: Frame attention loss and Ambiguity augmentation, aiming to use videos in EVA for disambiguation fully.Experiments on EVA show that visual information and the proposed methods can boost translation performance, and our model performs significantly better than existing MMT models.The EVA dataset and the SAFA model are available at: https://github.com/ku-nlp/video-helpful-MMT.git. Shuichiro Shimizu, Chenhui Chu, Sadao Kurohashi |
EMNLP | 4 |
| 2023 | GPT-RE: In-context Learning for Relation Extraction using Large Language ModelsabstractIn spite of the potential for ground-breaking achievements offered by large language models (LLMs) (e.g., GPT-3) via in-context learning (ICL), they still lag significantly behind fullysupervised baselines (e.g., fine-tuned BERT) in relation extraction (RE).This is due to the two major shortcomings of ICL for RE: (1) low relevance regarding entity and relation in existing sentence-level demonstration retrieval approaches for ICL; and (2) the lack of explaining input-label mappings of demonstrations leading to poor ICL effectiveness.In this paper, we propose GPT-RE to successfully address the aforementioned issues by (1) incorporating task-aware representations in demonstration retrieval; and (2) enriching the demonstrations with gold label-induced reasoning logic.We evaluate GPT-RE on four widely-used RE datasets and observe that GPT-RE achieves improvements over not only existing GPT-3 baselines, but also fully-supervised baselines as in Figure 1.Specifically, GPT-RE achieves SOTA performances on the Semeval and SciERC datasets, and competitive performances on the TACRED and ACE05 datasets.Additionally, a critical issue of LLMs revealed by previous work, the strong inclination to wrongly classify NULL examples into other predefined labels, is substantially alleviated by our method.We show an empirical analysis.1 Fei Cheng 0002, Zhuoyuan Mao, Qianying Liu, Haiyue Song, Jiwei Li 0001, Sadao Kurohashi |
EMNLP | 7 |
| 2023 | Hierarchical Softmax for End-To-End Low-Resource Multilingual Speech RecognitionabstractLow-resource speech recognition has been long-suffering from insufficient training data. In this paper, we propose an approach that leverages neighboring languages to improve low-resource scenario performance, founded on the hypothesis that similar linguistic units in neighboring languages exhibit comparable term frequency distributions, which enables us to construct a Huffman tree for performing multilingual hierarchical Softmax decoding. This hierarchical structure enables cross-lingual knowledge sharing among similar tokens, thereby enhancing low-resource training outcomes. Empirical analyses demonstrate that our method is effective in improving the accuracy and efficiency of low-resource speech recognition. Qianying Liu, Zhuo Gong, Zhengdong Yang, Sheng Li 0010, Chenchen Ding, Nobuaki Minematsu, Hao Huang 0009, Fei Cheng 0002, Chenhui Chu, Sadao Kurohashi |
ICASSP | 11 |
| 2023 | Toward Game-Based Learning of Japanese Writing for Elementary School StudentsabstractIt is a long-standing problem that many elementary school students in Japan have an aversion to writing compositions. To address the problem, we designed an AI educational game for elementary school students to study Japanese writing utilizing existing language resources. In the game, players construct simple and complex sentences by connecting given word cards with particle marks. The constructed sentences are automatically scored using large-scale language resources, allowing players to receive on-the-spot feedback, such as how to improve the use of a case marker. We also developed smartphone and web applications of the game and conducted a user study to assess it. The results of the user study demonstrated that our application can be used as a good introduction to studying Japanese writing. Kazumasa Omura, Kei Kubo, Frédéric Bergéron, Sadao Kurohashi |
ICCE | 4 |
| 2023 | SelfSeg: A Self-supervised Sub-word Segmentation Method for Neural Machine TranslationabstractSub-word segmentation is an essential pre-processing step for Neural Machine Translation (NMT). Existing work has shown that neural sub-word segmenters are better than Byte-Pair Encoding (BPE), however, they are inefficient, as they require parallel corpora, days to train, and hours to decode. This article introduces SelfSeg, a self-supervised neural sub-word segmentation method that is much faster to train/decode and requires only monolingual dictionaries instead of parallel corpora. SelfSeg takes as input a word in the form of a partially masked character sequence, optimizes the word generation probability, and generates the segmentation with the maximum posterior probability, which is calculated using a dynamic programming algorithm. The training time of SelfSeg depends on word frequencies, and we explore several word frequency normalization strategies to accelerate the training phase. Additionally, we propose a regularization mechanism that allows the segmenter to generate various segmentations for one word. To show the effectiveness of our approach, we conduct MT experiments in low-, middle-, and high-resource scenarios, where we compare the performance of using different segmentation methods. The experimental results demonstrate that, on the low-resource ALT dataset, our method achieves more than 1.2 BLEU score improvement compared with BPE and SentencePiece, and a 1.1 score improvement over Dynamic Programming Encoding (DPE) and Vocabulary Learning via Optimal Transport (VOLT), on average. The regularization method achieves approximately a 4.3 BLEU score improvement over BPE and a 1.2 BLEU score improvement over BPE-dropout, the regularized version of BPE. We also observed significant improvements on IWSLT15 Vi→En, WMT16 Ro→En, and WMT15 Fi→En datasets and competitive results on the WMT14 De→En and WMT14 Fr→En datasets. Furthermore, our method is 17.8× faster during training and up to 36.8× faster during decoding in a high-resource scenario compared to DPE. We provide extensive analysis, including why monolingual word-level data is enough to train SelfSeg. Haiyue Song, Raj Dabre, Chenhui Chu, Sadao Kurohashi, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2022 | Minimally-Supervised Joint Learning of Event Volitionality and Subject Animacy ClassificationabstractVolitionality and subject animacy are fundamental and closely related properties of an event. Their classification is challenging because it requires contextual text understanding and a huge amount of labeled data. This paper proposes a novel method that jointly learns volitionality and subject animacy at a low cost, heuristically labeling events in a raw corpus. Volitionality labels are assigned using a small lexicon of volitional and non-volitional adverbs such as deliberately and accidentally; subject animacy labels are assigned using a list of animate and inanimate nouns obtained from ontological knowledge. We then consider the problem of learning a classifier from the labeled events so that it can perform well on unlabeled events without the words used for labeling. We view the problem as a bias reduction or unsupervised domain adaptation problem and apply the techniques. We conduct experiments with crowdsourced gold data in Japanese and English and show that our method effectively learns volitionality and subject animacy without manually labeled data. Hirokazu Kiyomaru, Sadao Kurohashi |
AAAI | 2 |
| 2022 | Improving Commonsense Contingent Reasoning by Pseudo-data and Its Application to the Related TasksabstractContingent reasoning is one of the essential abilities in natural language understanding, and many language resources annotated with contingent relations have been constructed. However, despite the recent advances in deep learning, the task of contingent reasoning is still difficult for computers. In this study, we focus on the reasoning of contingent relation between basic events. Based on the existing data construction method, we automatically generate large-scale pseudo-problems and incorporate the generated data into training. We also investigate the generality of contingent knowledge through quantitative evaluation by performing transfer learning on the related tasks: discourse relation analysis, the Japanese Winograd Schema Challenge, and the JCommonsenseQA. The experimental results show the effectiveness of utilizing pseudo-problems for both the commonsense contingent reasoning task and the related tasks, which suggests the importance of contingent reasoning. Kazumasa Omura, Sadao Kurohashi |
COLING | 2 |
| 2022 | Rescue Implicit and Long-tail Cases: Nearest Neighbor Relation ExtractionabstractRelation extraction (RE) has achieved remarkable progress with the help of pre-trained language models.However, existing RE models are usually incapable of handling two situations: implicit expressions and long-tail relation types, caused by language complexity and data sparsity.In this paper, we introduce a simple enhancement of RE using k nearest neighbors (kNN-RE).kNN-RE allows the model to consult training relations at test time through a nearest-neighbor search and provides a simple yet effective means to tackle the two issues above.Additionally, we observe that kNN-RE serves as an effective way to leverage distant supervision (DS) data for RE.Experimental results show that the proposed kNN-RE achieves state-of-the-art performances on a variety of supervised RE datasets, i.e., ACE05, SciERC, and Wiki80, along with outperforming the best model to date on the i2b2 and Wiki80 datasets in the setting of allowing using DS.Our code and models are available at: https://github.com/YukinoWan/kNN-RE. Qianying Liu, Zhuoyuan Mao, Fei Cheng 0002, Sadao Kurohashi, Jiwei Li 0001 |
EMNLP | 5 |
| 2022 | JaMIE: A Pipeline Japanese Medical Information Extraction System with Novel Relation AnnotationabstractIn the field of Japanese medical information extraction, few analyzing tools are available and relation extraction is still an under-explored topic. In this paper, we first propose a novel relation annotation schema for investigating the medical and temporal relations between medical entities in Japanese medical reports. We experiment with the practical annotation scenarios by separately annotating two different types of reports. We design a pipeline system with three components for recognizing medical entities, classifying entity modalities, and extracting relations. The empirical results show accurate analyzing performance and suggest the satisfactory annotation quality, the superiority of the latest contextual embedding models. and the feasible annotation strategy for high-accuracy demand. Fei Cheng 0002, Shuntaro Yada, Ribeka Tanaka, Eiji Aramaki, Sadao Kurohashi |
LREC | 5 |
| 2022 | VISA: An Ambiguous Subtitles Dataset for Visual Scene-aware Machine TranslationabstractExisting multimodal machine translation (MMT) datasets consist of images and video captions or general subtitles which rarely contain linguistic ambiguity, making visual information not so effective to generate appropriate translations. We introduce VISA, a new dataset that consists of 40k Japanese-English parallel sentence pairs and corresponding video clips with the following key features: (1) the parallel sentences are subtitles from movies and TV episodes; (2) the source subtitles are ambiguous, which means they have multiple possible translations with different meanings; (3) we divide the dataset into Polysemy and Omission according to the cause of ambiguity. We show that VISA is challenging for the latest MMT system, and we hope that the dataset can facilitate MMT research. Shuichiro Shimizu, Weiqi Gu, Chenhui Chu, Sadao Kurohashi |
LREC | 5 |
| 2022 | Constructing a Culinary Interview Dialogue Corpus with Video Conferencing ToolabstractInterview is an efficient way to elicit knowledge from experts of different domains. In this paper, we introduce CIDC, an interview dialogue corpus in the culinary domain in which interviewers play an active role to elicit culinary knowledge from the cooking expert. The corpus consists of 308 interview dialogues (each about 13 minutes in length), which add up to a total of 69,000 utterances. We use a video conferencing tool for data collection, which allows us to obtain the facial expressions of the interlocutors as well as the screen-sharing contents. To understand the impact of the interlocutors’ skill level, we divide the experts into “semi-professionals’” and “enthusiasts” and the interviewers into “skilled interviewers” and “unskilled interviewers.” For quantitative analysis, we report the statistics and the results of the post-interview questionnaire. We also conduct qualitative analysis on the collected interview dialogues and summarize the salient patterns of how interviewers elicit knowledge from the experts. The corpus serves the purpose to facilitate future research on the knowledge elicitation mechanism in interview dialogues. Taro Okahisa, Ribeka Tanaka, Takashi Kodama, Yin Jou Huang, Sadao Kurohashi |
LREC | 5 |
| 2022 | Improving Event Duration Question Answering by Leveraging Existing Temporal Information Extraction DataabstractUnderstanding event duration is essential for understanding natural language. However, the amount of training data for tasks like duration question answering, i.e., McTACO, is very limited, suggesting a need for external duration information to improve this task. The duration information can be obtained from existing temporal information extraction tasks, such as UDS-T and TimeBank, where more duration data is available. A straightforward two-stage fine-tuning approach might be less likely to succeed given the discrepancy between the target duration question answering task and the intermediary duration classification task. This paper resolves this discrepancy by automatically recasting an existing event duration classification task from UDS-T to a question answering task similar to the target McTACO. We investigate the transferability of duration information by comparing whether the original UDS-T duration classification or the recast UDS-T duration question answering can be transferred to the target task. Our proposed model achieves a 13% Exact Match score improvement over the baseline on the McTACO duration question answering task, showing that the two-stage fine-tuning approach succeeds when the discrepancy between the target and intermediary tasks are resolved. Felix Giovanni Virgo, Fei Cheng 0002, Sadao Kurohashi |
LREC | 3 |
| 2022 | All-in-One: Emotion, Sentiment and Intensity Prediction Using a Multi-Task Ensemble FrameworkabstractWe propose a multi-task ensemble framework that jointly learns multiple related problems. The ensemble model aims to leverage the learned representations of three deep learning models (i.e., CNN, LSTM and GRU) and a hand-crafted feature representation for the predictions. Through multi-task framework, we address four problems of emotion and sentiment analysis, i.e., “emotionclassification&intensity”, “valence,arousal&dominancefor emotion”, “valence&arousalfor sentiment”, and “3-class categorical&5-class ordinal classificationfor sentiment”. The underlying problems cover two granularity (i.e.,coarse-grainedandfine-grained) and a diverse range of domains (i.e.,tweets,Facebook posts,news headlines,blogs,lettersetc.). Experimental results suggest that the proposed multi-task framework outperforms the single-task frameworks in all experiments. Md. Shad Akhtar, Deepanway Ghosal, Asif Ekbal, Pushpak Bhattacharyya, Sadao Kurohashi |
IEEE Trans. Affect. Comput. | 5 |
| 2022 | Linguistically Driven Multi-Task Pre-Training for Low-Resource Neural Machine TranslationabstractIn the present study, we propose novel sequence-to-sequence pre-training objectives for low-resource machine translation (NMT): Japanese-specific sequence to sequence (JASS) for language pairs involving Japanese as the source or target language, and English-specific sequence to sequence (ENSS) for language pairs involving English. JASS focuses on masking and reordering Japanese linguistic units known as bunsetsu, whereas ENSS is proposed based on phrase structure masking and reordering tasks. Experiments on ASPEC Japanese–English & Japanese–Chinese, Wikipedia Japanese–Chinese, News English–Korean corpora demonstrate that JASS and ENSS outperform MASS and other existing language-agnostic pre-training methods by up to +2.9 BLEU points for the Japanese–English tasks, up to +7.0 BLEU points for the Japanese–Chinese tasks and up to +1.3 BLEU points for English–Korean tasks. Empirical analysis, which focuses on the relationship between individual parts in JASS and ENSS, reveals the complementary nature of the subtasks of JASS and ENSS. Adequacy evaluation using LASER, human evaluation, and case studies reveals that our proposed methods significantly outperform pre-training methods without injected linguistic knowledge and they have a larger positive impact on the adequacy as compared to the fluency. Zhuoyuan Mao, Chenhui Chu, Sadao Kurohashi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2022 | RODA: Reverse Operation Based Data Augmentation for Solving Math Word ProblemsabstractAutomatically solving math word problems is a critical task in the field of natural language processing. Recent models have reached their performance bottleneck and require more high-quality data for training. We propose a novel data augmentation method that reverses the mathematical logic of math word problems to produce new high-quality math problems and introduce new knowledge points that can benefit learning the mathematical reasoning logic. We apply the augmented data on two SOTA math word problem solving models and compare our results with a strong data augmentation baseline. Experimental results show the effectiveness of our approach (we release our code and data athttps://github.com/yiyunya/RODA). Qianying Liu, Wenyu Guan, Sujian Li, Fei Cheng 0002, Daisuke Kawahara, Sadao Kurohashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2021 | Lightweight Cross-Lingual Sentence Representation LearningabstractZhuoyuan Mao, Prakhar Gupta, Chenhui Chu, Martin Jaggi, Sadao Kurohashi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zhuoyuan Mao, Prakhar Gupta, Chenhui Chu, Martin Jaggi, Sadao Kurohashi |
ACL/IJCNLP (1) | 5 |
| 2021 | Extractive Summarization Considering Discourse and Coreference Relations based on Heterogeneous GraphabstractModeling the relations between text spans in a document is a crucial yet challenging problem for extractive summarization.Various kinds of relations exist among text spans of different granularity, such as discourse relations between elementary discourse units and coreference relations between phrase mentions.In this paper, we propose a heterogeneous graph based model for extractive summarization that incorporates both discourse and coreference relations.The heterogeneous graph contains three types of nodes, each corresponds to text spans of different granularity.Experimental results on a benchmark summarization dataset verify the effectiveness of our proposed method. Yin Jou Huang, Sadao Kurohashi |
EACL | 2 |
| 2021 | Contextualized and Generalized Sentence Representations by Contrastive Self-Supervised Learning: A Case Study on Discourse Relation AnalysisabstractWe propose a method to learn contextualized and generalized sentence representations using contrastive self-supervised learning.In the proposed method, a model is given a text consisting of multiple sentences.One sentence is randomly selected as a target sentence.The model is trained to maximize the similarity between the representation of the target sentence with its context and that of the masked target sentence with the same context.Simultaneously, the model minimize the similarity between the latter representation and the representation of a random sentence with the same context.We apply our method to discourse relation analysis in English and Japanese and show that it outperforms strong baseline methods based on BERT, XLNet, and RoBERTa. Hirokazu Kiyomaru, Sadao Kurohashi |
NAACL-HLT | 2 |
| 2021 | Frustratingly Easy Edit-based Linguistic Steganography with a Masked Language ModelabstractWith advances in neural language models, the focus of linguistic steganography has shifted from edit-based approaches to generationbased ones.While the latter's payload capacity is impressive, generating genuine-looking texts remains challenging.In this paper, we revisit edit-based linguistic steganography, with the idea that a masked language model offers an off-the-shelf solution.The proposed method eliminates painstaking rule construction and has a high payload capacity for an edit-based model.It is also shown to be more secure against automatic detection than a generation-based method while offering better control of the security/payload capacity tradeoff. Honai Ueoka, Yugo Murawaki, Sadao Kurohashi |
NAACL-HLT | 3 |
| 2021 | Flexibly Focusing on Supporting Facts, Using Bridge Links, and Jointly Training Specialized Modules for Multi-Hop Question AnsweringabstractWith the help of the detailed annotated question answering dataset HotpotQA, recent question answering models are trained to justify their predicted answers with supporting facts from context documents. Some related works train the same model to find supporting facts and answers jointly without having specialized models for each task. The others train separate models for each task, but do not use supporting facts effectively to find the answer; they either use only the predicted sentences and ignore the remaining context, or do not use them at all. Furthermore, while complex graph-based models consider the bridge/connection between documents in the multi-hop setting, simple BERT-based models usually drop it. We propose FlexibleFocusedReader (FFReader), a model that 1) Flexibly focuses on predicted supporting facts (SFs) without ignoring the important remaining context, 2) Focuses on the bridge between documents, despite not using graph architectures, and 3) Jointly learns predicting SFs and answering with two specialized models. Our model achieves consistent improvement over the baseline. In particular, we find that flexibly focusing on SFs is important, rather than ignoring remaining context or not using SFs at all for finding the answer. We also find that tagging the entity that links the documents at hand is very beneficial. Finally, we show that joint training is crucial for FFReader. Tareq Alkhaldi 0001, Chenhui Chu, Sadao Kurohashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Automatically Neutralizing Subjective Bias in TextabstractTexts like news, encyclopedias, and some social media strive for objectivity. Yet bias in the form of inappropriate subjectivity — introducing attitudes via framing, presupposing truth, and casting doubt — remains ubiquitous. This kind of bias erodes our collective trust and fuels social conflict. To address this issue, we introduce a novel testbed for natural language generation: automatically bringing inappropriately subjective text into a neutral point of view (“neutralizing” biased text). We also offer the first parallel corpus of biased language. The corpus contains 180,000 sentence pairs and originates from Wikipedia edits that removed various framings, presuppositions, and attitudes from biased sentences. Last, we propose two strong encoder-decoder baselines for the task. A straightforward yet opaque concurrent system uses a BERT encoder to identify subjective words as part of the generation process. An interpretable and controllable modular algorithm separates these steps, using (1) a BERT-based classifier to identify problematic words and (2) a novel join embedding through which the classifier can edit the hidden states of the encoder. Large-scale human evaluation across four domains (encyclopedias, news headlines, books, and political speeches) suggests that these algorithms are a first step towards the automatic identification and reduction of bias. Reid Pryzant, Richard Diehl Martinez, Nathan Dass, Sadao Kurohashi, Daniel Jurafsky, Diyi Yang |
AAAI | 4 |
| 2020 | Native-like Expression Identification by Contrasting Native and Proficient Second Language SpeakersabstractWe propose a novel task of native-like expression identification by contrasting texts written by native speakers and those by proficient second language speakers.This task is highly challenging mainly because 1) the combinatorial nature of expressions prevents us from choosing candidate expressions a priori and 2) the distributions of the two types of texts overlap considerably.Our solution to the first problem is to combine a powerful neural network-based classifier of sentencelevel nativeness with an explainability method that measures an approximate contribution of a given expression to the classifier's prediction.To address the second problem, we introduce a special label neutral and reformulate the classification task as complementary-label learning.Our crowdsourcing-based evaluation and in-depth analysis suggest that our method successfully uncovers linguistically interesting usages distinctive of native speech. Oleksandr Harust, Yugo Murawaki, Sadao Kurohashi |
COLING | 3 |
| 2020 | BERT-based Cohesion Analysis of Japanese TextsabstractThe meaning of natural language text is supported by cohesion among various kinds of entities, including coreference relations, predicate-argument structures, and bridging anaphora relations. However, predicate-argument structures for nominal predicates and bridging anaphora relations have not been studied well, and their analyses have been still very difficult. Recent advances in neural networks, in particular, self training-based language models including BERT (Devlin et al., 2019), have significantly improved many natural language processing tasks, making it possible to dive into the study on analysis of cohesion in the whole text. In this study, we tackle an integrated analysis of cohesion in Japanese texts. Our results significantly outperformed existing studies in each task, especially about 10 to 20 point improvement both for zero anaphora and coreference resolution. Furthermore, we also showed that coreference resolution is different in nature from the other tasks and should be treated specially. Nobuhiro Ueda, Daisuke Kawahara, Sadao Kurohashi |
COLING | 3 |
| 2020 | A Method for Building a Commonsense Inference Dataset based on Basic EventsabstractWe present a scalable, low-bias, and low-cost method for building a commonsense inference dataset that combines automatic extraction from a corpus and crowdsourcing.Each problem is a multiple-choice question that asks contingency between basic events.We applied the proposed method to a Japanese corpus and acquired 104k problems.While humans can solve the resulting problems with high accuracy (88.9%), the accuracy of a highperformance transfer learning model is reasonably low (76.0%).We also confirmed through dataset analysis that the resulting dataset contains low bias.We released the datatset to facilitate language understanding research.1 Kazumasa Omura, Daisuke Kawahara, Sadao Kurohashi |
EMNLP (1) | 3 |
| 2020 | Acquiring Social Knowledge about Personality and Driving-related BehaviorabstractIn this paper, we introduce our psychological approach to collect human-specific social knowledge from a text corpus, using NLP techniques. It is often not explicitly described but shared among people, which we call social knowledge. We focus on the social knowledge, especially personality and driving. We used the language resources that were developed based on psychological research methods; a Japanese personality dictionary (317 words) and a driving experience corpus (8,080 sentences) annotated with behavior and subjectivity. Using them, we automatically extracted collocations between personality descriptors and driving-related behavior from a driving behavior and subjectivity corpus (1,803,328 sentences after filtering) and obtained unique 5,334 collocations. To evaluate the collocations as social knowledge, we designed four step-by-step crowdsourcing tasks. They resulted in 266 pieces of social knowledge. They include the knowledge that might be difficult to recall by themselves but easy to agree with. We discuss the acquired social knowledge and the contribution to implementations into systems. Ritsuko Iwai, Daisuke Kawahara, Takatsune Kumada, Sadao Kurohashi |
LREC | 4 |
| 2020 | Development of a Japanese Personality Dictionary based on Psychological MethodsabstractWe propose a new approach to constructing a personality dictionary with psychological evidence. In this study, we collect personality words, using word embeddings, and construct a personality dictionary with weights for Big Five traits. The weights are calculated based on the responses of the large sample (N=1,938, female = 1,004, M=49.8years old:20-78, SD=16.3). All the respondents answered a 20-item personality questionnaire and 537 personality items derived from word embeddings. We present the procedures to examine the qualities of responses with psychological methods and to calculate the weights. These result in a personality dictionary with two sub-dictionaries. We also discuss an application of the acquired resources. Ritsuko Iwai, Daisuke Kawahara, Takatsune Kumada, Sadao Kurohashi |
LREC | 4 |
| 2020 | Adapting BERT to Implicit Discourse Relation Classification with a Focus on Discourse ConnectivesabstractBERT, a neural network-based language model pre-trained on large corpora, is a breakthrough in natural language processing, significantly outperforming previous state-of-the-art models in numerous tasks. However, there have been few reports on its application to implicit discourse relation classification, and it is not clear how BERT is best adapted to the task. In this paper, we test three methods of adaptation. (1) We perform additional pre-training on text tailored to discourse classification. (2) In expectation of knowledge transfer from explicit discourse relations to implicit discourse relations, we add a task named explicit connective prediction at the additional pre-training step. (3) To exploit implicit connectives given by treebank annotators, we add a task named implicit connective prediction at the fine-tuning step. We demonstrate that these three techniques can be combined straightforwardly in a single training pipeline. Through comprehensive experiments, we found that the first and second techniques provide additional gain while the last one did not. Yudai Kishimoto, Yugo Murawaki, Sadao Kurohashi |
LREC | 3 |
| 2020 | JASS: Japanese-specific Sequence to Sequence Pre-training for Neural Machine TranslationabstractNeural machine translation (NMT) needs large parallel corpora for state-of-the-art translation quality. Low-resource NMT is typically addressed by transfer learning which leverages large monolingual or parallel corpora for pre-training. Monolingual pre-training approaches such as MASS (MAsked Sequence to Sequence) are extremely effective in boosting NMT quality for languages with small parallel corpora. However, they do not account for linguistic information obtained using syntactic analyzers which is known to be invaluable for several Natural Language Processing (NLP) tasks. To this end, we propose JASS, Japanese-specific Sequence to Sequence, as a novel pre-training alternative to MASS for NMT involving Japanese as the source or target language. JASS is joint BMASS (Bunsetsu MASS) and BRSS (Bunsetsu Reordering Sequence to Sequence) pre-training which focuses on Japanese linguistic units called bunsetsus. In our experiments on ASPEC Japanese–English and News Commentary Japanese–Russian translation we show that JASS can give results that are competitive with if not better than those given by MASS. Furthermore, we show for the first time that joint MASS and JASS pre-training gives results that significantly surpass the individual methods indicating their complementary nature. We will release our code, pre-trained models and bunsetsu annotated data as resources for researchers to use in their own NLP tasks. Zhuoyuan Mao, Fabien Cromières, Raj Dabre, Haiyue Song, Sadao Kurohashi |
LREC | 5 |
| 2020 | Coursera Corpus Mining and Multistage Fine-Tuning for Improving Lectures TranslationabstractLectures translation is a case of spoken language translation and there is a lack of publicly available parallel corpora for this purpose. To address this, we examine a framework for parallel corpus mining which is a quick and effective way to mine a parallel corpus from publicly available lectures at Coursera. Our approach determines sentence alignments, relying on machine translation and cosine similarity over continuous-space sentence representations. We also show how to use the resulting corpora in a multistage fine-tuning based domain adaptation for high-quality lectures translation. For Japanese–English lectures translation, we extracted parallel data of approximately 40,000 lines and created development and test sets through manual filtering for benchmarking translation performance. We demonstrate that the mined corpus greatly enhances the quality of translation when used in conjunction with out-of-domain parallel corpora via multistage training. This paper also suggests some guidelines to gather and clean corpora, mine parallel sentences, address noise in the mined data, and create high-quality evaluation splits. For the sake of reproducibility, we have released our code for parallel data creation. Haiyue Song, Raj Dabre, Atsushi Fujita, Sadao Kurohashi |
LREC | 4 |
| 2020 | Towards a Versatile Medical-Annotation Guideline Feasible Without Heavy Medical Knowledge: Starting From Critical Lung DiseasesabstractApplying natural language processing (NLP) to medical and clinical texts can bring important social benefits by mining valuable information from unstructured text. A popular application for that purpose is named entity recognition (NER), but the annotation policies of existing clinical corpora have not been standardized across clinical texts of different types. This paper presents an annotation guideline aimed at covering medical documents of various types such as radiography interpretation reports and medical records. Furthermore, the annotation was designed to avoid burdensome requirements related to medical knowledge, thereby enabling corpus development without medical specialists. To achieve these design features, we specifically focus on critical lung diseases to stabilize linguistic patterns in corpora. After annotating around 1100 electronic medical records following the annotation scheme, we demonstrated its feasibility using an NER task. Results suggest that our guideline is applicable to large-scale clinical NLP projects. Shuntaro Yada, Ayami Joh, Ribeka Tanaka, Fei Cheng 0002, Eiji Aramaki, Sadao Kurohashi |
LREC | 6 |
| 2019 | Minimally Supervised Learning of Affective Events Using Discourse RelationsabstractJun Saito, Yugo Murawaki, Sadao Kurohashi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jun Saito, Yugo Murawaki, Sadao Kurohashi |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Emotion helps Sentiment: A Multi-task Model for Sentiment and Emotion AnalysisabstractIn this paper, we propose a two-layered multi-task attention based neural network that performs sentiment analysis through emotion analysis. The proposed approach is based on Bidirectional Long Short-Term Memory and uses Distributional Thesaurus as a source of external knowledge to improve the sentiment and emotion prediction. The proposed system has two levels of attention to hierarchically build a meaningful representation. We evaluate our system on the benchmark dataset of SemEval 2016 Task 6 and also compare it with the state-of-the-art systems on Stance Sentiment Emotion Corpus. Experimental results show that the proposed system improves the performance of sentiment analysis by 3.2 F-score points on SemEval 2016 Task 6 dataset. Our network also boosts the performance of emotion analysis by 5 F-score points on Stance Sentiment Emotion Corpus. Asif Ekbal, Daisuke Kawahara, Sadao Kurohashi |
IJCNN | 4 |
| 2019 | Applying Machine Translation to Psychology: Automatic Translation of Personality Adjectives
Ritsuko Iwai, Daisuke Kawahara, Takatsune Kumada, Sadao Kurohashi |
MTSummit (2) | 4 |
| 2019 | FAQ Retrieval using Query-Question Similarity and BERT-Based Query-Answer RelevanceabstractFrequently Asked Question (FAQ) retrieval is an important task where the objective is to retrieve an appropriate Question-Answer (QA) pair from a database based on a user's query. We propose a FAQ retrieval system that considers the similarity between a user's query and a question as well as the relevance between the query and an answer. Although a common approach to FAQ retrieval is to construct labeled data for training, it takes annotation costs. Therefore, we use a traditional unsupervised information retrieval system to calculate the similarity between the query and question. On the other hand, the relevance between the query and answer can be learned by using QA pairs in a FAQ database. The recently-proposed BERT model is used for the relevance calculation. Since the number of QA pairs in FAQ page is not enough to train a model, we cope with this issue by leveraging FAQ sets that are similar to the one in question. We evaluate our approach on two datasets. The first one is localgovFAQ, a dataset we construct in a Japanese administrative municipality domain. The second is StackExchange dataset, which is the public dataset in English. We demonstrate that our proposed method outperforms baseline methods on these datasets. Wataru Sakata, Tomohide Shibata, Ribeka Tanaka, Sadao Kurohashi |
SIGIR | 4 |
| 2018 | Neural Adversarial Training for Semi-supervised Japanese Predicate-argument Structure AnalysisabstractJapanese predicate-argument structure (PAS) analysis involves zero anaphora resolution, which is notoriously difficult.To improve the performance of Japanese PAS analysis, it is straightforward to increase the size of corpora annotated with PAS.However, since it is prohibitively expensive, it is promising to take advantage of a large amount of raw corpora.In this paper, we propose a novel Japanese PAS analysis model based on semi-supervised adversarial training with a raw corpus.In our experiments, our model outperforms existing state-of-the-art models for Japanese PAS analysis. Shuhei Kurita, Daisuke Kawahara, Sadao Kurohashi |
ACL (1) | 3 |
| 2018 | Entity-Centric Joint Modeling of Japanese Coreference Resolution and Predicate Argument Structure AnalysisabstractPredicate argument structure analysis is a task of identifying structured events.To improve this field, we need to identify a salient entity, which cannot be identified without performing coreference resolution and predicate argument structure analysis simultaneously.This paper presents an entity-centric joint model for Japanese coreference resolution and predicate argument structure analysis.Each entity is assigned an embedding, and when the result of both analyses refers to an entity, the entity embedding is updated.The analyses take the entity embedding into consideration to access the global information of entities.Our experimental results demonstrate the proposed method can improve the performance of the intersentential zero anaphora resolution drastically, which is a notoriously difficult task in predicate argument structure analysis. Tomohide Shibata, Sadao Kurohashi |
ACL (1) | 2 |
| 2018 | A Knowledge-Augmented Neural Network Model for Implicit Discourse Relation ClassificationabstractIdentifying discourse relations that are not overtly marked with discourse connectives remains a challenging problem. The absence of explicit clues indicates a need for the combination of world knowledge and weak contextual clues, which can hardly be learned from a small amount of manually annotated data. In this paper, we address this problem by augmenting the input text with external knowledge and context and by adopting a neural network model that can effectively handle the augmented text. Experiments show that external knowledge did improve the classification accuracy. Contextual information provided no significant gain for implicit discourse relations, but it did for explicit ones. Yudai Kishimoto, Yugo Murawaki, Sadao Kurohashi |
COLING | 3 |
| 2018 | Cross-lingual Knowledge Projection Using Machine Translation and Target-side Knowledge Base CompletionabstractConsiderable effort has been devoted to building commonsense knowledge bases. However, they are not available in many languages because the construction of KBs is expensive. To bridge the gap between languages, this paper addresses the problem of projecting the knowledge in English, a resource-rich language, into other languages, where the main challenge lies in projection ambiguity. This ambiguity is partially solved by machine translation and target-side knowledge base completion, but neither of them is adequately reliable by itself. We show their combination can project English commonsense knowledge into Japanese and Chinese with high precision. Our method also achieves a top-10 accuracy of 90% on the crowdsourced English–Japanese benchmark. Furthermore, we use our method to obtain 18,747 facts of accurate Japanese commonsense within a very short period. Naoki Otani, Hirokazu Kiyomaru, Daisuke Kawahara, Sadao Kurohashi |
COLING | 4 |
| 2018 | Improving Crowdsourcing-Based Annotation of Japanese Discourse Relations
Yudai Kishimoto, Shinnosuke Sawada, Yugo Murawaki, Daisuke Kawahara, Sadao Kurohashi |
LREC | 5 |
| 2018 | Comprehensive Annotation of Various Types of Temporal Information on the Time Axis
Tomohiro Sakaguchi, Daisuke Kawahara, Sadao Kurohashi |
LREC | 3 |
| 2018 | Annotating a Driving Experience Corpus with Behavior and Subjectivity
Ritsuko Iwai, Daisuke Kawahara, Takatsune Kumada, Sadao Kurohashi |
PACLIC | 4 |
| 2017 | Neural Joint Model for Transition-based Chinese Syntactic AnalysisabstractWe present neural network-based joint models for Chinese word segmentation, POS tagging and dependency parsing.Our models are the first neural approaches for fully joint Chinese analysis that is known to prevent the error propagation problem of pipeline models.Although word embeddings play a key role in dependency parsing, they cannot be applied directly to the joint task in the previous work.To address this problem, we propose embeddings of character strings, in addition to words.Experiments show that our models outperform existing systems in Chinese word segmentation and POS tagging, and perform preferable accuracies in dependency parsing.We also explore bi-LSTM models with fewer features. Shuhei Kurita, Daisuke Kawahara, Sadao Kurohashi |
ACL (1) | 3 |
| 2017 | Timeline Generation Based on a Two-Stage Event-Time Anchoring Model
Tomohiro Sakaguchi, Sadao Kurohashi |
CICLing (2) | 2 |
| 2017 | Improving Chinese Semantic Role Labeling using High-quality Surface and Deep Case FramesabstractThis paper presents a method for improving semantic role labeling (SRL) using a large amount of automatically acquired knowledge.We acquire two varieties of knowledge, which we call surface case frames and deep case frames.Although the surface case frames are compiled from syntactic parses and can be used as rich syntactic knowledge, they have limited capability for resolving semantic ambiguity.To compensate the deficiency of the surface case frames, we compile deep case frames from automatic semantic roles.We also consider quality management for both types of knowledge in order to get rid of the noise brought from the automatic analyses.The experimental results show that Chinese SRL can be improved using automatically acquired knowledge and the quality management shows a positive effect on this task. Gongye Jin, Daisuke Kawahara, Sadao Kurohashi |
EACL (1) | 3 |
| 2017 | Enabling Multi-Source Neural Machine Translation By Concatenating Source Sentences In Multiple Languages
Raj Dabre, Fabien Cromières, Sadao Kurohashi |
MTSummit (1) | 3 |
| 2016 | Neural Network-Based Model for Japanese Predicate Argument Structure Analysis
Tomohide Shibata, Daisuke Kawahara, Sadao Kurohashi |
ACL (1) | 3 |
| 2016 | Age Related Differences in Episodic Memory Recollections: Applying Latent Dirichlet Allocation to Free-Writings on Driving Incidents by Older and Young Drivers
Ritsuko Iwai, Takatsune Kumada, Daisuke Kawahara, Sadao Kurohashi |
CogSci | 4 |
| 2016 | Consistent Word Segmentation, Part-of-Speech Tagging and Dependency Labelling Annotation for Chinese LanguageabstractIn this paper, we propose a new annotation approach to Chinese word segmentation, part-of-speech (POS) tagging and dependency labelling that aims to overcome the two major issues in traditional morphology-based annotation: Inconsistency and data sparsity. We re-annotate the Penn Chinese Treebank 5.0 (CTB5) and demonstrate the advantages of this approach compared to the original CTB5 annotation through word segmentation, POS tagging and machine translation experiments. Mo Shen, Wingmui Li, HyunJeong Choe, Chenhui Chu, Daisuke Kawahara, Sadao Kurohashi |
COLING | 6 |
| 2016 | Insertion Position Selection Model for Flexible Non-Terminals in Dependency Tree-to-Tree Machine Translation
Toshiaki Nakazawa, Sadao Kurohashi |
EMNLP | 3 |
| 2016 | IRT-based Aggregation Model of Crowdsourced Pairwise Comparison for Evaluating Machine TranslationsabstractRecent work on machine translation has used crowdsourcing to reduce costs of manual evaluations.However, crowdsourced judgments are often biased and inaccurate.In this paper, we present a statistical model that aggregates many manual pairwise comparisons to robustly measure a machine translation system's performance.Our method applies graded response model from item response theory (IRT), which was originally developed for academic tests.We conducted experiments on a public dataset from the Workshop on Statistical Machine Translation 2013, and found that our approach resulted in highly interpretable estimates and was less affected by noisy judges than previously proposed methods. Naoki Otani, Toshiaki Nakazawa, Daisuke Kawahara, Sadao Kurohashi |
EMNLP | 4 |
| 2016 | Simultaneous Sentence Boundary Detection and Alignment with Pivot-based Machine Translation Generated Lexicons
Antoine Bourlon, Chenhui Chu, Toshiaki Nakazawa, Sadao Kurohashi |
LREC | 4 |
| 2016 | Parallel Sentence Extraction from Comparable Corpora with Neural Network Features
Chenhui Chu, Raj Dabre, Sadao Kurohashi |
LREC | 3 |
| 2016 | Paraphrasing Out-of-Vocabulary Words with Word Embeddings and Semantic Lexicons for Low Resource Statistical Machine Translation
Chenhui Chu, Sadao Kurohashi |
LREC | 2 |
| 2016 | ASPEC: Asian Scientific Paper Excerpt Corpus
Toshiaki Nakazawa, Manabu Yaguchi, Kiyotaka Uchimoto, Masao Utiyama, Eiichiro Sumita, Sadao Kurohashi, Hitoshi Isahara |
LREC | 6 |
| 2016 | Flexible Non-Terminals for Dependency Tree-to-Tree ReorderingabstractJohn Richardson, Fabien Cromières, Toshiaki Nakazawa, Sadao Kurohashi. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Fabien Cromières, Toshiaki Nakazawa, Sadao Kurohashi |
HLT-NAACL | 4 |
| 2016 | Integrated Parallel Sentence and Fragment Extraction from Comparable Corpora: A Case Study on Chinese-Japanese WikipediaabstractParallel corpora are crucial for statistical machine translation (SMT); however, they are quite scarce for most language pairs and domains. As comparable corpora are far more available, many studies have been conducted to extract either parallel sentences or fragments from them for SMT. In this article, we propose an integrated system to extract both parallel sentences and fragments from comparable corpora. We first apply parallel sentence extraction to identify parallel sentences from comparable sentences. We then extract parallel fragments from the comparable sentences. Parallel sentence extraction is based on a parallel sentence candidate filter and classifier for parallel sentence identification. We improve it by proposing a novel filtering strategy and three novel feature sets for classification. Previous studies have found it difficult to accurately extract parallel fragments from comparable sentences. We propose an accurate parallel fragment extraction method that uses an alignment model to locate the parallel fragment candidates and an accurate lexicon-based filter to identify the truly parallel fragments. A case study on the Chinese--Japanese Wikipedia indicates that our proposed methods outperform previously proposed methods, and the parallel data extracted by our system significantly improves SMT performance. Chenhui Chu, Toshiaki Nakazawa, Sadao Kurohashi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2015 | Morphological Analysis for Unsegmented Languages using Recurrent Neural Network Language ModelabstractWe present a new morphological analy-sis model that considers semantic plausi-bility of word sequences by using a re-current neural network language model (RNNLM). In unsegmented languages, since language models are learned from automatically segmented texts and in-evitably contain errors, it is not apparent that conventional language models con-tribute to morphological analysis. To solve this problem, we do not use language mod-els based on raw word sequences but use a semantically generalized language model, RNNLM, in morphological analysis. In our experiments on two Japanese corpora, our proposed model significantly outper-formed baseline models. This result indi-cates the effectiveness of RNNLM in mor-phological analysis. 1 Hajime Morita, Daisuke Kawahara, Sadao Kurohashi |
EMNLP | 3 |
| 2015 | Korean-Chinese word translation using Chinese character knowledge
Yuanmei Lu, Toshiaki Nakazawa, Sadao Kurohashi |
MTSummit | 3 |
| 2015 | Leveraging Small Multilingual Corpora for SMT Using Many Pivot LanguagesabstractRaj Dabre, Fabien Cromieres, Sadao Kurohashi, Pushpak Bhattacharyya. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Raj Dabre, Fabien Cromières, Sadao Kurohashi, Pushpak Bhattacharyya |
HLT-NAACL | 3 |
| 2015 | Large-scale Dictionary Construction via Pivot-based Statistical Machine Translation with Significance Pruning and Neural Network Features
Raj Dabre, Chenhui Chu, Fabien Cromières, Toshiaki Nakazawa, Sadao Kurohashi |
PACLIC | 5 |
| 2015 | Pivot-Based Topic Models for Low-Resource Lexicon Extraction
Toshiaki Nakazawa, Sadao Kurohashi |
PACLIC | 3 |
| 2015 | Cross-language Projection of Dependency Trees for Tree-to-tree Machine Translation
Chenhui Chu, Fabien Cromières, Sadao Kurohashi |
PACLIC | 4 |
| 2015 | Preordering using a Target-Language Parser via Cross-Language Syntactic Projection for Statistical Machine TranslationabstractWhen translating between languages with widely different word orders, word reordering can present a major challenge. Although some word reordering methods do not employ source-language syntactic structures, such structures are inherently useful for word reordering. However, high-quality syntactic parsers are not available for many languages. We propose a preordering method using a target-language syntactic parser to process source-language syntactic structures without a source-language syntactic parser. To train our preordering model based on ITG, we produced syntactic constituent structures for source-language training sentences by (1) parsing target-language training sentences, (2) projecting constituent structures of the target-language sentences to the corresponding source-language sentences, (3) selecting parallel sentences with highly synchronized parallel structures, (4) producing probabilistic models for parsing using the projected partial structures and the Pitman-Yor process, and (5) parsing to produce full binary syntactic structures maximally synchronized with the corresponding target-language syntactic structures, using the constraints of the projected partial structures and the probabilistic models. Our ITG-based preordering model is trained using the produced binary syntactic structures and word alignments. The proposed method facilitates the learning of ITG by producing highly synchronized parallel syntactic structures based on cross-language syntactic projection and sentence selection. The preordering model jointly parses input sentences and identifies their reordered structures. Experiments with Japanese--English and Chinese--English patent translation indicate that our method outperforms existing methods, including string-to-tree syntax-based SMT, a preordering method that does not require a parser, and a preordering method that uses a source-language dependency parser. Isao Goto, Masao Utiyama, Eiichiro Sumita, Sadao Kurohashi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2014 | Iterative Bilingual Lexicon Extraction from Comparable Corpora with Topical and Contextual Knowledge
Chenhui Chu, Toshiaki Nakazawa, Sadao Kurohashi |
CICLing (2) | 3 |
| 2014 | Rapid Development of a Corpus with Discourse Annotations using Two-stage Crowdsourcing
Daisuke Kawahara, Yuichiro Machida, Tomohide Shibata, Sadao Kurohashi, Hayato Kobayashi, Manabu Sassano |
COLING | 4 |
| 2014 | Translation Rules with Right-Hand Side LatticesabstractIn Corpus-Based Machine Translation, the search space of the translation candidates for a given input sentence is often defined by a set of (cyclefree) context-free grammar rules. This happens naturally in Syntax-Based Machine Translation and Hierarchical Phrase-Based Machine Translation (where the representation will be the set of the target-side half of the synchronous rules used to parse the input sentence). But it is also possible to describe Phrase-Based Machine Translation in this framework. We propose a natural extension to this representation by using lattice-rules that allow to easily encode an exponential number of variations of each rules. We also demonstrate how the representation of the search space has an impact on decoding efficiency, and how it is possible to optimize this representation. Fabien Cromières, Sadao Kurohashi |
EMNLP | 2 |
| 2014 | Constructing a Chinese―Japanese Parallel Corpus from Wikipedia
Chenhui Chu, Toshiaki Nakazawa, Sadao Kurohashi |
LREC | 3 |
| 2014 | Constructing a Corpus of Japanese Predicate Phrases for Synonym/Antonym Relations
Tomoko Izumi, Tomohide Shibata, Hisako Asano, Yoshihiro Matsuo, Sadao Kurohashi |
LREC | 5 |
| 2014 | A Framework for Compiling High Quality Knowledge Resources From Raw Corpora
Gongye Jin, Daisuke Kawahara, Sadao Kurohashi |
LREC | 3 |
| 2014 | Bilingual Dictionary Construction with Transliteration Filtering
Toshiaki Nakazawa, Sadao Kurohashi |
LREC | 3 |
| 2014 | A Large Scale Database of Strongly-related Events in Japanese
Tomohide Shibata, Shotaro Kohama, Sadao Kurohashi |
LREC | 3 |
| 2014 | Improving Statistical Machine Translation Accuracy Using Bilingual Lexicon Extractionwith Paraphrases
Chenhui Chu, Toshiaki Nakazawa, Sadao Kurohashi |
PACLIC | 3 |
| 2014 | Distortion Model Based on Word Sequence Labeling for Statistical Machine TranslationabstractThis article proposes a new distortion model for phrase-based statistical machine translation. In decoding, a distortion model estimates the source word position to be translated next (subsequent position; SP) given the last translated source word position (current position; CP). We propose a distortion model that can simultaneously consider the word at the CP, the word at an SP candidate, the context of the CP and an SP candidate, relative word order among the SP candidates, and the words between the CP and an SP candidate. These considered elements are called rich context . Our model considers rich context by discriminating label sequences that specify spans from the CP to each SP candidate. It enables our model to learn the effect of relative word order among SP candidates as well as to learn the effect of distances from the training data. In contrast to the learning strategy of existing methods, our learning strategy is that the model learns preference relations among SP candidates in each sentence of the training data. This leaning strategy enables consideration of all of the rich context simultaneously. In our experiments, our model had higher BLUE and RIBES scores for Japanese-English, Chinese-English, and German-English translation compared to the lexical reordering models. Isao Goto, Masao Utiyama, Eiichiro Sumita, Akihiro Tamura, Sadao Kurohashi |
ACM Trans. Asian Lang. Inf. Process. | 5 |
| 2014 | Dependency Parse Reranking with Rich Subtree FeaturesabstractIn pursuing machine understanding of human language, highly accurate syntactic analysis is a crucial step. In this work, we focus on dependency grammar, which models syntax by encoding transparent predicate-argument structures. Recent advances in dependency parsing have shown that employing higher-order subtree structures in graph-based parsers can substantially improve the parsing accuracy. However, the inefficiency of this approach increases with the order of the subtrees. This work explores a new reranking approach for dependency parsing that can utilize complex subtree representations by applying efficient subtree selection methods. We demonstrate the effectiveness of the approach in experiments conducted on the Penn Treebank and the Chinese Treebank. Our system achieves the best performance among known supervised systems evaluated on these datasets, improving the baseline accuracy from 91.88% to 93.42% for English, and from 87.39% to 89.25% for Chinese. Mo Shen, Daisuke Kawahara, Sadao Kurohashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Distortion Model Considering Rich Context for Statistical Machine Translation
Isao Goto, Masao Utiyama, Eiichiro Sumita, Akihiro Tamura, Sadao Kurohashi |
ACL (1) | 5 |
| 2013 | Japanese Zero Reference Resolution Considering Exophora and Author/Reader MentionsabstractIn Japanese, zero references often occur and many of them are categorized into zero exophora, in which a referent is not mentioned in the document.However, previous studies have focused on only zero endophora, in which a referent explicitly appears.We present a zero reference resolution model considering zero exophora and author/reader of a document.To deal with zero exophora, our model adds pseudo entities corresponding to zero exophora to candidate referents of zero pronouns.In addition, we automatically detect mentions that refer to the author and reader of a document by using lexico-syntactic patterns.We represent their particular behavior in a discourse as a feature vector of a machine learning model.The experimental results demonstrate the effectiveness of our model for not only zero exophora but also zero endophora. Masatsugu Hangyo, Daisuke Kawahara, Sadao Kurohashi |
EMNLP | 3 |
| 2013 | Automatic Knowledge Acquisition for Case Alternation between the Passive and Active Voices in JapaneseabstractWe present a method for automatically acquiring knowledge for case alternation between the passive and active voices in Japanese.By leveraging several linguistic constraints on alternation patterns and lexical case frames obtained from a large Web corpus, our method aligns a case frame in the passive voice to a corresponding case frame in the active voice and finds an alignment between their cases.We then apply the acquired knowledge to a case alternation task and prove its usefulness. Ryohei Sasano, Daisuke Kawahara, Sadao Kurohashi, Manabu Okumura |
EMNLP | 3 |
| 2013 | Accurate Parallel Fragment Extraction from Quasi-Comparable Corpora using Alignment Model and Translation Lexicon
Chenhui Chu, Toshiaki Nakazawa, Sadao Kurohashi |
IJCNLP | 3 |
| 2013 | High Quality Dependency Selection from Automatic Parses
Gongye Jin, Daisuke Kawahara, Sadao Kurohashi |
IJCNLP | 3 |
| 2013 | Precise Information Retrieval Exploiting Predicate-Argument Structures
Daisuke Kawahara, Keiji Shinzato, Tomohide Shibata, Sadao Kurohashi |
IJCNLP | 4 |
| 2013 | Robust Transliteration Mining from Comparable Corpora with Bilingual Topic Models
Toshiaki Nakazawa, Sadao Kurohashi |
IJCNLP | 3 |
| 2013 | A Simple Approach to Unknown Word Processing in Japanese Morphological Analysis
Ryohei Sasano, Sadao Kurohashi, Manabu Okumura |
IJCNLP | 2 |
| 2013 | Chinese Word Segmentation by Mining Maximized Substrings
Mo Shen, Daisuke Kawahara, Sadao Kurohashi |
IJCNLP | 3 |
| 2013 | Chinese-Japanese Machine Translation Exploiting Chinese CharactersabstractThe Chinese and Japanese languages share Chinese characters. Since the Chinese characters in Japanese originated from ancient China, many common Chinese characters exist between these two languages. Since Chinese characters contain significant semantic information and common Chinese characters share the same meaning in the two languages, they can be quite useful in Chinese-Japanese machine translation (MT). We therefore propose a method for creating a Chinese character mapping table for Japanese, traditional Chinese, and simplified Chinese, with the aim of constructing a complete resource of common Chinese characters. Furthermore, we point out two main problems in Chinese word segmentation for Chinese-Japanese MT, namely, unknown words and word segmentation granularity, and propose an approach exploiting common Chinese characters to solve these problems. We also propose a statistical method for detecting other semantically equivalent Chinese characters other than the common ones and a method for exploiting shared Chinese characters in phrase alignment. Results of the experiments carried out on a state-of-the-art phrase-based statistical MT system and an example-based MT system show that our proposed approaches can improve MT performance significantly, thereby verifying the effectiveness of shared Chinese characters for Chinese-Japanese MT. Chenhui Chu, Toshiaki Nakazawa, Daisuke Kawahara, Sadao Kurohashi |
ACM Trans. Asian Lang. Inf. Process. | 4 |
| 2012 | Flexible Japanese Sentence Compression by Relaxing Unit Constraints
Jun Harashima, Sadao Kurohashi |
COLING | 2 |
| 2012 | Semi-Supervised Noun Compound Analysis with Edge and Span Features
Yugo Murawaki, Sadao Kurohashi |
COLING | 2 |
| 2012 | Alignment by Bilingual Generation and Monolingual Derivation
Toshiaki Nakazawa, Sadao Kurohashi |
COLING | 2 |
| 2012 | Exploiting Shared Chinese Characters in Chinese Word Segmentation Optimization for Chinese-Japanese Machine Translation
Chenhui Chu, Toshiaki Nakazawa, Daisuke Kawahara, Sadao Kurohashi |
EAMT | 4 |
| 2012 | Chinese Characters Mapping Table of Japanese, Traditional Chinese and Simplified Chinese
Chenhui Chu, Toshiaki Nakazawa, Sadao Kurohashi |
LREC | 3 |
| 2012 | Building a Diverse Document Leads Corpus Annotated with Semantic Relations
Masatsugu Hangyo, Daisuke Kawahara, Sadao Kurohashi |
PACLIC | 3 |
| 2012 | A Reranking Approach for Dependency Parsing with Variable-sized Subtree Features
Mo Shen, Daisuke Kawahara, Sadao Kurohashi |
PACLIC | 3 |
| 2012 | Predicate-Argument Structure-Based Textual Entailment Recognition System Exploiting Wide-Coverage Lexical KnowledgeabstractThis article proposes a predicate-argument structure based Textual Entailment Recognition system exploiting wide-coverage lexical knowledge. Different from conventional machine learning approaches where several features obtained from linguistic analysis and resources are utilized, our proposed method regards a predicate-argument structure as a basic unit, and performs the matching/alignment between a text and hypothesis. In matching between predicate-arguments, wide-coverage relations between words/phrases such as synonym and is-a are utilized, which are automatically acquired from a dictionary, Web corpus, and Wikipedia. Tomohide Shibata, Sadao Kurohashi |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2011 | Extracting Paraphrases from Definition Sentences on the Web
Chikara Hashimoto, Kentaro Torisawa, Stijn De Saeger, Jun'ichi Kazama, Sadao Kurohashi |
ACL | 5 |
| 2011 | Efficient retrieval of tree translation examples for Syntax-Based Machine Translation
Fabien Cromières, Sadao Kurohashi |
EMNLP | 2 |
| 2011 | Non-parametric Bayesian Segmentation of Japanese Noun Phrases
Yugo Murawaki, Sadao Kurohashi |
EMNLP | 2 |
| 2011 | Relevance Feedback using Latent Information
Jun Harashima, Sadao Kurohashi |
IJCNLP | 2 |
| 2011 | Generative Modeling of Coordination by Factoring Parallelism and Selectional Preferences
Daisuke Kawahara, Sadao Kurohashi |
IJCNLP | 2 |
| 2011 | Bayesian Subtree Alignment Model based on Dependency Trees
Toshiaki Nakazawa, Sadao Kurohashi |
IJCNLP | 2 |
| 2011 | A Discriminative Approach to Japanese Zero Anaphora Resolution with Large-scale Lexicalized Case Frames
Ryohei Sasano, Sadao Kurohashi |
IJCNLP | 2 |
| 2011 | Acquiring Strongly-related Events using Predicate-argument Co-occurring Statistics and Case Frames
Tomohide Shibata, Sadao Kurohashi |
IJCNLP | 2 |
| 2011 | Japanese-Chinese Phrase Alignment Using Common Chinese Characters Information
Chenhui Chu, Toshiaki Nakazawa, Sadao Kurohashi |
MTSummit | 3 |
| 2011 | Web Spam Detection by Exploring Densely Connected SubgraphsabstractIn this paper, we present a Web spam detection algorithm that relies on link analysis. The method consists of three steps: (1) decomposition of web graphs in densely connected sub graphs and calculation of the features for each sub graph, (2) use of SVM classifiers to identify sub graphs composed of Web spam, and (3) propagation of predictions over web graphs by a biased Page Rank algorithm to expand the scope of identification. We performed experiments on a public benchmark. An empirical study of the core structure of web graphs suggests that highly ranked non-spam hosts can be identified by viewing the coreness of the web graph elements. Yutaka I. Leon-Suematsu, Kentaro Inui, Sadao Kurohashi, Yutaka Kidawara |
Web Intelligence | 3 |
| 2010 | Using Smaller Constituents Rather Than Sentences in Active Learning for Japanese Dependency Parsing
Manabu Sassano, Sadao Kurohashi |
ACL | 2 |
| 2010 | Acquiring Reliable Predicate-argument Structures from Raw Corpora for Case Frame Compilation
Daisuke Kawahara, Sadao Kurohashi |
LREC | 2 |
| 2010 | Online Japanese Unknown Morpheme Detection using Orthographic Variation
Yugo Murawaki, Sadao Kurohashi |
LREC | 2 |
| 2010 | Dependency Tree-based Sentiment Classification using CRFs with Hidden Variables
Tetsuji Nakagawa, Kentaro Inui, Sadao Kurohashi |
HLT-NAACL | 3 |
| 2009 | An Alignment Algorithm Using Belief Propagation and a Structure-Based Distortion Model
Fabien Cromières, Sadao Kurohashi |
EACL | 2 |
| 2009 | A Probabilistic Model for Associative Anaphora Resolution
Ryohei Sasano, Sadao Kurohashi |
EMNLP | 2 |
| 2009 | The Effect of Corpus Size on Case Frame Acquisition for Discourse Analysis
Ryohei Sasano, Daisuke Kawahara, Sadao Kurohashi |
HLT-NAACL | 3 |
| 2009 | Identifying Information Sender Configuration of Web PagesabstractThe source of a piece of information is a crucial element to consider when judging the credibility of that information. In this paper, we address the task of identifying the information source which is cast as a problem of identifying the {\em information sender configuration (ISC)} of a Web page. An information sender of a Web page is an entity which is involved in the publication of the information on the page. An ISC of a Web page describes the information senders of the page and the relationship among them. Information sender extraction is thus a subtask of identifying ISC, and we present a method for extracting information senders from Web pages and offer preliminary evaluation. The ISC provides a basis for deeper analysis of information on the Web. Yoshikiyo Kato, Daisuke Kawahara, Kentaro Inui, Sadao Kurohashi, Tomohide Shibata |
Web Intelligence | 4 |
| 2009 | Web Information Organization Using Keyword Distillation Based ClusteringabstractThis paper describes a system that conducts search result clustering for several thousands of Web pages, and elaborates cluster labels through keyword distillation. Keyword distillation is a method that properly handles spelling variations, transliterations, synonyms, inclusion relations and word ambiguity, using linguistic resources and contexts of a user's query. The system provides a clustering result from 1,000 pages in less than one minute by taking advantage of a search engine infrastructure and grid computing environment. Experimental results show that the system correctly merged synonymous keywords and is useful for finding topics hidden in the lower-ranked pages in a search result. Tomohide Shibata, Yasuo Bamba, Keiji Shinzato, Sadao Kurohashi |
Web Intelligence | 4 |
| 2008 | Coordination Disambiguation without Any Similarities
Daisuke Kawahara, Sadao Kurohashi |
COLING | 2 |
| 2008 | A Fully-Lexicalized Probabilistic Model for Japanese Zero Anaphora Resolution
Ryohei Sasano, Daisuke Kawahara, Sadao Kurohashi |
COLING | 3 |
| 2008 | Chinese Dependency Parsing with Large Scale Automatically Constructed Case Structures
Daisuke Kawahara, Sadao Kurohashi |
COLING | 3 |
| 2008 | Online Acquisition of Japanese Unknown Morphemes using Morphological Constraints
Yugo Murawaki, Sadao Kurohashi |
EMNLP | 2 |
| 2008 | Japanese Named Entity Recognition Using Structural Natural Language Processing
Ryohei Sasano, Sadao Kurohashi |
IJCNLP | 2 |
| 2008 | SYNGRAPH: A Flexible Matching Method based on Synonymous Expression Extraction from an Ordinary Dictionary and a Web Corpus
Tomohide Shibata, Michitaka Odani, Jun Harashima, Takashi Oonishi, Sadao Kurohashi |
IJCNLP | 5 |
| 2008 | TSUBAKI: An Open Search Engine Infrastructure for Developing New Information Access Methodology
Keiji Shinzato, Tomohide Shibata, Daisuke Kawahara, Chikara Hashimoto, Sadao Kurohashi |
IJCNLP | 5 |
| 2008 | A Large-Scale Web Data Collection as a Natural Language Processing Infrastructure
Keiji Shinzato, Daisuke Kawahara, Chikara Hashimoto, Sadao Kurohashi |
LREC | 4 |
| 2008 | Grasping Major Statements and Their Contradictions Toward Information Credibility Analysis of Web ContentsabstractThe World Wide Web contains wide variety of news reports, arguments, opinions, etc. that vary widely in quality. People judge the credibility of information on the Web for decision making in daily life. At present, while the quantity of information on the Web is explosively increasing, it is necessary to develop a system that supports such judgments. We have been developing an information credibility analysis system, WISDOM that considers the viewpoints of information contents, information senders, and information appearances. In this paper, as a viewpoint of information contents, we propose a method for providing a bird's eye view of major statements on a given topic and their contradictions. We evaluate the obtained statements in our experiments, and confirm the effectiveness of our approach. Furthermore, we discuss our future objectives. Daisuke Kawahara, Sadao Kurohashi, Kentaro Inui |
Web Intelligence | 2 |
| 2007 | Construction of Domain Dictionary for Fundamental Vocabulary
Chikara Hashimoto, Sadao Kurohashi |
ACL | 2 |
| 2007 | Probabilistic Coordination Disambiguation in a Fully-Lexicalized Japanese Parser
Daisuke Kawahara, Sadao Kurohashi |
EMNLP-CoNLL | 2 |
| 2007 | Automatic object model acquisition and object recognition by integrating linguistic and visual informationabstractIn order to make the best use of multimedia contents effectively, the crucial point is the structural analysis of the contents, in which several media processing techniques, including image, audio and text analyses, should be integrated. To understand utterances in videos in accordance with the scene, it is essential to recognize what object appears in the videos. In this paper, we focus on Japanese cooking TV videos, and propose a method for acquiring object models of foods in an unsupervised manner and performing object recognition based on the acquired object models. First, a topic of each video segment is identified based on HMMs to obtain good examples for the object model acquisition. After that, close-up images are extracted from image sequences, and an attention region on the close-up image is determined. Then, an important word is extracted as a keyword from utterances around the close-up image, and is made correspond to the close-up image. By collecting a set of close-up image and keyword from a large amount of videos, object models are acquired. After acquiring the object models, object recognition is performed based on the acquired object models and linguistic information. We conducted experiments on two kinds of cooking TV programs. We acquired the object models of around 100 foods with an accuracy 77.8%. The F measure of object recognition was 0.727. Tomohide Shibata, Norio Kato, Sadao Kurohashi |
ACM Multimedia | 3 |
| 2007 | Development of a Japanese-Chinese machine translation system
Hitoshi Isahara, Sadao Kurohashi, Jun'ichi Tsujii, Kiyotaka Uchimoto, Hiroshi Nakagawa, Hiroyuki Kaji, Shun'ichi Kikuchi |
MTSummit | 2 |
| 2007 | Structural phrase alignment based on consistency criteria
Toshiaki Nakazawa, Sadao Kurohashi |
MTSummit | 3 |
| 2006 | Unsupervised Topic Identification by Integrating Linguistic and Visual Information Based on Hidden Markov Models
Tomohide Shibata, Sadao Kurohashi |
ACL | 2 |
| 2006 | Case Frame Compilation from the Web using High-Performance Computing
Daisuke Kawahara, Sadao Kurohashi |
LREC | 2 |
| 2006 | A Fully-Lexicalized Probabilistic Model for Japanese Syntactic and Case Structure Analysis
Daisuke Kawahara, Sadao Kurohashi |
HLT-NAACL | 2 |
| 2006 | Cards-to-presentation on the web: generating multimedia contents featuring agent animations
Yukiko I. Nakano, Toshihiro Murayama, Masashi Okamoto, Daisuke Kawahara, Sadao Kurohashi, Toyoaki Nishida |
J. Netw. Comput. Appl. | 6 |
| 2005 | Lexical Choice via Topic Adaptation for Paraphrasing Written Language to Spoken Language
Nobuhiro Kaji, Sadao Kurohashi |
IJCNLP | 2 |
| 2005 | PP-Attachment Disambiguation Boosted by a Gigantic Volume of Unambiguous Examples
Daisuke Kawahara, Sadao Kurohashi |
IJCNLP | 2 |
| 2005 | Automatic Acquisition of Basic Katakana Lexicon from a Given Corpus
Toshiaki Nakazawa, Daisuke Kawahara, Sadao Kurohashi |
IJCNLP | 3 |
| 2005 | Automatic Slide Generation Based on Discourse Structure Analysis
Tomohide Shibata, Sadao Kurohashi |
IJCNLP | 2 |
| 2005 | Probabilistic Model for Example-based Machine TranslationabstractExample-based machine translation (EBMT) systems, so far, rely on heuristic measures in retrieving translation examples. Such a heuristic measure costs time to adjust, and might make its algorithm unclear. This paper presents a probabilistic model for EBMT. Under the proposed model, the system searches the translation example combination which has the highest probability. The proposed model clearly formalizes EBMT process. In addition, the model can naturally incorporate the context similarity of translation examples. The experimental results demonstrate that the proposed model has a slightly better translation quality than state-of-the-art EBMT systems. Eiji Aramaki, Sadao Kurohashi, Hideki Kashioka, Naoto Katoh |
MTSummit | 2 |
| 2004 | Improving Japanese Zero Pronoun Resolution by Global Word Sense Disambiguation
Daisuke Kawahara, Sadao Kurohashi |
COLING | 2 |
| 2004 | Automatic Construction of Nominal Case Frames and its Application to Indirect Anaphora Resolution
Ryohei Sasano, Daisuke Kawahara, Sadao Kurohashi |
COLING | 3 |
| 2004 | Example-Based Machine Translation Without Saying Inferable Predicate
Eiji Aramaki, Sadao Kurohashi, Hideki Kashioka, Hideki Tanaka |
IJCNLP | 2 |
| 2004 | Zero Pronoun Resolution Based on Automatically Constructed Case Frames and Structural Preference of Antecedents
Daisuke Kawahara, Sadao Kurohashi |
IJCNLP | 2 |
| 2004 | Resolution of Modifier-Head Relation Gaps Using Automatically Extracted Metonymic Expressions
Yoji Kiyota, Sadao Kurohashi, Fuyuko Kido |
IJCNLP | 2 |
| 2004 | Video Content Manipulation by Means of Content Annotation and Nonsymbolic Gestural Interfaces
Burin Anuchitkittikul, Masashi Okamoto, Sadao Kurohashi, Toyoaki Nishida, Yoichi Sato 0001 |
KES | 3 |
| 2004 | Structural Analysis of Instruction Utterances Using Linguistic and Visual Information
Tomohide Shibata, Masato Tachiki, Daisuke Kawahara, Masashi Okamoto, Sadao Kurohashi, Toyoaki Nishida |
KES | 5 |
| 2004 | Toward Text Understanding: Integrating Relevance-tagged Corpus and Automatically Constructed Case Frames
Daisuke Kawahara, Ryohei Sasano, Sadao Kurohashi |
LREC | 3 |
| 2004 | Paraphrasing Predicates from Written Language to Spoken Language Using the Web
Nobuhiro Kaji, Masashi Okamoto, Sadao Kurohashi |
HLT-NAACL | 3 |
| 2003 | Embodied Conversational Agents for Presenting Intellectual Multimedia Contents
Yukiko I. Nakano, Toshihiro Murayama, Daisuke Kawahara, Sadao Kurohashi, Toyoaki Nishida |
KES | 4 |
| 2003 | Structural Analysis of Instruction Utterances
Tomohide Shibata, Daisuke Kawahara, Masashi Okamoto, Sadao Kurohashi, Toyoaki Nishida |
KES | 4 |
| 2002 | Verb Paraphrase based on Case Frame AlignmentabstractThis paper describes a method of translating a predicate-argument structure of a verb into that of an equivalent verb, which is a core component of the dictionary-based paraphrasing. Our method grasps several usages of a headword and those of the def-heads as a form of their case frames and aligns those case frames, which means the acquisition of word sense disambiguation rules and the detection of the appropriate equivalent and case marker transformation. Nobuhiro Kaji, Daisuke Kawahara, Sadao Kurohashi, Satoshi Sato |
ACL | 3 |
| 2002 | Fertilization of Case Frame Dictionary for Robust Japanese Case Analysis
Daisuke Kawahara, Sadao Kurohashi |
COLING | 2 |
| 2002 | "Dialog Navigator": A Question Answering System Based on Large Text Knowledge Base
Yoji Kiyota, Sadao Kurohashi, Fuyuko Kido |
COLING | 2 |
| 2002 | Construction of a Japanese Relevance-tagged Corpus
Daisuke Kawahara, Sadao Kurohashi, Kôiti Hasida |
LREC | 2 |
| 2001 | Finding translation correspondences from parallel parsed corpus for example-based translationabstractThis paper describes a system for finding phrasal translation correspondences from parallel parsed corpus that are collections paired English and Japanese sentences. First, the system finds phrasal correspondences by Japanese-English translation dictionary consultation. Then, the system finds correspondences in remaining phrases by using sentences dependency structures and the balance of all correspondences. The method is based on an assumption that in parallel corpus most fragments in a source sentence have corresponding fragments in a target sentence. Eiji Aramaki, Sadao Kurohashi, Satoshi Sato, Hideo Watanabe |
MTSummit | 2 |
| 2001 | Building domain-independent text generation system
Xinyu Deng, Sadao Kurohashi, Jun-ichi Nakamura |
PACLIC | 2 |
| 2000 | Japanese Case Structure Analysis
Daisuke Kawahara, Nobuhiro Kaji, Sadao Kurohashi |
COLING | 3 |
| 2000 | Finding Structural Correspondences from Bilingual Parsed Corpus for Corpus-based Translation
Hideo Watanabe, Sadao Kurohashi, Eiji Aramaki |
COLING | 2 |
| 2000 | Nonlocal Language Modeling based on Context Co-occurrence VectorsabstractThis paper presents a novel nonlocal language model which utilizes contextual information. A reduced vector space model calculated from co-occurrences of word pairs provides word co-occurrence vectors. The sum of word co-occurrence vectors represents the context of a document, and the cosine similarity between the context vector and the word co-occurrence vectors represents the long-distance lexical dependencies. Experiments on the Mainichi Newspaper corpus show significant improvement in perplexity (5.0% overall and 27.2% on target vocabulary) Sadao Kurohashi, Manabu Ori |
EMNLP | 1 |
| 2000 | Discourse Structure Analysis for News Video by Checking Surface Information in the TranscriptabstractVarious kinds of video recordings have discourse structures. Therefore, it is important to determine how video segments are combined and what kind of coherence relations they are connected with. We propose a method for estimating the discourse structure of video news reports. Yasuhiko Watanabe, Yoshihiro Okada, Eiichi Iwanari, Sadao Kurohashi |
ICPR | 4 |
| 2000 | A Parallel English-Japanese Query Collection for the Evaluation of On-Line Help Systems
Richard F. E. Sutcliffe, Sadao Kurohashi |
LREC | 2 |
| 1999 | Semantic Analysis of Japanese Noun Phrases - A New Approach to Dictionary-Based UnderstandingabstractThis paper presents a new method of analyzing Japanese noun phrases of the form N1 no N2. The Japanese postposition no roughly corresponds to of, but it has much broader usage. The method exploits a definition of N2 in a dictionary. For example, rugby no coach can be interpreted as a person who teaches technique in rugby. We illustrate the effectiveness of the method by the analysis of 300 test noun phrases. Sadao Kurohashi, Yasuyuki Sakai |
ACL | 1 |
| 1999 | Automatic Discovery of Definition Patterns Based on the MDL Principle
Masatoshi Tsuchiya, Sadao Kurohashi |
Discovery Science | 2 |
| 1994 | Automatic Detection of Discourse Structure by Checking Surface Information in Sentences
Sadao Kurohashi, Makoto Nagao |
COLING | 1 |
| 1994 | A Syntactic Analysis Method of Long Japanese Sentences Based on the Detection of Conjunctive Structures
Sadao Kurohashi, Makoto Nagao |
Comput. Linguistics | 1 |
| 1992 | Dynamic Programming Method for Analyzing Conjunctive Structures in Japanese
Sadao Kurohashi, Makoto Nagao |
COLING | 1 |