EDBT 2026 Demo / reviewers in the wild / expert
Daisuke Kawahara
dblp:72/4981
· DBLP profile ↗
83ranked-venue papers
23as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 82 · 23 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary ReconstructionabstractRecent advances in multimodal large language models (MLLMs) have significantly enhanced video understanding capabilities, opening new possibilities for practical applications. Yet current video benchmarks focus largely on indoor scenes or short-range outdoor activities, leaving the challenges associated with long-distance travel largely unexplored. Mastering extended geospatial-temporal trajectories is critical for next-generation MLLMs, underpinning real-world tasks such as embodied-AI planning and navigation. To bridge this gap, we present VIR-Bench, a novel benchmark consisting of 200 travel videos that frames itinerary reconstruction as a challenging task designed to evaluate and push forward MLLMs' geospatial-temporal intelligence. Experimental results reveal that state-of-the-art MLLMs, including proprietary ones, struggle to achieve high scores, underscoring the difficulty of handling videos that span extended spatial and temporal scales. Moreover, we conduct an in-depth case study in which we develop a prototype travel-planning agent that leverages the insights gained from VIR-Bench. The agent’s markedly improved itinerary recommendations verify that our evaluation protocol not only benchmarks models effectively but also translates into concrete performance gains in user-facing applications. Eiki Murata, Lingfang Zhang, Ayako Sato, So Fukuda, Keisuke Nakao, Yusuke Nakamura, Sebastian Zwirner, Yi-Chia Chen, Hiroyuki Otomo, Hiroki Ouchi, Daisuke Kawahara |
AAAI | 14 |
| 2026 | Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images
Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara |
ICDAR (2) | 3 |
| 2026 | Building Effective Japanese Medical LLMs with an Open Recipe for Domain Adaptation through Continued Pre-training
Akiko Aizawa, Yuki Arase, Fei Cheng 0002, Teruhito Kanazawa, Daisuke Kawahara, Kazuma Kobayashi, Takashi Kodama, Sadao Kurohashi, Yusuke Oda, Tsuta Yuma, Zhishen Yang, Rio Yokota |
LREC | 8 |
| 2026 | JMTEB and JMTEB-lite: Japanese Massive Text Embedding Benchmark and Its Lightweight Version
Shengzhe Li, Masaya Ohagi, Ryokan Ri, Akihiko Fukuchi, Tomohide Shibata, Daisuke Kawahara |
LREC | 6 |
| 2026 | Construction of a Japanese RAG Benchmark Using Synthetic Documents on Non-existent Entities and Events
Shengzhe Li, Masaya Ohagi, Hayato Tsukagoshi, Akihiko Fukuchi, Tomohide Shibata, Daisuke Kawahara |
LREC | 6 |
| 2026 | Constructing a Japanese Claim Decomposition Dataset for Fact-Checking of LLM-Generated Texts
Miwa Masano, Ribeka Keyaki, Atsushi Keyaki, Rei Minamoto, Kaito Horio, Hirokazu Kiyomaru, Kouta Nakayama, Hideyuki Tachibana, Daisuke Kawahara |
LREC | 9 |
| 2026 | Evaluating Multimodal Large Language Models on Vertically Written Japanese Text
Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara |
LREC | 3 |
| 2024 | Time-aware COMET: A Commonsense Knowledge Model with Temporal KnowledgeabstractTo better handle commonsense knowledge, which is difficult to acquire in ordinary training of language models, commonsense knowledge graphs and commonsense knowledge models have been constructed. The former manually and symbolically represents commonsense, and the latter stores these graphs’ knowledge in the models’ parameters. However, the existing commonsense knowledge models that deal with events do not consider granularity or time axes. In this paper, we propose a time-aware commonsense knowledge model, TaCOMET. The construction of TaCOMET consists of two steps. First, we create TimeATOMIC using ChatGPT, which is a commonsense knowledge graph with time. Second, TaCOMET is built by continually finetuning an existing commonsense knowledge model on TimeATOMIC. TimeATOMIC and continual finetuning let the model make more time-aware generations with rich commonsense than the existing commonsense models. We also verify the applicability of TaCOMET on a robotic decision-making task. TaCOMET outperformed the existing commonsense knowledge model when proper times are input. Our dataset and models will be made publicly available. Eiki Murata, Daisuke Kawahara |
LREC/COLING | 2 |
| 2024 | A Comprehensive Analysis of Memorization in Large Language ModelsabstractThis paper presents a comprehensive study that investigates memorization in large language models (LLMs) from multiple perspectives.Experiments are conducted with the Pythia and LLM-jp model suites, both of which offer LLMs with over 10B parameters and full access to their pre-training corpora.Our findings include: (1) memorization is more likely to occur with larger model sizes, longer prompt lengths, and frequent texts, which aligns with findings in previous studies; (2) memorization is less likely to occur for texts not trained during the latter stages of training, even if they frequently appear in the training corpus; (3) the standard methodology for judging memorization can yield false positives, and texts that are infrequent yet flagged as memorized typically result from causes other than true memorization 1 . Hirokazu Kiyomaru, Issa Sugiura, Daisuke Kawahara, Sadao Kurohashi |
INLG | 3 |
| 2024 | Advantages of Persistent Cohomology in Estimating Animal Location From Grid Cell Population ActivityabstractMany cognitive functions are represented as cell assemblies. In the case of spatial navigation, the population activity of place cells in the hippocampus and grid cells in the entorhinal cortex represents self-location in the environment. The brain cannot directly observe self-location information in the environment. Instead, it relies on sensory information and memory to estimate self-location. Therefore, estimating low-dimensional dynamics, such as the movement trajectory of an animal exploring its environment, from only the high-dimensional neural activity is important in deciphering the information represented in the brain. Most previous studies have estimated the low-dimensional dynamics (i.e., latent variables) behind neural activity by unsupervised learning with Bayesian population decoding using artificial neural networks or gaussian processes. Recently, persistent cohomology has been used to estimate latent variables from the phase information (i.e., circular coordinates) of manifolds created by neural activity. However, the advantages of persistent cohomology over Bayesian population decoding are not well understood. We compared persistent cohomology and Bayesian population decoding in estimating the animal location from simulated and actual grid cell population activity. We found that persistent cohomology can estimate the animal location with fewer neurons than Bayesian population decoding and robustly estimate the animal location from actual noisy data. Daisuke Kawahara, Shigeyoshi Fujisawa |
Neural Comput. | 1 |
| 2022 | JGLUE: Japanese General Language Understanding EvaluationabstractTo develop high-performance natural language understanding (NLU) models, it is necessary to have a benchmark to evaluate and analyze NLU ability from various perspectives. While the English NLU benchmark, GLUE, has been the forerunner, benchmarks are now being released for languages other than English, such as CLUE for Chinese and FLUE for French; but there is no such benchmark for Japanese. We build a Japanese NLU benchmark, JGLUE, from scratch without translation to measure the general NLU ability in Japanese. We hope that JGLUE will facilitate NLU research in Japanese. Kentaro Kurihara, Daisuke Kawahara, Tomohide Shibata |
LREC | 2 |
| 2022 | RODA: Reverse Operation Based Data Augmentation for Solving Math Word ProblemsabstractAutomatically solving math word problems is a critical task in the field of natural language processing. Recent models have reached their performance bottleneck and require more high-quality data for training. We propose a novel data augmentation method that reverses the mathematical logic of math word problems to produce new high-quality math problems and introduce new knowledge points that can benefit learning the mathematical reasoning logic. We apply the augmented data on two SOTA math word problem solving models and compare our results with a strong data augmentation baseline. Experimental results show the effectiveness of our approach (we release our code and data athttps://github.com/yiyunya/RODA). Qianying Liu, Wenyu Guan, Sujian Li, Fei Cheng 0002, Daisuke Kawahara, Sadao Kurohashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | BERT-based Cohesion Analysis of Japanese TextsabstractThe meaning of natural language text is supported by cohesion among various kinds of entities, including coreference relations, predicate-argument structures, and bridging anaphora relations. However, predicate-argument structures for nominal predicates and bridging anaphora relations have not been studied well, and their analyses have been still very difficult. Recent advances in neural networks, in particular, self training-based language models including BERT (Devlin et al., 2019), have significantly improved many natural language processing tasks, making it possible to dive into the study on analysis of cohesion in the whole text. In this study, we tackle an integrated analysis of cohesion in Japanese texts. Our results significantly outperformed existing studies in each task, especially about 10 to 20 point improvement both for zero anaphora and coreference resolution. Furthermore, we also showed that coreference resolution is different in nature from the other tasks and should be treated specially. Nobuhiro Ueda, Daisuke Kawahara, Sadao Kurohashi |
COLING | 2 |
| 2020 | A Method for Building a Commonsense Inference Dataset based on Basic EventsabstractWe present a scalable, low-bias, and low-cost method for building a commonsense inference dataset that combines automatic extraction from a corpus and crowdsourcing.Each problem is a multiple-choice question that asks contingency between basic events.We applied the proposed method to a Japanese corpus and acquired 104k problems.While humans can solve the resulting problems with high accuracy (88.9%), the accuracy of a highperformance transfer learning model is reasonably low (76.0%).We also confirmed through dataset analysis that the resulting dataset contains low bias.We released the datatset to facilitate language understanding research.1 Kazumasa Omura, Daisuke Kawahara, Sadao Kurohashi |
EMNLP (1) | 2 |
| 2020 | Acquiring Social Knowledge about Personality and Driving-related BehaviorabstractIn this paper, we introduce our psychological approach to collect human-specific social knowledge from a text corpus, using NLP techniques. It is often not explicitly described but shared among people, which we call social knowledge. We focus on the social knowledge, especially personality and driving. We used the language resources that were developed based on psychological research methods; a Japanese personality dictionary (317 words) and a driving experience corpus (8,080 sentences) annotated with behavior and subjectivity. Using them, we automatically extracted collocations between personality descriptors and driving-related behavior from a driving behavior and subjectivity corpus (1,803,328 sentences after filtering) and obtained unique 5,334 collocations. To evaluate the collocations as social knowledge, we designed four step-by-step crowdsourcing tasks. They resulted in 266 pieces of social knowledge. They include the knowledge that might be difficult to recall by themselves but easy to agree with. We discuss the acquired social knowledge and the contribution to implementations into systems. Ritsuko Iwai, Daisuke Kawahara, Takatsune Kumada, Sadao Kurohashi |
LREC | 2 |
| 2020 | Development of a Japanese Personality Dictionary based on Psychological MethodsabstractWe propose a new approach to constructing a personality dictionary with psychological evidence. In this study, we collect personality words, using word embeddings, and construct a personality dictionary with weights for Big Five traits. The weights are calculated based on the responses of the large sample (N=1,938, female = 1,004, M=49.8years old:20-78, SD=16.3). All the respondents answered a 20-item personality questionnaire and 537 personality items derived from word embeddings. We present the procedures to examine the qualities of responses with psychological methods and to calculate the weights. These result in a personality dictionary with two sub-dictionaries. We also discuss an application of the acquired resources. Ritsuko Iwai, Daisuke Kawahara, Takatsune Kumada, Sadao Kurohashi |
LREC | 2 |
| 2019 | Tree-structured Decoding for Solving Math Word ProblemsabstractQianying Liu, Wenyv Guan, Sujian Li, Daisuke Kawahara. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Qianying Liu, Wenyu Guan, Sujian Li, Daisuke Kawahara |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Emotion helps Sentiment: A Multi-task Model for Sentiment and Emotion AnalysisabstractIn this paper, we propose a two-layered multi-task attention based neural network that performs sentiment analysis through emotion analysis. The proposed approach is based on Bidirectional Long Short-Term Memory and uses Distributional Thesaurus as a source of external knowledge to improve the sentiment and emotion prediction. The proposed system has two levels of attention to hierarchically build a meaningful representation. We evaluate our system on the benchmark dataset of SemEval 2016 Task 6 and also compare it with the state-of-the-art systems on Stance Sentiment Emotion Corpus. Experimental results show that the proposed system improves the performance of sentiment analysis by 3.2 F-score points on SemEval 2016 Task 6 dataset. Our network also boosts the performance of emotion analysis by 5 F-score points on Stance Sentiment Emotion Corpus. Asif Ekbal, Daisuke Kawahara, Sadao Kurohashi |
IJCNN | 3 |
| 2019 | Applying Machine Translation to Psychology: Automatic Translation of Personality Adjectives
Ritsuko Iwai, Daisuke Kawahara, Takatsune Kumada, Sadao Kurohashi |
MTSummit (2) | 2 |
| 2018 | Neural Adversarial Training for Semi-supervised Japanese Predicate-argument Structure AnalysisabstractJapanese predicate-argument structure (PAS) analysis involves zero anaphora resolution, which is notoriously difficult.To improve the performance of Japanese PAS analysis, it is straightforward to increase the size of corpora annotated with PAS.However, since it is prohibitively expensive, it is promising to take advantage of a large amount of raw corpora.In this paper, we propose a novel Japanese PAS analysis model based on semi-supervised adversarial training with a raw corpus.In our experiments, our model outperforms existing state-of-the-art models for Japanese PAS analysis. Shuhei Kurita, Daisuke Kawahara, Sadao Kurohashi |
ACL (1) | 2 |
| 2018 | Cross-lingual Knowledge Projection Using Machine Translation and Target-side Knowledge Base CompletionabstractConsiderable effort has been devoted to building commonsense knowledge bases. However, they are not available in many languages because the construction of KBs is expensive. To bridge the gap between languages, this paper addresses the problem of projecting the knowledge in English, a resource-rich language, into other languages, where the main challenge lies in projection ambiguity. This ambiguity is partially solved by machine translation and target-side knowledge base completion, but neither of them is adequately reliable by itself. We show their combination can project English commonsense knowledge into Japanese and Chinese with high precision. Our method also achieves a top-10 accuracy of 90% on the crowdsourced English–Japanese benchmark. Furthermore, we use our method to obtain 18,747 facts of accurate Japanese commonsense within a very short period. Naoki Otani, Hirokazu Kiyomaru, Daisuke Kawahara, Sadao Kurohashi |
COLING | 3 |
| 2018 | Improving Crowdsourcing-Based Annotation of Japanese Discourse Relations
Yudai Kishimoto, Shinnosuke Sawada, Yugo Murawaki, Daisuke Kawahara, Sadao Kurohashi |
LREC | 4 |
| 2018 | JFCKB: Japanese Feature Change Knowledge Base
Tetsuaki Nakamura, Daisuke Kawahara |
LREC | 2 |
| 2018 | JDCFC: A Japanese Dialogue Corpus with Feature Changes
Tetsuaki Nakamura, Daisuke Kawahara |
LREC | 2 |
| 2018 | Comprehensive Annotation of Various Types of Temporal Information on the Time Axis
Tomohiro Sakaguchi, Daisuke Kawahara, Sadao Kurohashi |
LREC | 2 |
| 2018 | Annotating a Driving Experience Corpus with Behavior and Subjectivity
Ritsuko Iwai, Daisuke Kawahara, Takatsune Kumada, Sadao Kurohashi |
PACLIC | 2 |
| 2017 | Neural Joint Model for Transition-based Chinese Syntactic AnalysisabstractWe present neural network-based joint models for Chinese word segmentation, POS tagging and dependency parsing.Our models are the first neural approaches for fully joint Chinese analysis that is known to prevent the error propagation problem of pipeline models.Although word embeddings play a key role in dependency parsing, they cannot be applied directly to the joint task in the previous work.To address this problem, we propose embeddings of character strings, in addition to words.Experiments show that our models outperform existing systems in Chinese word segmentation and POS tagging, and perform preferable accuracies in dependency parsing.We also explore bi-LSTM models with fewer features. Shuhei Kurita, Daisuke Kawahara, Sadao Kurohashi |
ACL (1) | 2 |
| 2017 | Improving Chinese Semantic Role Labeling using High-quality Surface and Deep Case FramesabstractThis paper presents a method for improving semantic role labeling (SRL) using a large amount of automatically acquired knowledge.We acquire two varieties of knowledge, which we call surface case frames and deep case frames.Although the surface case frames are compiled from syntactic parses and can be used as rich syntactic knowledge, they have limited capability for resolving semantic ambiguity.To compensate the deficiency of the surface case frames, we compile deep case frames from automatic semantic roles.We also consider quality management for both types of knowledge in order to get rid of the noise brought from the automatic analyses.The experimental results show that Chinese SRL can be improved using automatically acquired knowledge and the quality management shows a positive effect on this task. Gongye Jin, Daisuke Kawahara, Sadao Kurohashi |
EACL (1) | 2 |
| 2016 | Neural Network-Based Model for Japanese Predicate Argument Structure Analysis
Tomohide Shibata, Daisuke Kawahara, Sadao Kurohashi |
ACL (1) | 2 |
| 2016 | Age Related Differences in Episodic Memory Recollections: Applying Latent Dirichlet Allocation to Free-Writings on Driving Incidents by Older and Young Drivers
Ritsuko Iwai, Takatsune Kumada, Daisuke Kawahara, Sadao Kurohashi |
CogSci | 3 |
| 2016 | Consistent Word Segmentation, Part-of-Speech Tagging and Dependency Labelling Annotation for Chinese LanguageabstractIn this paper, we propose a new annotation approach to Chinese word segmentation, part-of-speech (POS) tagging and dependency labelling that aims to overcome the two major issues in traditional morphology-based annotation: Inconsistency and data sparsity. We re-annotate the Penn Chinese Treebank 5.0 (CTB5) and demonstrate the advantages of this approach compared to the original CTB5 annotation through word segmentation, POS tagging and machine translation experiments. Mo Shen, Wingmui Li, HyunJeong Choe, Chenhui Chu, Daisuke Kawahara, Sadao Kurohashi |
COLING | 5 |
| 2016 | IRT-based Aggregation Model of Crowdsourced Pairwise Comparison for Evaluating Machine TranslationsabstractRecent work on machine translation has used crowdsourcing to reduce costs of manual evaluations.However, crowdsourced judgments are often biased and inaccurate.In this paper, we present a statistical model that aggregates many manual pairwise comparisons to robustly measure a machine translation system's performance.Our method applies graded response model from item response theory (IRT), which was originally developed for academic tests.We conducted experiments on a public dataset from the Workshop on Statistical Machine Translation 2013, and found that our approach resulted in highly interpretable estimates and was less affected by noisy judges than previously proposed methods. Naoki Otani, Toshiaki Nakazawa, Daisuke Kawahara, Sadao Kurohashi |
EMNLP | 3 |
| 2015 | Morphological Analysis for Unsegmented Languages using Recurrent Neural Network Language ModelabstractWe present a new morphological analy-sis model that considers semantic plausi-bility of word sequences by using a re-current neural network language model (RNNLM). In unsegmented languages, since language models are learned from automatically segmented texts and in-evitably contain errors, it is not apparent that conventional language models con-tribute to morphological analysis. To solve this problem, we do not use language mod-els based on raw word sequences but use a semantically generalized language model, RNNLM, in morphological analysis. In our experiments on two Japanese corpora, our proposed model significantly outper-formed baseline models. This result indi-cates the effectiveness of RNNLM in mor-phological analysis. 1 Hajime Morita, Daisuke Kawahara, Sadao Kurohashi |
EMNLP | 2 |
| 2014 | A Step-wise Usage-based Method for Inducing Polysemy-aware Verb ClassesabstractWe present an unsupervised method for in-ducing verb classes from verb uses in giga-word corpora. Our method consists of two clustering steps: verb-specific seman-tic frames are first induced by clustering verb uses in a corpus and then verb classes are induced by clustering these frames. By taking this step-wise approach, we can not only generate verb classes based on a massive amount of verb uses in a scalable manner, but also deal with verb polysemy, which is bypassed by most of the previous studies on verb clustering. In our exper-iments, we acquire semantic frames and verb classes from two giga-word corpora, the larger comprising 20 billion words. The effectiveness of our approach is veri-fied through quantitative evaluations based on polysemy-aware gold-standard data. 1 Daisuke Kawahara, Daniel W. Peterson, Martha Palmer |
ACL (1) | 1 |
| 2014 | Rapid Development of a Corpus with Discourse Annotations using Two-stage Crowdsourcing
Daisuke Kawahara, Yuichiro Machida, Tomohide Shibata, Sadao Kurohashi, Hayato Kobayashi, Manabu Sassano |
COLING | 1 |
| 2014 | Inducing Example-based Semantic Frames from a Massive Amount of Verb UsesabstractWe present an unsupervised method for inducing semantic frames from verb uses in giga-word corpora.Our semantic frames are verb-specific example-based frames that are distinguished according to their senses.We use the Chinese Restaurant Process to automatically induce these frames from a massive amount of verb instances.In our experiments, we acquire broad-coverage semantic frames from two giga-word corpora, the larger comprising 20 billion words.Our experimental results indicate the effectiveness of our approach. Daisuke Kawahara, Daniel W. Peterson, Octavian Popescu, Martha Palmer |
EACL | 1 |
| 2014 | A Framework for Compiling High Quality Knowledge Resources From Raw Corpora
Gongye Jin, Daisuke Kawahara, Sadao Kurohashi |
LREC | 2 |
| 2014 | Single Classifier Approach for Verb Sense Disambiguation based on Generalized Features
Daisuke Kawahara, Martha Palmer |
LREC | 1 |
| 2014 | Dependency Parse Reranking with Rich Subtree FeaturesabstractIn pursuing machine understanding of human language, highly accurate syntactic analysis is a crucial step. In this work, we focus on dependency grammar, which models syntax by encoding transparent predicate-argument structures. Recent advances in dependency parsing have shown that employing higher-order subtree structures in graph-based parsers can substantially improve the parsing accuracy. However, the inefficiency of this approach increases with the order of the subtrees. This work explores a new reranking approach for dependency parsing that can utilize complex subtree representations by applying efficient subtree selection methods. We demonstrate the effectiveness of the approach in experiments conducted on the Penn Treebank and the Chinese Treebank. Our system achieves the best performance among known supervised systems evaluated on these datasets, improving the baseline accuracy from 91.88% to 93.42% for English, and from 87.39% to 89.25% for Chinese. Mo Shen, Daisuke Kawahara, Sadao Kurohashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Japanese Zero Reference Resolution Considering Exophora and Author/Reader MentionsabstractIn Japanese, zero references often occur and many of them are categorized into zero exophora, in which a referent is not mentioned in the document.However, previous studies have focused on only zero endophora, in which a referent explicitly appears.We present a zero reference resolution model considering zero exophora and author/reader of a document.To deal with zero exophora, our model adds pseudo entities corresponding to zero exophora to candidate referents of zero pronouns.In addition, we automatically detect mentions that refer to the author and reader of a document by using lexico-syntactic patterns.We represent their particular behavior in a discourse as a feature vector of a machine learning model.The experimental results demonstrate the effectiveness of our model for not only zero exophora but also zero endophora. Masatsugu Hangyo, Daisuke Kawahara, Sadao Kurohashi |
EMNLP | 2 |
| 2013 | Automatic Knowledge Acquisition for Case Alternation between the Passive and Active Voices in JapaneseabstractWe present a method for automatically acquiring knowledge for case alternation between the passive and active voices in Japanese.By leveraging several linguistic constraints on alternation patterns and lexical case frames obtained from a large Web corpus, our method aligns a case frame in the passive voice to a corresponding case frame in the active voice and finds an alignment between their cases.We then apply the acquired knowledge to a case alternation task and prove its usefulness. Ryohei Sasano, Daisuke Kawahara, Sadao Kurohashi, Manabu Okumura |
EMNLP | 2 |
| 2013 | High Quality Dependency Selection from Automatic Parses
Gongye Jin, Daisuke Kawahara, Sadao Kurohashi |
IJCNLP | 2 |
| 2013 | Precise Information Retrieval Exploiting Predicate-Argument Structures
Daisuke Kawahara, Keiji Shinzato, Tomohide Shibata, Sadao Kurohashi |
IJCNLP | 1 |
| 2013 | Chinese Word Segmentation by Mining Maximized Substrings
Mo Shen, Daisuke Kawahara, Sadao Kurohashi |
IJCNLP | 2 |
| 2013 | Chinese-Japanese Machine Translation Exploiting Chinese CharactersabstractThe Chinese and Japanese languages share Chinese characters. Since the Chinese characters in Japanese originated from ancient China, many common Chinese characters exist between these two languages. Since Chinese characters contain significant semantic information and common Chinese characters share the same meaning in the two languages, they can be quite useful in Chinese-Japanese machine translation (MT). We therefore propose a method for creating a Chinese character mapping table for Japanese, traditional Chinese, and simplified Chinese, with the aim of constructing a complete resource of common Chinese characters. Furthermore, we point out two main problems in Chinese word segmentation for Chinese-Japanese MT, namely, unknown words and word segmentation granularity, and propose an approach exploiting common Chinese characters to solve these problems. We also propose a statistical method for detecting other semantically equivalent Chinese characters other than the common ones and a method for exploiting shared Chinese characters in phrase alignment. Results of the experiments carried out on a state-of-the-art phrase-based statistical MT system and an example-based MT system show that our proposed approaches can improve MT performance significantly, thereby verifying the effectiveness of shared Chinese characters for Chinese-Japanese MT. Chenhui Chu, Toshiaki Nakazawa, Daisuke Kawahara, Sadao Kurohashi |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2012 | Exploiting Shared Chinese Characters in Chinese Word Segmentation Optimization for Chinese-Japanese Machine Translation
Chenhui Chu, Toshiaki Nakazawa, Daisuke Kawahara, Sadao Kurohashi |
EAMT | 3 |
| 2012 | Building a Diverse Document Leads Corpus Annotated with Semantic Relations
Masatsugu Hangyo, Daisuke Kawahara, Sadao Kurohashi |
PACLIC | 2 |
| 2012 | A Reranking Approach for Dependency Parsing with Variable-sized Subtree Features
Mo Shen, Daisuke Kawahara, Sadao Kurohashi |
PACLIC | 2 |
| 2011 | Generative Modeling of Coordination by Factoring Parallelism and Selectional Preferences
Daisuke Kawahara, Sadao Kurohashi |
IJCNLP | 1 |
| 2010 | Acquiring Reliable Predicate-argument Structures from Raw Corpora for Case Frame Compilation
Daisuke Kawahara, Sadao Kurohashi |
LREC | 1 |
| 2009 | Mining Parallel Texts from Mixed-Language Web Pages
Masao Utiyama, Daisuke Kawahara, Keiji Yasuda, Eiichiro Sumita |
MTSummit | 2 |
| 2009 | The Effect of Corpus Size on Case Frame Acquisition for Discourse Analysis
Ryohei Sasano, Daisuke Kawahara, Sadao Kurohashi |
HLT-NAACL | 2 |
| 2009 | Identifying Information Sender Configuration of Web PagesabstractThe source of a piece of information is a crucial element to consider when judging the credibility of that information. In this paper, we address the task of identifying the information source which is cast as a problem of identifying the {\em information sender configuration (ISC)} of a Web page. An information sender of a Web page is an entity which is involved in the publication of the information on the page. An ISC of a Web page describes the information senders of the page and the relationship among them. Information sender extraction is thus a subtask of identifying ISC, and we present a method for extracting information senders from Web pages and offer preliminary evaluation. The ISC provides a basis for deeper analysis of information on the Web. Yoshikiyo Kato, Daisuke Kawahara, Kentaro Inui, Sadao Kurohashi, Tomohide Shibata |
Web Intelligence | 2 |
| 2009 | Using Short Dependency Relations from Auto-Parsed Data for Chinese Dependency ParsingabstractDependency parsing has become increasingly popular for a surge of interest lately for applications such as machine translation and question answering. Currently, several supervised learning methods can be used for training high-performance dependency parsers if sufficient labeled data are available. However, currently used statistical dependency parsers provide poor results for words separated by long distances. In order to solve this problem, this article presents an effective dependency parsing approach of incorporating short dependency information from unlabeled data. The unlabeled data is automatically parsed by using a deterministic dependency parser, which exhibits a relatively high performance for short dependencies between words. We then train another parser that uses the information on short dependency relations extracted from the output of the first parser. The proposed approach achieves an unlabeled attachment score of 86.52%, an absolute 1.24% improvement over the baseline system on the Chinese Treebank data set. The results indicate that the proposed approach improves the parsing performance for longer distance words. Wenliang Chen, Daisuke Kawahara, Kiyotaka Uchimoto, Hitoshi Isahara |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2008 | Coordination Disambiguation without Any Similarities
Daisuke Kawahara, Sadao Kurohashi |
COLING | 1 |
| 2008 | A Fully-Lexicalized Probabilistic Model for Japanese Zero Anaphora Resolution
Ryohei Sasano, Daisuke Kawahara, Sadao Kurohashi |
COLING | 2 |
| 2008 | Chinese Dependency Parsing with Large Scale Automatically Constructed Case Structures
Daisuke Kawahara, Sadao Kurohashi |
COLING | 2 |
| 2008 | Construction of an Idiom Corpus and its Application to Idiom Identification based on WSD Incorporating Idiom-Specific Features
Chikara Hashimoto, Daisuke Kawahara |
EMNLP | 2 |
| 2008 | Dependency Parsing with Short Dependency Relations in Unlabeled Data
Wenliang Chen, Daisuke Kawahara, Kiyotaka Uchimoto, Hitoshi Isahara |
IJCNLP | 2 |
| 2008 | Learning Reliability of Parses for Domain Adaptation of Dependency Parsing
Daisuke Kawahara, Kiyotaka Uchimoto |
IJCNLP | 1 |
| 2008 | TSUBAKI: An Open Search Engine Infrastructure for Developing New Information Access Methodology
Keiji Shinzato, Tomohide Shibata, Daisuke Kawahara, Chikara Hashimoto, Sadao Kurohashi |
IJCNLP | 3 |
| 2008 | A Method for Automatically Constructing Case Frames for English
Daisuke Kawahara, Kiyotaka Uchimoto |
LREC | 1 |
| 2008 | A Large-Scale Web Data Collection as a Natural Language Processing Infrastructure
Keiji Shinzato, Daisuke Kawahara, Chikara Hashimoto, Sadao Kurohashi |
LREC | 2 |
| 2008 | Grasping Major Statements and Their Contradictions Toward Information Credibility Analysis of Web ContentsabstractThe World Wide Web contains wide variety of news reports, arguments, opinions, etc. that vary widely in quality. People judge the credibility of information on the Web for decision making in daily life. At present, while the quantity of information on the Web is explosively increasing, it is necessary to develop a system that supports such judgments. We have been developing an information credibility analysis system, WISDOM that considers the viewpoints of information contents, information senders, and information appearances. In this paper, as a viewpoint of information contents, we propose a method for providing a bird's eye view of major statements on a given topic and their contradictions. We evaluate the obtained statements in our experiments, and confirm the effectiveness of our approach. Furthermore, we discuss our future objectives. Daisuke Kawahara, Sadao Kurohashi, Kentaro Inui |
Web Intelligence | 1 |
| 2008 | Enriching Multilingual Language Resources by Discovering Missing Cross-Language Links in WikipediaabstractWe present a novel method for discovering missing cross-language links between English and Japanese Wikipedia articles. We collect candidates of missing cross-language links -- a pair of English and Japanese Wikipedia articles, which could be connected by cross-language links. Then we select the correct cross-language links among the candidates by using a classifier trained with various types of features. Our method has three desirable characteristics for discovering missing links. First, our method can discover cross-language links with high accuracy (92\% precision with 78\% recall rates). Second, the features used in a classifier are language-independent. Third, without relying on any external knowledge, we generate the features based on resources automatically obtained from Wikipedia. In this work, we discover approximately $10^5$ missing cross-language links from Wikipedia, which are almost two-thirds as many as the existing cross-language links in Wikipedia. Jong-Hoon Oh, Daisuke Kawahara, Kiyotaka Uchimoto, Jun'ichi Kazama, Kentaro Torisawa |
Web Intelligence | 2 |
| 2007 | Minimally Lexicalized Dependency Parsing
Daisuke Kawahara, Kiyotaka Uchimoto |
ACL | 1 |
| 2007 | Probabilistic Coordination Disambiguation in a Fully-Lexicalized Japanese Parser
Daisuke Kawahara, Sadao Kurohashi |
EMNLP-CoNLL | 1 |
| 2006 | Case Frame Compilation from the Web using High-Performance Computing
Daisuke Kawahara, Sadao Kurohashi |
LREC | 1 |
| 2006 | A Fully-Lexicalized Probabilistic Model for Japanese Syntactic and Case Structure Analysis
Daisuke Kawahara, Sadao Kurohashi |
HLT-NAACL | 1 |
| 2006 | Cards-to-presentation on the web: generating multimedia contents featuring agent animations
Yukiko I. Nakano, Toshihiro Murayama, Masashi Okamoto, Daisuke Kawahara, Sadao Kurohashi, Toyoaki Nishida |
J. Netw. Comput. Appl. | 4 |
| 2005 | PP-Attachment Disambiguation Boosted by a Gigantic Volume of Unambiguous Examples
Daisuke Kawahara, Sadao Kurohashi |
IJCNLP | 1 |
| 2005 | Automatic Acquisition of Basic Katakana Lexicon from a Given Corpus
Toshiaki Nakazawa, Daisuke Kawahara, Sadao Kurohashi |
IJCNLP | 2 |
| 2004 | Improving Japanese Zero Pronoun Resolution by Global Word Sense Disambiguation
Daisuke Kawahara, Sadao Kurohashi |
COLING | 1 |
| 2004 | Automatic Construction of Nominal Case Frames and its Application to Indirect Anaphora Resolution
Ryohei Sasano, Daisuke Kawahara, Sadao Kurohashi |
COLING | 2 |
| 2004 | Zero Pronoun Resolution Based on Automatically Constructed Case Frames and Structural Preference of Antecedents
Daisuke Kawahara, Sadao Kurohashi |
IJCNLP | 1 |
| 2004 | Structural Analysis of Instruction Utterances Using Linguistic and Visual Information
Tomohide Shibata, Masato Tachiki, Daisuke Kawahara, Masashi Okamoto, Sadao Kurohashi, Toyoaki Nishida |
KES | 3 |
| 2004 | Toward Text Understanding: Integrating Relevance-tagged Corpus and Automatically Constructed Case Frames
Daisuke Kawahara, Ryohei Sasano, Sadao Kurohashi |
LREC | 1 |
| 2003 | Embodied Conversational Agents for Presenting Intellectual Multimedia Contents
Yukiko I. Nakano, Toshihiro Murayama, Daisuke Kawahara, Sadao Kurohashi, Toyoaki Nishida |
KES | 3 |
| 2003 | Structural Analysis of Instruction Utterances
Tomohide Shibata, Daisuke Kawahara, Masashi Okamoto, Sadao Kurohashi, Toyoaki Nishida |
KES | 2 |
| 2002 | Verb Paraphrase based on Case Frame AlignmentabstractThis paper describes a method of translating a predicate-argument structure of a verb into that of an equivalent verb, which is a core component of the dictionary-based paraphrasing. Our method grasps several usages of a headword and those of the def-heads as a form of their case frames and aligns those case frames, which means the acquisition of word sense disambiguation rules and the detection of the appropriate equivalent and case marker transformation. Nobuhiro Kaji, Daisuke Kawahara, Sadao Kurohashi, Satoshi Sato |
ACL | 2 |
| 2002 | Fertilization of Case Frame Dictionary for Robust Japanese Case Analysis
Daisuke Kawahara, Sadao Kurohashi |
COLING | 1 |
| 2002 | Construction of a Japanese Relevance-tagged Corpus
Daisuke Kawahara, Sadao Kurohashi, Kôiti Hasida |
LREC | 1 |
| 2000 | Japanese Case Structure Analysis
Daisuke Kawahara, Nobuhiro Kaji, Sadao Kurohashi |
COLING | 1 |