Masayuki Asahara

dblp:22/5249 · DBLP profile ↗
← Back
51ranked-venue papers
11as first author
15since 2021 · last 2026
0000-0002-5178-7275ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 11 first-author · 15 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 All-words pronunciation estimation of Japanese homographs
Kanako Komiya, Taichiro Kobayashi, Masayuki Asahara, Hiroyuki Shinnou
Data Knowl. Eng.3
2025 Exploring spatial and temporal dynamics of language comprehension in the brain with CCG
Shinnosuke Isono, Kohei Kajikawa, Yushi Sugimoto, Masayuki Asahara, Yohei Oseki
CogSci4
2025 Large-Scale Japanese Metaphor Corpus Construction: Expanding BCCWJ-Metaphor with Automated Annotation
Rowan Hall Maudslay, Kanako Komiya, Sachi Kato, Masayuki Asahara
PACLIC5
2024 Assigning Impression Rating Information to the 'Balanced Corpus of Contemporary Written Japanese'
Sachi Kato, Masayuki Asahara
PACLIC2
2023 Investigation of Information Processing Mechanisms in the Human Brain During Reading Tanka Poetry
Anna Sato, Junichi Chikazoe, Shotaro Funai, Daichi Mochihashi, Yutaka Shikano, Masayuki Asahara, Satoshi Iso, Ichiro Kobayashi 0001
ICANN (8)6
2023 All-Words Word Sense Disambiguation for Historical Japanese
Soma Asada, Kanako Komiya, Masayuki Asahara
PACLIC3
2023 Word Familiarity Rate Estimation for Japanese Functional Words Using a Bayesian Linear Mixed Model
Bocheng Chen, Masayuki Asahara
PACLIC2
2023 Spatial Information Annotation Based on the Double Cross Model
Yoshiko Kawabata, Mai Omura, Masayuki Asahara, Johane Takeuchi
PACLIC3
2023 UD_Japanese-CEJC: Dependency Relation Annotation on Corpus of Everyday Japanese Conversation
abstract
In this study, we have developed Universal Dependencies (UD) resources for spoken Japanese in the Corpus of Everyday Japanese Conversation (CEJC).The CEJC is a large corpus of spoken language that encompasses various everyday conversations in Japanese, and includes word delimitation and part-ofspeech annotation.We have newly annotated Long Word Unit delimitation and Bunsetsu (Japanese phrase)-based dependencies, including Bunsetsu boundaries, for CEJC.The UD of Japanese resources was constructed in accordance with hand-maintained conversion rules from the CEJC with two types of word delimitation, part-of-speech tags and Bunsetsu-based syntactic dependency relations.Furthermore, we examined various issues pertaining to the construction of UD in the CEJC by comparing it with the written Japanese corpus and evaluating UD parsing accuracy.
Mai Omura, Hiroshi Matsuda, Masayuki Asahara, Aya Wakasa
SIGDIAL3
2022 Reading Time and Vocabulary Rating in the Japanese Language: Large-Scale Japanese Reading Time Data Collection Using Crowdsourcing
abstract
This study examines how differences in human vocabulary affect reading time. Specifically, we assumed vocabulary to be the random effect of research participants when applying a generalized linear mixed model to the ratings of participants in the word familiarity survey. Thereafter, we asked the participants to take part in a self-paced reading task to collect their reading times. Through fixed effect of vocabulary when applying a generalized linear mixed model to reading time, we clarified the tendency that vocabulary differences give to reading time.
Masayuki Asahara
LREC1
2022 Word Sense Disambiguation of Corpus of Historical Japanese Using Japanese BERT Trained with Contemporary Texts
Kanako Komiya, Nagi Oki, Masayuki Asahara
PACLIC3
2021 Lower Perplexity is Not Always Human-Like
abstract
Tatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida, Masayuki Asahara, Kentaro Inui. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Tatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida, Masayuki Asahara, Kentaro Inui
ACL/IJCNLP (1)5
2021 Dependency Enhanced Contextual Representations for Japanese Temporal Relation Classification
Chenjing Geng, Fei Cheng 0002, Masayuki Asahara, Lis Pereira, Ichiro Kobayashi 0001
PACLIC3
2021 The Annotation of Antonym Information in the 'Word List by Semantic Principles'
Sachi Kato, Masayuki Asahara, Nanami Moriyama, Makoto Yamazaki, Asami Ogiwara
PACLIC2
2021 ALICE++: Adversarial Training for Robust and Effective Temporal Reasoning
Lis Pereira, Fei Cheng 0002, Masayuki Asahara, Ichiro Kobayashi 0001
PACLIC3
2020 KOTONOHA: A Corpus Concordance System for Skewer-Searching NINJAL Corpora
abstract
The National Institute for Japanese Language and Linguistics, Japan (NINJAL, Japan), has developed several types of corpora. For each corpus NINJAL provided an online search environment, ‘Chunagon’, which is a morphological-information-annotation-based concordance system made publicly available in 2011. NINJAL has now provided a skewer-search system ‘Kotonoha’ based on the ‘Chunagon’ systems. This system enables querying of multiple corpora by certain categories, such as register type and period.
Teruaki Oka, Yuichi Ishimoto, Yutaka Yagi, Takenori Nakamura, Masayuki Asahara, Kikuo Maekawa, Toshinobu Ogiso, Hanae Koiso, Kumiko Sakoda, Nobuko Kibe
LREC5
2020 Design of BCCWJ-EEG: Balanced Corpus with Human Electroencephalography
abstract
The past decade has witnessed the happy marriage between natural language processing (NLP) and the cognitive science of language. Moreover, given the historical relationship between biological and artificial neural networks, the advent of deep learning has re-sparked strong interests in the fusion of NLP and the neuroscience of language. Importantly, this inter-fertilization between NLP, on one hand, and the cognitive (neuro)science of language, on the other, has been driven by the language resources annotated with human language processing data. However, there remain several limitations with those language resources on annotations, genres, languages, etc. In this paper, we describe the design of a novel language resource called BCCWJ-EEG, the Balanced Corpus of Contemporary Written Japanese (BCCWJ) experimentally annotated with human electroencephalography (EEG). Specifically, after extensively reviewing the language resources currently available in the literature with special focus on eye-tracking and EEG, we summarize the details concerning (i) participants, (ii) stimuli, (iii) procedure, (iv) data preprocessing, (v) corpus evaluation, (vi) resource release, and (vii) compilation schedule. In addition, potential applications of BCCWJ-EEG to neuroscience and NLP will also be discussed.
Yohei Oseki, Masayuki Asahara
LREC2
2020 Composing Word Vectors for Japanese Compound Words Using Bilingual Word Embeddings
Teruo Hirabayashi, Kanako Komiya, Masayuki Asahara, Hiroyuki Shinnou
PACLIC3
2020 Generation and Evaluation of Concept Embeddings Via Fine-Tuning Using Automatically Tagged Corpus
Kanako Komiya, Daiki Yaginuma, Masayuki Asahara, Hiroyuki Shinnou
PACLIC3
2018 Universal Dependencies Version 2 for Japanese
Masayuki Asahara, Hiroshi Kanayama, Takaaki Tanaka, Yusuke Miyao, Sumire Uematsu, Shinsuke Mori, Yuji Matsumoto 0001, Mai Omura, Yugo Murawaki
LREC1
2018 All-words Word Sense Disambiguation Using Concept Embeddings
Rui Suzuki, Kanako Komiya, Masayuki Asahara, Minoru Sasaki, Hiroyuki Shinnou
LREC3
2018 Between Reading Time and Clause Boundaries in Japanese - Wrap-up Effect in a Head-Final Language
Masayuki Asahara
PACLIC1
2018 Annotation of 'Word List by Semantic Principles' Labels for the Balanced Corpus of Contemporary Written Japanese
Sachi Kato, Masayuki Asahara, Makoto Yamazaki
PACLIC2
2017 Between Reading Time and Syntactic/Semantic Categories
abstract
This article presents a contrastive analysis between reading time and syntactic/semantic categories in Japanese. We overlaid the reading time annotation of BCCWJ-EyeTrack and a syntactic/semantic category information annotation on the ‘Balanced Corpus of Contemporary Written Japanese’. Statistical analysis based on a mixed linear model showed that verbal phrases tend to have shorter reading times than adjectives, adverbial phrases, or nominal phrases. The results suggest that the preceding phrases associated with the presenting phrases promote the reading process to shorten the gazing time.
Masayuki Asahara, Sachi Kato
IJCNLP(1)1
2017 Between Reading Time and Information Structure
Masayuki Asahara
PACLIC1
2016 Reading-Time Annotations for "Balanced Corpus of Contemporary Written Japanese"
abstract
The Dundee Eyetracking Corpus contains eyetracking data collected while native speakers of English and French read newspaper editorial articles. Similar resources for other languages are still rare, especially for languages in which words are not overtly delimited with spaces. This is a report on a project to build an eyetracking corpus for Japanese. Measurements were collected while 24 native speakers of Japanese read excerpts from the Balanced Corpus of Contemporary Written Japanese Texts were presented with or without segmentation (i.e. with or without space at the boundaries between bunsetsu segmentations) and with two types of methodologies (eyetracking and self-paced reading presentation). Readers’ background information including vocabulary-size estimation and Japanese reading-span score were also collected. As an example of the possible uses for the corpus, we also report analyses investigating the phenomena of anti-locality.
Masayuki Asahara, Hajime Ono, Edson T. Miyamoto
COLING1
2016 Universal Dependencies for Japanese
Takaaki Tanaka, Yusuke Miyao, Masayuki Asahara, Sumire Uematsu, Hiroshi Kanayama, Shinsuke Mori, Yuji Matsumoto 0001
LREC3
2013 BCCWJ-TimeBank: Temporal and Event Information Annotation on Japanese Text
Masayuki Asahara, Sachi Yasuda, Hikari Konishi, Mizuho Imada, Kikuo Maekawa
PACLIC1
2012 Head-driven Transition-based Parsing with Top-down Prediction
Katsuhiko Hayashi 0001, Taro Watanabe, Masayuki Asahara, Yuji Matsumoto 0001
ACL (1)3
2012 Mining Rules for Rewriting States in a Transition-Based Dependency Parser
Akihiro Inokuchi, Ayumu Yamaoka, Takashi Washio, Yuji Matsumoto 0001, Masayuki Asahara, Masakazu Iwatate, Hideto Kazawa
PRICAI5
2011 Third-order Variational Reranking on Packed-Shared Dependency Forests
Katsuhiko Hayashi 0001, Taro Watanabe, Masayuki Asahara, Yuji Matsumoto 0001
EMNLP3
2011 Jointly Extracting Japanese Predicate-Argument Relation with Markov Logic
Katsumasa Yoshikawa, Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP2
2009 Jointly Identifying Temporal Relations with Markov Logic
Katsumasa Yoshikawa, Sebastian Riedel 0001, Masayuki Asahara, Yuji Matsumoto 0001
ACL/IJCNLP3
2008 Japanese Dependency Parsing Using a Tournament Model
Masakazu Iwatate, Masayuki Asahara, Yuji Matsumoto 0001
COLING2
2008 A Pipeline Approach for Syntactic and Semantic Dependency Parsing
Yotaro Watanabe, Masakazu Iwatate, Masayuki Asahara, Yuji Matsumoto 0001
CoNLL3
2008 Use of Event Types for Temporal Relation Identification in Chinese Text
Yuchang Cheng, Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP2
2008 Analyzing Chinese Synthetic Words with Tree-based Information and a Survey on Chinese Morphologically Derived Words
Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP2
2008 Japanese-Spanish Thesaurus Construction Using English as a Pivot
Jessica C. Ramírez, Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP2
2007 A Graph-Based Approach to Named Entity Categorization in Wikipedia Using Conditional Random Fields
Yotaro Watanabe, Masayuki Asahara, Yuji Matsumoto 0001
EMNLP-CoNLL2
2007 Constructing a Temporal Relation Tagged Corpus of Chinese Based on Dependency Structure Analysis
abstract
This paper describes an annotation guideline for a temporal relation tagged corpus. Our goal is to construct a machine learnable model that automatically analyzes temporal events and relations between events. Since analyzing all combinations of events is inefficient, we examine use of dependency structure analysis to efficiently recognize meaningful temporal relations. We survey a small tagged data set to investigate the coverage of our method. Although the coverage of our methods is about 49%, we find that the dependency structure appears useful for reducing manual efforts in constructing a tagged corpus with temporal relations.
Yuchang Cheng, Masayuki Asahara, Yuji Matsumoto 0001
TIME2
2006 Multi-lingual Dependency Parsing at NAIST
Yuchang Cheng, Masayuki Asahara, Yuji Matsumoto 0001
CoNLL2
2006 An Annotated Corpus Management Tool: ChaKi
Yuji Matsumoto 0001, Masayuki Asahara, Kiyota Hashimoto, Yukio Tono, Akira Ohtani, Toshio Morita
LREC2
2006 The Construction of a Dictionary for a Two-layer Chinese Morphological Analyzer
Chooi-Ling Goh, Yuchang Cheng, Masayuki Asahara, Yuji Matsumoto 0001
PACLIC4
2005 Building a Japanese-Chinese Dictionary Using Kanji/Hanzi Conversion
Chooi-Ling Goh, Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP2
2005 Automatic Extraction of Fixed Multiword Expressions
Campbell Hore, Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP2
2004 Japanese Unknown Word Identification by Character-based Chunking
Masayuki Asahara, Yuji Matsumoto 0001
COLING1
2004 Deterministic Dependency Structure Analyzer for Chinese
Yuchang Cheng, Masayuki Asahara, Yuji Matsumoto 0001
IJCNLP2
2004 Pruning False Unknown Words to Improve Chinese Word Segmentation
Chooi-Ling Goh, Masayuki Asahara, Yuji Matsumoto 0001
PACLIC2
2003 Japanese Named Entity Extraction with Redundant Morphological Analysis
Masayuki Asahara, Yuji Matsumoto 0001
HLT-NAACL1
2002 Use of XML and Relational Databases for Consistent Development and Maintenance of Lexicons and Annotated Corpora
Masayuki Asahara, Ryuichi Yoneda, Akiko Yamashita, Yasuharu Den, Yuji Matsumoto 0001
LREC1
2000 Extended Models and Tools for High-performance Part-of-speech
Masayuki Asahara, Yuji Matsumoto 0001
COLING1