VLDB 2026 Research / reviewers in the wild / expert
Kanako Komiya
dblp:92/3903
· DBLP profile ↗
38ranked-venue papers
14as first author
13since 2021 · last 2026
0000-0001-6405-1067ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 14 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | All-words pronunciation estimation of Japanese homographs
Kanako Komiya, Taichiro Kobayashi, Masayuki Asahara, Hiroyuki Shinnou |
Data Knowl. Eng. | 1 |
| 2025 | Similarity-Based Scoring Model for Handwritten Answers in Japanese Workbooks
Takahiro Saito, Hung Tuan Nguyen, Kanako Komiya, Tsunenori Ishioka, Masaki Nakagawa |
AIED (5) | 3 |
| 2025 | Investigating the Influence of Automated Transcription Error Correction on Automated Grading for Handwritten Japanese Answers
Rina Suzuki, Takahiro Saito, Hisao Usui, Hiroaki Ozaki, Hung Tuan Nguyen, Kanako Komiya, Tsunenori Ishioka, Masaki Nakagawa |
AIED (6) | 6 |
| 2025 | Large-Scale Japanese Metaphor Corpus Construction: Expanding BCCWJ-Metaphor with Automated Annotation
Rowan Hall Maudslay, Kanako Komiya, Sachi Kato, Masayuki Asahara |
PACLIC | 3 |
| 2025 | Answer Generation for Large-Scale Official Trial Tests in Japanese University Entrance Exams Using RAG
Taketsuna Ichiyanagi, Kanako Komiya, Tsunenori Ishioka, Masaki Nakagawa |
PKAW | 2 |
| 2024 | Error Correction of Japanese Character-Recognition in Answers to Writing-Type Questions Using T5
Rina Suzuki, Hisao Usui, Hiroaki Ozaki, Hung Tuan Nguyen, Kanako Komiya, Tsunenori Ishioka, Masaki Nakagawa |
DAS | 5 |
| 2024 | All-Words Pronunciation Estimation of Japanese Homographs Using Automatically Tagged Data
Taichiro Kobayashi, Kanako Komiya, Hiroyuki Shinnou |
NLDB (1) | 2 |
| 2024 | Analysis of cross-linguality of XL-WSD dataset: A comparative study of Japanese and Dutch
Naranbuuvei Ganbat, Soma Asada, Kanako Komiya |
PACLIC | 3 |
| 2023 | All-Words Word Sense Disambiguation for Historical Japanese
Soma Asada, Kanako Komiya, Masayuki Asahara |
PACLIC | 2 |
| 2023 | Word Segmentation of Hiragana Sentences Using Hiragana BERTabstractAbstract Unlike Western languages, word segmentation is necessary for Japanese sentences because they do not have word boundaries. The performances of existing morphological analyzers for Japanese sentences are very high. However, it is difficult to segment sentences mostly written in Hiragana, which is a Japanese writing system simpler than Kanji, because clues to segment the sentences decrease. In this study, we created a word segmentation model of Hiragana sentences using two types of BERT: unigram and bigram BERT models. We pre-trained the BERT models with Wikipedia and fine-tuned them with the core data of the Balanced Corpus of Contemporary Written Japanese for word segmentation. In addition to the two types of BERT-based word segmentation systems, we developed a word segmentation system for Hiragana sentences using KyTea, a toolkit developed for analyzing text, with a focus on languages requiring word segmentation. We compared them in word segmentation of Hiragana sentences. The experiments revealed that the unigram BERT-based word segmentation system outperformed the bigram BERT-based word segmentation system and the KyTea-based word segmentation system. Jun Izutsu, Kanako Komiya, Hiroyuki Shinnou |
PRICAI (2) | 2 |
| 2023 | Composing Word Embeddings for Compound Words Using Linguistic KnowledgeabstractIn recent years, the use of distributed representations has been a fundamental technology for natural language processing. However, Japanese has multiple compound words, and often we must compare the meanings of a word and a compound word. Moreover, word boundaries in Japanese are unspecific because Japanese does not have delimiters between words, e.g., “ぶどう狩り” (grape picking) is one word according to one dictionary, whereas “ぶどう” and “狩り” are different words according to another dictionary. This study describes an attempt to compose word embeddings of a compound word from its constituent words in Japanese. We used “short unit” and “long unit,” both of which are the units of terms in UniDic—a Japanese dictionary compiled by the National Institute for Japanese Language and Linguistics—for constituent and compound words, respectively. Furthermore, we composed a word embedding of a compound word from the word embeddings of two constituent words using a neural network. The training data for the word embedding of compound words was created using a corpus generated by concatenating the corpora divided by constituent and compound words. We propose using linguistic knowledge for compositing word embedding to demonstrate how it improves the composition performance. We compared cosine similarity between composed and correct word embeddings of compound words to assess models with and without linguistic knowledge. Furthermore, we evaluated our methods by the ranking of synonyms using a thesaurus. We compared several frameworks and algorithms that use three types of linguistic knowledge—semantic patterns, parts of speech patterns, and compositionality score—and then investigated which linguistic knowledge improves the composition performance. The experiments demonstrated that the multitask models with the classification task of the parts of speech patterns and the estimation task of compositionality scores achieved high performances. Kanako Komiya, Shinji Kono, Takumi Seitou, Teruo Hirabayashi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2022 | Reputation Analysis Using Key Phrases and Sentiment Scores Extracted from Reviews
Yipu Huang, Minoru Sasaki, Kanako Komiya |
PACLIC | 3 |
| 2022 | Word Sense Disambiguation of Corpus of Historical Japanese Using Japanese BERT Trained with Contemporary Texts
Kanako Komiya, Nagi Oki, Masayuki Asahara |
PACLIC | 1 |
| 2020 | Composing Word Vectors for Japanese Compound Words Using Bilingual Word Embeddings
Teruo Hirabayashi, Kanako Komiya, Masayuki Asahara, Hiroyuki Shinnou |
PACLIC | 2 |
| 2020 | Generation and Evaluation of Concept Embeddings Via Fine-Tuning Using Automatically Tagged Corpus
Kanako Komiya, Daiki Yaginuma, Masayuki Asahara, Hiroyuki Shinnou |
PACLIC | 1 |
| 2020 | Neural Machine Translation from Historical Japanese to Contemporary Japanese Using Diachronically Domain-Adapted Word Embeddings
Masashi Takaku, Tosho Hirasawa, Mamoru Komachi, Kanako Komiya |
PACLIC | 4 |
| 2019 | Composing Word Vectors for Japanese Compound Words Using Dependency Relations
Kanako Komiya, Takumi Seitou, Minoru Sasaki, Hiroyuki Shinnou |
CICLing (1) | 1 |
| 2018 | All-words Word Sense Disambiguation Using Concept Embeddings
Rui Suzuki, Kanako Komiya, Masayuki Asahara, Minoru Sasaki, Hiroyuki Shinnou |
LREC | 2 |
| 2018 | Domain Adaptation for Sentiment Analysis using Keywords in the Target Domain as the Learning Weight
Jing Bai 0012, Hiroyuki Shinnou, Kanako Komiya |
PACLIC | 3 |
| 2018 | Domain Adaptation Using a Combination of Multiple Embeddings for Sentiment Analysis
Hiroyuki Shinnou, Kanako Komiya |
PACLIC | 3 |
| 2018 | Fine-tuning for Named Entity Recognition Using Part-of-Speech Tagging
Masaya Suzuki, Kanako Komiya, Minoru Sasaki, Hiroyuki Shinnou |
PACLIC | 2 |
| 2018 | Comparison of Methods to Annotate Named Entity CorporaabstractThe authors compared two methods for annotating a corpus for the named entity (NE) recognition task using non-expert annotators: (i) revising the results of an existing NE recognizer and (ii) manually annotating the NEs completely. The annotation time, degree of agreement, and performance were evaluated based on the gold standard. Because there were two annotators for one text for each method, two performances were evaluated: the average performance of both annotators and the performance when at least one annotator is correct. The experiments reveal that semi-automatic annotation is faster, achieves better agreement, and performs better on average. However, they also indicate that sometimes, fully manual annotation should be used for some texts whose document types are substantially different from the training data document types. In addition, the machine learning experiments using semi-automatic and fully manually annotated corpora as training data indicate that the F-measures could be better for some texts when manual instead of semi-automatic annotation was used. Finally, experiments using the annotated corpora for training as additional corpora show that (i) the NE recognition performance does not always correspond to the performance of the NE tag annotation and (ii) the system trained with the manually annotated corpus outperforms the system trained with the semi-automatically annotated corpus with respect to newswires, even though the existing NE recognizer was mainly trained with newswires. Kanako Komiya, Masaya Suzuki, Tomoya Iwakura, Minoru Sasaki, Hiroyuki Shinnou |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2017 | Domain Adaptation for Word Sense Disambiguation Using Word Embeddings
Kanako Komiya, Shota Suzuki, Minoru Sasaki, Hiroyuki Shinnou, Manabu Okumura |
CICLing (1) | 1 |
| 2017 | Japanese all-words WSD system using the Kyoto Text Analysis ToolKit
Hiroyuki Shinnou, Kanako Komiya, Minoru Sasaki, Shinsuke Mori |
PACLIC | 2 |
| 2016 | Supervised Word Sense Disambiguation with Sentences Similarities from Context Word Embeddings
Shoma Yamaki, Hiroyuki Shinnou, Kanako Komiya, Minoru Sasaki |
PACLIC | 3 |
| 2016 | Selecting Training Data for Unsupervised Domain Adaptation in Word Sense Disambiguation
Kanako Komiya, Minoru Sasaki, Hiroyuki Shinnou, Yoshiyuki Kotani, Manabu Okumura |
PRICAI | 1 |
| 2015 | Surrounding Word Sense Model for Japanese All-words Word Sense Disambiguation
Kanako Komiya, Yuto Sasaki, Hajime Morita, Minoru Sasaki, Hiroyuki Shinnou, Yoshiyuki Kotani |
PACLIC | 1 |
| 2015 | Unsupervised Domain Adaptation for Word Sense Disambiguation using Stacked Denoising Autoencoder
Kazuhei Kouno, Hiroyuki Shinnou, Minoru Sasaki, Kanako Komiya |
PACLIC | 4 |
| 2015 | Learning under Covariate Shift for Domain Adaptation for Word Sense Disambiguation
Hiroyuki Shinnou, Minoru Sasaki, Kanako Komiya |
PACLIC | 3 |
| 2015 | Hybrid Method of Semi-supervised Learning and Feature Weighted Learning for Domain Adaptation of Document Classification
Hiroyuki Shinnou, Liying Xiao, Minoru Sasaki, Kanako Komiya |
PACLIC | 4 |
| 2014 | Cross-Lingual Product Recommendation Using Collaborative Filtering with Translation Pairs
Kanako Komiya, Shohei Shibata, Yoshiyuki Kotani |
CICLing (2) | 1 |
| 2012 | Automatic Domain Adaptation for Word Sense Disambiguation Based on Comparison of Multiple Classifiers
Kanako Komiya, Manabu Okumura |
PACLIC | 1 |
| 2012 | The Transliteration from Alphabet Queries to Japanese Product Names
Rieko Tsuji, Yoshinori Nemoto, Wimvipa Luangpiensamut, Yuji Abe, Takeshi Kimura, Kanako Komiya, Koji Fujimoto, Yoshiyuki Kotani |
PACLIC | 6 |
| 2012 | Chinese Morphological Analysis Using Morpheme and Character Features
Kanako Komiya, Haixia Hou, Kazutomo Shibahara, Koji Fujimoto, Yoshiyuki Kotani |
PRICAI | 1 |
| 2012 | Using Tagged and Untagged Corpora to Improve Thai Morphological Analysis with Unknown Word Boundary Detections
Wimvipa Luangpiensamut, Kanako Komiya, Yoshiyuki Kotani |
PRICAI | 2 |
| 2012 | Nested Monte-Carlo Search with simulation reduction
Haruhiko Akiyama, Kanako Komiya, Yoshiyuki Kotani |
Knowl. Based Syst. | 2 |
| 2011 | Automatic Determination of a Domain Adaptation Method for Word Sense Disambiguation Using Decision Tree Learning
Kanako Komiya, Manabu Okumura |
IJCNLP | 1 |
| 2006 | Generating a Set of Rules to Determine Honorific Expression Using Decision Tree Learning
Kanako Komiya, Yasuhiro Tajima, Nobuo Inui, Yoshiyuki Kotani |
CICLing | 1 |