Kazuma Takaoka

dblp:97/7160 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 56% Geometric modeling and processing · 44%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Geometric modeling and processing › shape analysis
morphological analysis
0.012004
Morphological analysis of the corpus of spontaneous Japanese · IEEE Trans. Speech Audio Process. 2004
Audio and music processing
speech corpus
0.012004
Morphological analysis of the corpus of spontaneous Japanese · IEEE Trans. Speech Audio Process. 2004

Methods — techniques the papers use, named apart from their topics

semi-automatic analysis · 0.0
YearPublicationVenuePosition
2018 Sudachi: a Japanese Tokenizer for Business
Kazuma Takaoka, Sorami Hisamoto, Noriko Kawahara, Miho Sakamoto, Yoshitaka Uchida, Yuji Matsumoto 0001
LREC1
2004 Morphological analysis of the corpus of spontaneous Japanese
abstract
This paper describes two methods for detecting word segments and their morphological information in a Japanese spontaneous speech corpus, and describes how to tag a large spontaneous speech corpus accurately by using the two methods. The first method is used to detect any type of word segments. The second method is used when there are several definitions for word segments and their POS categories, and when one type of word segments includes another type of word segments. In this paper, we show that by using semi-automatic analysis, we achieve a precision of better than 99% for detecting and tagging short-unit words and 97% for long-unit words; the two types of words that comprise the corpus. We also show that better accuracy is achieved by using both methods than by using only the first.
Kiyotaka Uchimoto, Kazuma Takaoka, Chikashi Nobata, Atsushi Yamada, Satoshi Sekine, Hitoshi Isahara
IEEE Trans. Speech Audio Process.2