VLDB 2026 Research / reviewers in the wild / expert
Byeongseon Park
dblp:194/3760
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-7419-7999ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CAVIARES: Corpus for Audio-Visual Expressive Voice AgentabstractHigh-quality audio-visual corpora are essential for building voice agents capable of natural human-machine communication, but existing corpora commonly contain a limited amount of data per speaker, making personalized modeling difficult. We present CAVIARES, a new audio-visual corpus comprising 9.5 hours of expressive speech recorded by a single professional Japanese female speaker. CAVIARES consists of two subsets: acted dialogue and expressive reading, providing a diverse range of speaking styles for speech-to-facial motion modeling and multimodal learning tasks. In this paper, we describe the construction process of CAVIARES and the results of corpus analysis. CAVIARES will be released for research purposes only. Jinsheng Chen, Yuki Saito 0001, Naoko Tanji, Hironori Doi, Byeongseon Park, Yuma Shirahata, Kentaro Tachibana, Hiroshi Saruwatari |
ASRU | 6 |
| 2025 | Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
Hien Ohnaka, Yuma Shirahata, Byeongseon Park, Ryuichi Yamamoto |
INTERSPEECH | 3 |
| 2024 | Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
Yuma Shirahata, Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana |
INTERSPEECH | 2 |
| 2022 | A Unified Accent Estimation Method Based on Multi-Task Learning for Japanese Text-to-Speech
Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana |
INTERSPEECH | 1 |
| 2021 | Phrase Break Prediction with Bidirectional Encoder Representations in Japanese Text-to-Speech SynthesisabstractWe propose a novel phrase break prediction method that combines implicit features extracted from a pre-trained large language model, a.k.a BERT, and explicit features extracted from BiLSTM with linguistic features. In conventional BiLSTM based methods, word representations and/or sentence representations are used as independent components. The proposed method takes account of both representations to extract the latent semantics, which cannot be captured by previous methods. The objective evaluation results show that the proposed method obtains an absolute improvement of 3.2 points for the F1 score compared with BiLSTM-based conventional methods using linguistic features. Moreover, the perceptual listening test results verify that a TTS system that applied our proposed method achieved a mean opinion score of 4.39 in prosody naturalness, which is highly competitive with the score of 4.37 for synthesized speech with ground-truth phrase breaks. Kosuke Futamata, Byeongseon Park, Ryuichi Yamamoto, Kentaro Tachibana |
Interspeech | 2 |
| 2019 | Estimating Comic Content from the Book Cover Information Using Fine-Tuned VGG Model for Comic Search
Byeongseon Park, Mitsunori Matsushita |
MMM (2) | 1 |