VLDB 2026 Research / reviewers in the wild / expert
Satoshi Kobashikawa
dblp:09/3769
· DBLP profile ↗
31ranked-venue papers
8as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 22 · 6 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Talking Face Generation for Impression Conversion Considering Speech SemanticsabstractThis study investigates the talking face generation method to convert a speaker’s video to give a target impression, such as “favorable” or “considerate”. Such an impression conversion method needs to consider the input speech semantics because they affect the impression of a speaker’s video along with the facial expression. Conventional emotional talking face generation methods utilize speech information to synchronize the lip and speech of the output video. However, they cannot consider speech semantics because the speech representations contain only phonetic information. To solve this problem, we propose a facial expression conversion model that uses a semantic vector obtained from BERT embeddings of speech recognition results of input speech. We first constructed an audio-visual dataset with impression labels assigned to each utterance. The evaluation results based on the dataset showed that the proposed method could improve the estimation accuracy of the facial expressions of the target video. Saki Mizuno, Nobukatsu Hojo, Kazutoshi Shinoda, Keita Suzuki, Mana Ihori, Tomohiro Tanaka, Naotaka Kawata, Satoshi Kobashikawa, Ryo Masumura |
ICASSP | 9 |
| 2024 | Learning from Multiple Annotator Biased Labels in Multimodal Conversation
Kazutoshi Shinoda, Nobukatsu Hojo, Saki Mizuno, Keita Suzuki, Satoshi Kobashikawa, Ryo Masumura |
INTERSPEECH | 5 |
| 2023 | Next-Speaker Prediction Based on Non-Verbal Information in Multi-Party Video ConversationabstractWe propose a method for next-speaker prediction, a task to predict who speaks in the next turn among multiple current listeners, in multi-party video conversation. Previous studies used non-verbal features, such as head movements and gaze behavior, for next-speaker prediction in face-to-face conversation. However, in video conversation, these non-verbal features are vague and ineffective because they look at the screen displaying other participants. Since non-verbal features include participant characteristics, it is necessary to use training data with rich combinations of participants to robustly predict the next speaker. Previous studies used training data with a limited number of combinations of participants because the data consist only of recorded data. Therefore, the proposed method uses 1) novel non-verbal features for next-speaker prediction in video conversation, specifically facial expressions, hand movements and speech segments, and 2) data augmentation of participant combinations in the training data. We conducted experiments to evaluate the proposed method, and the results using video-conversation data indicate its effectiveness. Saki Mizuno, Nobukatsu Hojo, Satoshi Kobashikawa, Ryo Masumura |
ICASSP | 3 |
| 2023 | Audio-Visual Praise Estimation for Conversational Video based on Synchronization-Guided Multimodal Transformer
Nobukatsu Hojo, Saki Mizuno, Satoshi Kobashikawa, Ryo Masumura, Mana Ihori, Tomohiro Tanaka |
INTERSPEECH | 3 |
| 2022 | Multimodal Negotiation Corpus with Various Subjective Assessments for Social-Psychological Outcome Prediction from Non-Verbal CuesabstractThis study investigates social-psychological negotiation-outcome prediction (SPNOP), a novel task for estimating various subjective evaluation scores of negotiation, such as satisfaction and trust, from negotiation dialogue data. To investigate SPNOP, a corpus with various psychological measurements is beneficial because the interaction process of negotiation relates to many aspects of psychology. However, current negotiation corpora only include information related to objective outcomes or a single aspect of psychology. In addition, most use the “laboratory setting” that uses non-skilled negotiators and over simplified negotiation scenarios. There is a concern that such a gap with actual negotiation will intrinsically affect the behavior and psychology of negotiators in the corpus, which can degrade the performance of models trained from the corpus in real situations. Therefore, we created a negotiation corpus with three features; 1) was assessed with various psychological measurements, 2) used skilled negotiators, and 3) used scenarios of context-rich negotiation. We recorded video and audio of negotiations in Japanese to investigate SPNOP in the context of social signal processing. Experimental results indicate that social-psychological outcomes can be effectively estimated from multimodal information. Nobukatsu Hojo, Satoshi Kobashikawa, Saki Mizuno, Ryo Masumura |
LREC | 2 |
| 2021 | Analysis of Multimodal Features for Speaking Proficiency Scoring in an Interview DialogueabstractThis paper analyzes the effectiveness of different modalities in automated speaking proficiency scoring in an online dialogue task of non-native speakers. Conversational competence of a language learner can be assessed through the use of multimodal behaviors such as speech content, prosody, and visual cues. Although lexical and acoustic features have been widely studied, there has been no study on the usage of visual features, such as facial expressions and eye gaze. To build an automated speaking proficiency scoring system using multi-modal features, we first constructed an online video interview dataset of 210 Japanese English-learners with annotations of their speaking proficiency. We then examined two approaches for incorporating visual features and compared the effectiveness of each modality. Results show the end-to-end approach with deep neural networks achieves a higher correlation with human scoring than one with handcrafted features. Modalities are effective in the order of lexical, acoustic, and visual features. Mao Saeki, Yoichi Matsuyama, Satoshi Kobashikawa, Tetsuji Ogawa, Tetsunori Kobayashi |
SLT | 3 |
| 2020 | Improving Speaker-Attribute Estimation by Voting Based on Speaker Cluster InformationabstractThis paper proposes a general post-processing method for improving speaker-attribute estimation. Estimating speaker-specific attributes such as age and gender is an important task with a wide range of applications. While the recent proposed deep neural network-based end-to-end approach achieves high performance, the model tends to over-fit to specific speakers when the amount of training data is limited or imbalanced. To solve this over-fitting problem, we propose a general framework for correcting unreliable results. The proposed algorithm first clusters the target utterances into speaker clusters by speaker similarity based on i-vectors. Then, for each of the speaker cluster, the speaker-attribute class of the cluster is determined by voting on the utterances assigned to the cluster. By then replacing the result of each utterance with the clusters' speaker-attribute class, we can correct the result of unreliable utterances. We used two tasks to evaluate the proposed algorithm including age estimation using the NIST-SRE10 and age-gender classification using an in-house read speech corpus, yielding significant improvements in mean absolute and classification errors. Naohiro Tawara, Hosana Kamiyama, Satoshi Kobashikawa, Atsunori Ogawa |
ICASSP | 3 |
| 2020 | Customer Satisfaction Estimation in Contact Center Calls Based on a Hierarchical Multi-Task ModelabstractThis article presents a novel customer satisfaction (CS) estimation method that outputs both turn-level and call-level estimations simultaneously. Our key idea is to directly apply turn-level estimation results to call-level estimation and optimize them jointly; previous works treat both as being independent. Our proposal applies long short-term memory recurrent neural networks (LSTM-RNNs) to turn-level and call-level CS estimation to capture long-range sequential context in contact center calls. In addition, both networks are hierarchically stacked so as to use turn-level estimation results for call-level estimation directly. In order to learn the relationship between the two tasks, we also introduce joint optimization training to the stacked model. Several analyses of turn-level and call-level CS are provided on acted and real calls to support the proposed method. Experiments show that the proposed framework outperforms the conventional methods in both turn-level and call-level estimations. Atsushi Ando, Ryo Masumura, Hosana Kamiyama, Satoshi Kobashikawa, Yushi Aono, Tomoki Toda |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Does the Lombard Effect Improve Emotional Communication in Noise? - Analysis of Emotional Speech Acted in NoiseabstractSpeakers usually adjust their way of talking in noisy environments involuntarily for effective communication. This adaptation is known as the Lombard effect. Although speech accompanying the Lombard effect can improve the intelligibility of a speaker's voice, the changes in acoustic features (e.g. fundamental frequency, speech intensity, and spectral tilt) caused by the Lombard effect may also affect the listener's judgment of emotional content. To the best of our knowledge, there is no published study on the influence of the Lombard effect in emotional speech. Therefore, we recorded parallel emotional speech waveforms uttered by 12 speakers under both quiet and noisy conditions in a professional recording studio in order to explore how the Lombard effect interacts with emotional speech. By analyzing confusion matrices and acoustic features, we aim to answer the following questions: 1) Can speakers express their emotions correctly even under adverse conditions? 2) Can listeners recognize the emotion contained in speech signals even under noise? 3) How does emotional speech uttered in noise differ from emotional speech uttered in quiet conditions in terms of acoustic characteristic? Yi Zhao 0006, Atsushi Ando, Shinji Takaki, Junichi Yamagishi, Satoshi Kobashikawa |
INTERSPEECH | 5 |
| 2019 | Speech Emotion Recognition Based on Multi-Label Emotion Existence Model
Atsushi Ando, Ryo Masumura, Hosana Kamiyama, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 4 |
| 2019 | Improving Conversation-Context Language Models with Multiple Spoken Language Understanding Models
Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Hosana Kamiyama, Takanobu Oba, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 6 |
| 2018 | Soft-Target Training with Ambiguous Emotional Utterances for DNN-Based Speech Emotion ClassificationabstractThis paper presents a novel emotion classification method for natural speech. One of the problems in the state-of-the-art method based on Deep Neural Network (DNN) is the paucity of the training data compared to model complexity. To solve this problem, this paper utilizes the ambiguous emotional utterances, utterances that have no dominant target emotion label. While previous work ignored ambiguous emotional utterances for training, the proposed method leverages all annotated labels via soft-target training. In addition, this paper modifies the soft-target training in order to effectively handle both clear and ambiguous emotional utterances. Experiments show that the proposed method yields performance improvements in terms of both weighted and unweighted accuracies. Atsushi Ando, Satoshi Kobashikawa, Hosana Kamiyama, Ryo Masumura, Yusuke Ijima, Yushi Aono |
ICASSP | 2 |
| 2018 | Automatic Question Detection from Acoustic and Phonetic Features Using Feature-wise Pre-training
Atsushi Ando, Reine Asakawa, Ryo Masumura, Hosana Kamiyama, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 5 |
| 2017 | Hierarchical LSTMs with Joint Learning for Estimating Customer Satisfaction from Contact Center Calls
Atsushi Ando, Ryo Masumura, Hosana Kamiyama, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 4 |
| 2017 | Interaction and Transition Model for Speech Emotion Recognition in Dialogue
Ruo Zhang, Atsushi Ando, Satoshi Kobashikawa, Yushi Aono |
INTERSPEECH | 3 |
| 2014 | Efficient data selection for speech recognition based on prior confidence estimation using speech and monophone models
Satoshi Kobashikawa, Taichi Asami, Yoshikazu Yamaguchi, Hirokazu Masataki, Satoshi Takahashi |
Comput. Speech Lang. | 1 |
| 2013 | Unsupervised confidence calibration using examples of recognized words and their contexts
Taichi Asami, Satoshi Kobashikawa, Hirokazu Masataki, Osamu Yoshioka, Satoshi Takahashi |
INTERSPEECH | 2 |
| 2013 | Fast unsupervised adaptation based on efficient statistics accumulation using frame independent confidence within monophone states
Satoshi Kobashikawa, Atsunori Ogawa, Taichi Asami, Yoshikazu Yamaguchi, Hirokazu Masataki, Satoshi Takahashi |
Comput. Speech Lang. | 1 |
| 2012 | Speech Data Clustering Based on Phoneme Error Trend for Unsupervised Acoustic Model Adaptation
Taichi Asami, Satoshi Kobashikawa, Hirokazu Masataki, Osamu Yoshioka, Satoshi Takahashi |
INTERSPEECH | 2 |
| 2012 | Efficient Beam Width Control to Suppress Excessive Speech Recognition Computation Time Based on Prior Score Range Normalization
Satoshi Kobashikawa, Takaaki Hori, Yoshikazu Yamaguchi, Taichi Asami, Hirokazu Masataki, Satoshi Takahashi |
INTERSPEECH | 1 |
| 2012 | Efficient prior and incremental beam width control to suppress excessive speech recognition time based on score range estimationabstractThis paper proposes a technique that efficiently controls the beam width to yield practical computation times when auto-transcribing massive volumes of speeches. We focus on the fact that a lot of time is wasted by recognizing poor quality speeches that will yield, with inordinate slowness, erroneous transcriptions and provide no useful results. To stabilize the time regardless of quality, our proposal controls the beam width based on prolonged score spread against the target speech; it formulates the score range within the width and maximizes computation efficiency by regulating the range relevant to the hypotheses' survival rate. The proposed technique can control the width rapidly by using just monophones prior to decoding. It also restricts the width in decoding by using the processing speed and remaining data time to better handle stubborn speeches. Experiments with several SNRs and actual call-center speeches confirm a reduction in computation time while matching the accuracy of existing techniques. Satoshi Kobashikawa, Takaaki Hori, Yoshikazu Yamaguchi, Taichi Asami, Hirokazu Masataki, Satoshi Takahashi |
SLT | 1 |
| 2011 | Extracting call-reason segments from contact center dialogs by using automatically acquired boundary expressionsabstractTo improve the performance of call-reason analysis at contact centers, we introduce a novel method to extract call-reason segments from dialogs. It is based on the following two characteristics of contact center conversations; 1) customers state their requests at the beginning of the calls, 2) agents tend to use typical phrases at the end of the call-reason segments. Our proposal acquires these typical phrases from stored speech data automatically and extracts the call-reason segment precisely by detecting the typical phrases. Experiments show that it significantly improves the performance of call-reason information retrieval since it allows the search scope to be limited to the call-reason segments of calls. Takaaki Fukutomi, Satoshi Kobashikawa, Taichi Asami, Tsubasa Shinozaki, Hirokazu Masataki, Satoshi Takahashi |
ICASSP | 2 |
| 2011 | Spoken Document Confidence Estimation Using Contextual Coherence
Taichi Asami, Narichika Nomoto, Satoshi Kobashikawa, Yoshikazu Yamaguchi, Hirokazu Masataki, Satoshi Takahashi |
INTERSPEECH | 3 |
| 2011 | Morpheme Conversion for Connecting Speech Recognizer and Language Analyzers in Unsegmented Languages
Kenji Imamura, Tomoko Izumi, Kugatsu Sadamitsu, Kuniko Saito, Satoshi Kobashikawa, Hirokazu Masataki |
INTERSPEECH | 5 |
| 2010 | Efficient data selection for speech recognition based on prior confidence estimation using speech and context independent models
Satoshi Kobashikawa, Taichi Asami, Yoshikazu Yamaguchi, Hirokazu Masataki, Satoshi Takahashi |
INTERSPEECH | 1 |
| 2010 | Improving hmm-based extractive summarization for multi-domain contact center dialoguesabstractThis paper reports the improvements we made to our previously proposed hidden Markov model (HMM) based summarization method for multi-domain contact center dialogues. Since the method relied on Viterbi decoding for selecting utterances to include in a summary, it had the inability to control compression rates. We enhance our method by using the forward-backward algorithm together with integer linear programming (ILP) to enable the control of compression rates, realizing summaries that contain as many domain-related utterances and as many important words as possible within a predefined character length. Using call transcripts as input, we verify the effectiveness of our enhancement. Ryuichiro Higashinaka, Yasuhiro Minami, Hitoshi Nishikawa, Kohji Dohsaka, Toyomi Meguro, Satoshi Kobashikawa, Hirokazu Masataki, Osamu Yoshioka, Satoshi Takahashi, Gen-ichiro Kikui |
SLT | 6 |
| 2010 | Efficient data selection for spoken document retrieval based on prior confidence estimation using speech and context independent modelsabstractThis paper proposes an efficient speech sample selection technique that can identify those samples that will be well recognized. Conventional confidence measures can identify well-recognized speech samples, but they require speech recognition to estimate confidence scores. Speech samples with low confidence should not undergo recognition since they yield speech documents that will eventually be rejected. The proposed technique can select the samples that will justify the application of speech recognition. It is based on rapid prior confidence estimation by using speech and context independent models to calculate acoustic likelihood values on a frame-by-frame basis. Tests show that the proposed confidence estimation technique is over 50 times faster than the conventional posterior confidence measure while maintaining equivalent data selection performance for speech recognition and spoken document retrieval. Satoshi Kobashikawa, Taichi Asami, Yoshikazu Yamaguchi, Hirokazu Masataki, Satoshi Takahashi |
SLT | 1 |
| 2009 | Rapid unsupervised adaptation using frame independent output probabilities of gender and context independent phoneme models
Satoshi Kobashikawa, Atsunori Ogawa, Yoshikazu Yamaguchi, Satoshi Takahashi |
INTERSPEECH | 1 |
| 2005 | Rapid response and robust speech recognition by preliminary model adaptation for additive and convolutional noise
Satoshi Kobashikawa, Satoshi Takahashi, Yoshikazu Yamaguchi, Atsunori Ogawa |
INTERSPEECH | 1 |
| 2004 | Robust speech recognition based on HMM composition and modified wiener filter
Sumitaka Sakauchi, Yoshikazu Yamaguchi, Satoshi Takahashi, Satoshi Kobashikawa |
INTERSPEECH | 4 |
| 2002 | Acoustic modeling of sentence stress using differential features between syllables for English rhythm learning system development
Nobuaki Minematsu, Satoshi Kobashikawa, Keikichi Hirose, Donna Erickson |
INTERSPEECH | 2 |