Yun-Shao Lin

dblp:212/6294 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-1494-1413ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism
Hsing-Hang Chou, Yun-Shao Lin, Ching-Chin Sung, Yu Tsao 0001, Chi-Chun Lee
INTERSPEECH2
2024 A Cluster-based Personalized Federated Learning Strategy for End-to-End ASR of Dementia Patients
Wei-Tung Hsu, Chin-Po Chen, Yun-Shao Lin, Chi-Chun Lee
INTERSPEECH3
2023 Speaking State Decoder with Transition Detection for Next Speaker Prediction
Shao-Hao Lu, Yun-Shao Lin, Chi-Chun Lee
INTERSPEECH2
2023 An Interaction-process-guided Framework for Small-group Performance Prediction
abstract
A small group is a fundamental interaction unit for achieving a shared goal. Group performance can be automatically predicted using computational methods to analyze members’ verbal behavior in task-oriented interactions, as has been proven in several recent works. Most of the prior works focus on lower-level verbal behaviors, such as acoustics and turn-taking patterns, using either hand-crafted features or even advanced end-to-end methods. However, higher-level group-based communicative functions used between group members during conversations have not yet been considered. In this work, we propose a two-stage training framework that effectively integrates the communication function, as defined using Bales’s interaction process analysis (IPA) coding system, with the embedding learned from the low-level features in order to improve the group performance prediction. Our result shows a significant improvement compared to the state-of-the-art methods (4.241 MSE and 0.341 Pearson’s correlation on NTUBA-task1 and 3.794 MSE and 0.291 Pearson’s correlation on NTUBA-task2) on the National Taiwan University Business Administration (NTUBA) small-group interaction database. Furthermore, based on the design of IPA, our computational framework can provide a time-grained analysis of the group communication process and interpret the beneficial communicative behaviors for achieving better group performance.
Yun-Shao Lin, Yi-Ching Liu, Chi-Chun Lee
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Emotion-Shift Aware CRF for Decoding Emotion Sequence in Conversation
Yun-Shao Lin, Chi-Chun Lee
INTERSPEECH2
2021 Through the Words of Viewers: Using Comment-Content Entangled Network for Humor Impression Recognition
abstract
Research into understanding humor has been investigated over centuries. It has recently attracted various technical effort in computing humor automatically from data, especially for humor in speech. Comprehension on the same speech and the ability to realize a humor event vary depending on each individual audience's background and experience. Most previous works on automatic humor detection or impression recognition mainly model the produced textual content only without considering audience responses. We collect a corpus of TED Talks including audience comments for each of the presented TED speech. We propose a novel network architecture that considers the natural entanglement between speech transcripts and user's online feedbacks as an integrative graph structure, where the content speech and online feedbacks are nodes where the edges are connected though their common words. Our model achieves 61.2% of accuracy in a three-class classification on humor impression recognition on TED talks; our experiments further demonstrate viewers comments are essential in improving the recognition tasks, and a joint content-comment modeling achieves the best recognition.
Huan-Yu Chen, Yun-Shao Lin, Chi-Chun Lee
SLT2
2020 Predicting Performance Outcome with a Conversational Graph Convolutional Network for Small Group Interactions
abstract
Studying behaviors of members during small group interaction provides objective insights in improving the efficiency of the decision making process in our daily working life. By introducing the use of the graph structure in modeling the natural inter-member conversational ties during such an interaction, we aim to advance the state-of-art computational approach in predicting group performance scores. Specifically, we proposed a Conversational Graph Convolutional Network (CGCN) that utilizes conversation dynamic as the graph to aggregate group member's speech and lexical behaviors in predicting the group performance. Our result shows that Speech CGCN achieves the state-of-the-art performance at MSE 3.896 (0.323 Pearson correlation) outperform the current best method in ELEA dataset. Our model additionally reveals that an imbalance conversational graph structure is positively correlated to group performances.
Yun-Shao Lin, Chi-Chun Lee
ICASSP1
2020 A Dialogical Emotion Decoder for Speech Motion Recognition in Spoken Dialog
abstract
Developing a robust emotion speech recognition (SER) system for human dialog is important in advancing conversational agent design. In this paper, we proposed a novel inference algorithm, a dialogical emotion decoding (DED) algorithm, that treats a dialog as a sequence and consecutively decode the emotion states of each utterance over time with a given recognition engine. This decoder is trained by incorporating intra- and inter-speakers emotion influences within a conversation. Our approach achieves a 70.1% in four class emotion on the IEMOCAP database, which is 3% over the state-of-art model. The evaluation is further conducted on a multi-party interaction database, the MELD, which shows a similar effect. Our proposed DED is in essence a conversational emotion rescoring decoder that can also be flexibly combined with different SER engines.
Sung-Lin Yeh, Yun-Shao Lin, Chi-Chun Lee
ICASSP2
2020 Improving Speech Emotion Recognition Using Graph Attentive Bi-Directional Gated Recurrent Unit Network
Bo-Hao Su, Chun-Min Chang, Yun-Shao Lin, Chi-Chun Lee
INTERSPEECH3
2020 Speech Representation Learning for Emotion Recognition Using End-to-End ASR with Factorized Adaptation
Sung-Lin Yeh, Yun-Shao Lin, Chi-Chun Lee
INTERSPEECH2
2019 An Interaction-aware Attention Network for Speech Emotion Recognition in Spoken Dialogs
abstract
Obtaining robust speech emotion recognition (SER) in scenarios of spoken interactions is critical to the developments of next generation human-machine interface. Previous research has largely focused on performing SER by modeling each utterance of the dialog in isolation without considering the transactional and dependent nature of the human-human conversation. In this work, we propose an interaction-aware attention network (IAAN) that incorporate contextual information in the learned vocal representation through a novel attention mechanism. Our proposed method achieves 66.3% accuracy (7.9% over baseline methods) in four class emotion recognition and is also the current state-of-art recognition rates obtained on the benchmark database.
Sung-Lin Yeh, Yun-Shao Lin, Chi-Chun Lee
ICASSP2
2019 Enforcing Semantic Consistency for Cross Corpus Valence Regression from Speech Using Adversarial Discrepancy Learning
Gao-Yi Chao, Yun-Shao Lin, Chun-Min Chang, Chi-Chun Lee
INTERSPEECH2
2019 Predicting Group Performances Using a Personality Composite-Network Architecture During Collaborative Task
Shun-Chang Zhong, Yun-Shao Lin, Chun-Min Chang, Yi-Ching Liu, Chi-Chun Lee
INTERSPEECH2
2018 A Genre-Affect Relationship Network with Task-Specific Uncertainty Weighting foR Recognizing Induced Emotion in Music
abstract
Emotion is a core fundamental attribute of humans. Using music to induce emotional responses from subjects to better facilitate human behavior shaping have been effective across domains of health, education, and retail. Computationally model the musically-induced emotion provides necessary content-based analytics for large-scale and wide-applicability of such human-centered applications. In this work, we propose a relationship neural network architecture to learn to regress the induced emotion attributes with an auxiliary task of genre classification. Our proposed Genre-Affect Relationship Network with homoscedastic uncertainty weighting embeds the relationship between affect and genre as tensor normal prior within task-specific layers; the architecture is optimized further by incorporating task-specific uncertainty. The proposed architecture achieves a state-of-art 0.564 average Pearson correlation computed over nine induced emotion ratings in the Emotify database. Furthermore, we provide an analysis to understand the relationship between the induced emotions of these musical pieces and their associated genres.
Wei-Hao Chang, Jeng-Lin Li, Yun-Shao Lin, Chi-Chun Lee
ICME3
2018 Using Interlocutor-Modulated Attention BLSTM to Predict Personality Traits in Small Group Interaction
abstract
Small group interaction occurs often in workplace and education settings. Its dynamic progression is an essential factor in dictating the final group performance outcomes. The personality of each individual within the group is reflected in his/her interpersonal behaviors with other members of the group as they engage in these task-oriented interactions. In this work, we propose an interlocutor-modulated attention BSLTM (IM-aBLSTM) architecture that models an individual's vocal behaviors during small group interactions in order to automatically infer his/her personality traits. The interlocutor-modulated attention mechanism jointly optimize the relevant interpersonal vocal behaviors of other members of group during interactions. In specifics, we evaluate our proposed IM-aBLSTM in one of the largest small group interaction database, the ELEA corpus. Our framework achieves a promising unweighted recall accuracy of 87.9% in ten different binary personality trait prediction tasks, which outperforms the best results previously reported on the same database by 10.4% absolute. Finally, by analyzing the interpersonal vocal behaviors in the region of high attention weights, we observe several distinct intra- and inter-personal vocal behavior patterns that vary as a function of personality traits.
Yun-Shao Lin, Chi-Chun Lee
ICMI1
2018 An Interlocutor-Modulated Attentional LSTM for Differentiating between Subgroups of Autism Spectrum Disorder
Yun-Shao Lin, Susan Shur-Fen Gau, Chi-Chun Lee
INTERSPEECH1
2017 Deriving Dyad-Level Interaction Representation Using Interlocutors Structural and Expressive Multimodal Behavior Features
Yun-Shao Lin, Chi-Chun Lee
INTERSPEECH1