Jen-Chun Lin

dblp:70/10152 · DBLP profile ↗
← Back
31ranked-venue papers
11as first author
10since 2021 · last 2026
0000-0002-9237-4119ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 10 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author
YearPublicationVenuePosition
2026 Tai-Chi: Text-to-motion generation with locality-aware bipartite body-part motion prior
Jian-Kai Zhu, Wen-Li Wei, Jen-Chun Lin
Comput. Graph.3
2025 Birds of a Feather: Learning to Retrieve Dance Poses From Music Via Ground-Truth Annotation Lifting
abstract
Learning to retrieve dance poses from music, a cross-modal retrieval task, has gained prominence in assisting choreographers in creating dances that harmonize with music. The recent predominant approach is to map music into 3D pose and shape space, and then match it with dance poses [1]. However, we found that the mainstream choreography dataset lacks discriminative power in terms of 3D pose and shape annotations across different dance genres, hindering the model’s ability to learn effective mappings, which in turn reduces retrieval performance. To address the issue, we propose LiftNet, a deep-net model that uses dance genres as guidance to lift 3D pose and shape annotations, making them more discriminative and easier for the downstream retrieval model to learn. Experimental results demonstrate that using the lifted annotations from our LiftNet as new learning targets substantially enhances the performance of all existing cross-modal music-to-dance pose retrieval models.
Bo-Wei Tseng, Wen-Li Wei, Jen-Chun Lin
ICASSP3
2024 Music-to-Dance Poses: Learning to Retrieve Dance Poses from Music
abstract
Choreography is an artful blend of technique and creativity, requiring the meticulous design of movement sequences in harmony with music. To support choreographers in this intricate task, this work proposes a "music-to-dance pose retrieval" system that uses music snippets to retrieve dance poses, predicts 3D human poses and shapes, and then matches them within the 3D pose and shape space. Central to our method is the EDSA adapter, a Self-Attention adapter that utilizes an Encoder-Decoder transformation, allowing a large-scale pre-trained music model to be fine-tuned effectively and efficiently for learning projection from music snippets to 3D human poses and shapes. Experimental results demonstrate that our EDSA adapter outperforms existing techniques for fine-tuning a large-scale pre-trained model in cross-modal music-to-dance pose retrieval task.
Bo-Wei Tseng, Kenneth Yang, Yu-Hua Hu, Wen-Li Wei, Jen-Chun Lin
ICASSP5
2024 Multi-Candidate Motion Modeling for 3D Human Pose and Shape Estimation from Monocular Video
abstract
Estimating 3D human pose and shape from monocular video is an ill-posed problem due to depth ambiguity. Yet, most existing methods overlook the potential multiple motion hypotheses arising from this ambiguity. To tackle this, we propose a multi-candidate motion pose and shape network (MMPS-Net), which is designed to generate temporal representations of multiple plausible motion candidates and yield their adaptive fusion for 3D human pose and shape estimation. Specifically, we first propose a multi-candidate motion continuity attention (MMoCA) module to generate multiple kinematically compliant motion candidates. Second, we introduce a multi-candidate cross-attention (MCA) module to enable information passing among candidates to strengthen their relevance. Third, we develop a multi-candidate hierarchical attentive feature integration (MHAFI) module to refine the target frame’s feature representation by capturing temporal correlations within each motion candidate and adaptively integrating all candidates. By coupling these designs, MMPSNet surpasses video-based methods on the 3DPW, MPI-INF-3DHP, and Human3.6M benchmarks.
Wen-Li Wei, Jen-Chun Lin
ICME2
2024 Progressive Hypothesis Transformer for 3D Human Mesh Recovery
abstract
Recent advancements in Transformer-based human mesh reconstruction (HMR) are commendable. However, these models often lift 2D images directly to 3D vertices without explicit intermediate guidance. In addition, the global attention mechanism tends to spread attention across larger body areas and even unrelated background regions during human mesh estimation, rather than focusing on critical local regions such as human body joints. This tendency leads to inaccurate and unrealistic results for complex activities. To address these challenges, we introduce the Progressive Hypothesis Transformer, which employs 2D and 3D pose predictions to progressively guide our model. Moreover, we propose a mechanism that generates multiple plausible hypotheses for both 2D and 3D poses to mitigate potential inaccuracies arising from intermediate pose estimations. Our model also incorporates inter-intra attention to capture correlations between joints and hypotheses. Experimental results demonstrate that our method surpasses existing image-based approaches on Human3.6M [13] and 3DPW [36] with fewer parameters and relatively lower computational costs.
Huang-Ru Liao, Jen-Chun Lin, Chun-Yi Lee
WACV2
2024 Bridging Actions: Generate 3D Poses and Shapes In-Between Photos
abstract
Generating realistic 3D human motion has been a fundamental goal of the game/animation industry. This work presents a novel transition generation technique that can bridge the actions of people in the foreground by generating 3D poses and shapes in-between photos, allowing 3D animators/novice users to easily create/edit 3D motions. To achieve this, we propose an adaptive motion network (ADAM-Net) that effectively learns human motion from masked action sequences to generate kinematically compliant 3D poses and shapes in-between given temporally-sparse photos. Three core learning designs underpin ADAM-Net. First, we introduce a random masking process that randomly masks images from an action sequence and fills masked regions in latent space by interpolation of unmasked images to simulate various transitions under given temporally-sparse photos. Second, we propose a long-range adaptive motion (L-ADAM) attention module that leverages visual cues observed from human motion to adaptively recalibrate the range that needs attention in a sequence, along with a multi-head cross-attention. Third, we develop a short-range adaptive motion (S-ADAM) attention module that weightedly selects and integrates adjacent feature representations at different levels to strengthen temporal correlation. By coupling these designs, the results demonstrate that ADAM-Net excels not only in generating 3D poses and shapes in-between photos, but also in classic 3D human pose and shape estimation.
Wen-Li Wei, Jen-Chun Lin
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Global-Local Awareness Network for Image Super-Resolution
abstract
Deep-net models based on self-attention, such as Swin Transformer, have achieved great success for single image super-resolution (SISR). While self-attention excels at modeling global information, it is less effective at capturing high frequencies (e.g., edges etc.) that deliver local information primarily, which is crucial for SISR. To tackle this, we propose a global-local awareness network (GLA-Net) to effectively capture global and local information to learn comprehensive features with low- and high-frequency information. First, we design a GLA layer that combines a high-frequency-oriented Inception module with a low-frequency-oriented Swin Transformer module to simultaneously process local and global information. Second, we introduce dense connections in-between GLA blocks to strengthen feature propagation and alleviate the vanishing-gradient problem, where each GLA block is composed of several GLA layers. By coupling these core designs, GLA-Net achieves SOTA performance on SISR.
Pin-Chi Pan, Tzu-Hao Hsu, Wen-Li Wei, Jen-Chun Lin
ICIP4
2022 Capturing Humans in Motion: Temporal-Attentive 3D Human Pose and Shape Estimation from Monocular Video
abstract
Learning to capture human motion is essential to 3D human pose and shape estimation from monocular video. However, the existing methods mainly rely on recurrent or convolutional operation to model such temporal information, which limits the ability to capture non-local context relations of human motion. To address this problem, we propose a motion pose and shape network (MPS-Net) to effectively capture humans in motion to estimate accurate and temporally coherent 3D human pose and shape from a video. Specifically, we first propose a motion continuity attention (MoCA) module that leverages visual cues observed from human motion to adaptively recalibrate the range that needs attention in the sequence to better capture the motion continuity dependencies. Then, we develop a hierarchical attentive feature integration (HAFI) module to effectively combine adjacent past and future feature represen-tations to strengthen temporal correlation and refine the feature representation of the current frame. By coupling the MoCA and HAFI modules, the proposed MPS-Net excels in estimating 3D human pose and shape in the video. Though conceptually simple, our MPS-Net not only outperforms the state-of-the-art methods on the 3DPW, MPI-INF-3DHP, and Human3.6M benchmark datasets, but also uses fewer network parameters. The video demos can be found at https://mps-net.github.io/MPS-Net/.
Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hong-Yuan Mark Liao
CVPR2
2021 Positions, Channels, and Layers: Fully Generalized Non-Local Network for Singer Identification
abstract
Recently, a non-local (NL) operation has been designed as the central building block for deep-net models to capture long-range dependencies (Wang et al. 2018). Despite its excellent performance, it does not consider the interaction between positions across channels and layers, which is crucial in fine-grained classification tasks. To address the limitation, we target at singer identification (SID) task and present a fully generalized non-local (FGNL) module to help identify fine-grained vocals. Specifically, we first propose a FGNL operation, which extends the NL operation to explore the correlations between positions across channels and layers. Secondly, we further apply a depth-wise convolution with Gaussian kernel in the FGNL operation to smooth feature maps for better generalization. More, we modify the squeeze-and-excitation (SE) scheme into the FGNL module to adaptively emphasize correlated feature channels to help uncover relevant feature responses and eventually the target singer. Evaluating results on the benchmark artist20 dataset shows that the FGNL module significantly improves the accuracy of the deep-net models in SID. Codes are available at https://github.com/ian-k-1217/Fully-Generalized-Non-Local-Network.
I-Yuan Kuo, Wen-Li Wei, Jen-Chun Lin
AAAI3
2021 Learning to Visualize Music Through Shot Sequence for Automatic Concert Video Mashup
abstract
An experienced director usually switches among different types of shots to make visual storytelling more touching. When filming a musical performance, appropriate switching shots can produce some special effects, such as enhancing the expression of emotion or heating up the atmosphere. However, while the visual storytelling technique is often used in making professional recordings of a live concert, amateur recordings of audiences often lack such storytelling concepts and skills when filming the same event. Thus a versatile system that can perform video mashup to create a refined high-quality video from such amateur clips is desirable. To this end, we aim at translating the music into an attractive shot (type) sequence by learning the relation between music and visual storytelling of shots. The resulting shot sequence can then be used to better portray the visual storytelling of a song and guide the concert video mashup process. To achieve the task, we first introduces a novel probabilistic-based fusion approach, named as multi-resolution fused recurrent neural networks (MF-RNNs) with film-language, which integrates multi-resolution fused RNNs and a film-language model for boosting the translation performance. We then distill the knowledge in MF-RNNs with film-language into a lightweight RNN, which is more efficient and easier to deploy. The results from objective and subjective experiments demonstrate that both MF-RNNs with film-language and lightweight RNN can generate attractive shot sequences for music, thereby enhancing the viewing and listening experience.
Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hsiao-Rong Tyan, Hsin-Min Wang, Hong-Yuan Mark Liao
IEEE Trans. Multim.2
2020 Learning From Music to Visual Storytelling of Shots: A Deep Interactive Learning Mechanism
abstract
Learning from music to visual storytelling of shots is an interesting and emerging task. It produces a coherent visual story in the form of a shot type sequence, which not only expands the storytelling potential for a song but also facilitates automatic concert video mashup process and storyboard generation. In this study, we present a deep interactive learning (DIL) mechanism for building a compact yet accurate sequence-to-sequence model to accomplish the task. Different from the one-way transfer between a pre-trained teacher network (or ensemble network) and a student network in knowledge distillation (KD), the proposed method enables collaborative learning between an ensemble teacher network and a student network. Namely, the student network also teaches. Specifically, our method first learns a teacher network that is composed of several assistant networks to generate a shot type sequence and produce the soft target (shot types) distribution accordingly through KD. It then constructs the student network that learns from both the ground truth label (hard target) and the soft target distribution to alleviate the difficulty of optimization and improve generalization capability. As the student network gradually advances, it turns to feed back knowledge to the assistant networks, thereby improving the teacher network in each iteration. Owing to such interactive designs, the DIL mechanism bridges the gap between the teacher and student networks and produces more superior capability for both networks. Objective and subjective experimental results demonstrate that both the teacher and student networks can generate more attractive shot sequences from music, thereby enhancing the viewing and listening experience.
Jen-Chun Lin, Wen-Li Wei, Yen-Yu Lin, Tyng-Luh Liu, Hong-Yuan Mark Liao
ACM Multimedia1
2019 What Makes You Look Like You: Learning an Inherent Feature Representation for Person Re-Identification
abstract
In this work, we address person re-identification (ReID) by learning an inherent feature representation (inherent code) that is unique to each individual. This task is difficult because the appearance of a person may vary dramatically due to diverse factors, such as illuminations, viewpoints, and human pose changes. To tackle this issue, we propose new learning objectives to learn the inherent code for each person based on deep learning. Specifically, the proposed deep-net model is trained by jointly optimizing the multiple objectives that pulls the instances of the same person closer while pushing the instances belonging to different persons far from each other. Owing to such complementary designs, the deep-net model yields a robust code for each individual and hence better solve person ReID. Promising experimental results demonstrate the robustness and effectiveness of our proposed method.
Wen-Li Wei, Jen-Chun Lin, Yen-Yu Lin, Hong-Yuan Mark Liao
AVSS2
2019 Tell Me Where It is Still Blurry: Adversarial Blurred Region Mining and Refining
abstract
Mobile devices such as smart phones are ubiquitously being used to take photos and videos, thus increasing the importance of image deblurring. This study introduces a novel deep learning approach that can automatically and progressively achieve the task via adversarial blurred region mining and refining (adversarial BRMR). Starting with a collaborative mechanism of two coupled conditional generative adversarial networks (CGANs), our method first learns the image-scale CGAN, denoted as iGAN, to globally generate a deblurred image and locally uncover its still blurred regions through an adversarial mining process. Then, we construct the patch-scale CGAN, denoted as pGAN, to further improve sharpness of the most blurred region in each iteration. Owing to such complementary designs, the adversarial BRMR indeed functions as a bridge between iGAN and pGAN, and yields the performance synergy in better solving blind image deblurring. The overall formulation is self-explanatory and effective to globally and locally restore an underlying sharp image. Experimental results on benchmark datasets demonstrate that the proposed method outperforms the current state-of-the-art technique for blind image deblurring both quantitatively and qualitatively.
Jen-Chun Lin, Wen-Li Wei, Tyng-Luh Liu, C.-C. Jay Kuo, Hong-Yuan Mark Liao
ACM Multimedia1
2018 Seethevoice: Learning from Music to Visual Storytelling of Shots
abstract
Types of shots in the language of film are considered the key elements used by a director for visual storytelling. In filming a musical performance, manipulating shots could stimulate desired effects such as manifesting the emotion or deepening the atmosphere. However, while the visual storytelling technique is often employed in creating professional recordings of a live concert, audience recordings of the same event often lack such sophisticated manipulations. Thus it would be useful to have a versatile system that can perform video mashup to create a refined video from such amateur clips. To this end, we propose to translate the music into a near-professional shot (type) sequence by learning the relation between music and visual storytelling of shots. The resulting shot sequence can then be used to better portray the visual storytelling of a song and guide the concert video mashup process. Our method introduces a novel probabilistic-based fusion approach, named as multi-resolution fused recurrent neural networks (MF-RNNs) with film-language, which integrates multi-resolution fused RNNs and a film-language model for boosting the translation performance. The results from objective and subjective experiments demonstrate that MF-RNNs with film-language can generate an appealing shot sequence with better viewing experience.
Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao
ICME2
2018 Coherent Deep-Net Fusion To Classify Shots In Concert Videos
abstract
Varying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director. The technique is often used in creating professional recordings of a live concert, but meanwhile may not be appropriately applied in audience recordings of the same event. Such variations could cause the task of classifying shots in concert videos, professional or amateur, very challenging. To achieve more reliable shot classification, we propose a novel probabilistic-based approach, named as coherent classification net (CC-Net), by addressing three crucial issues. First, we focus on learning more effective features by fusing the layer-wise outputs extracted from a deep convolutional neural network (CNN), pretrained on a large-scale data set for object recognition. Second, we introduce a frame-wise classification scheme, the error weighted deep cross-correlation model (EW-Deep-CCM), to boost the classification accuracy. Specifically, the deep neural network-based cross-correlation model (deep-CCM) is constructed to not only model the extracted feature hierarchies of CNN independently, but also relate the statistical dependencies of paired features from different layers. Then, a Bayesian error weighting scheme for a classifier combination is adopted to explore the contributions from individual Deep-CCM classifiers to enhance the accuracy of shot classification in each image frame. Third, we feed the frame-wise classification results to a linear-chain conditional random field module to refine the shot predictions by taking into account the global and temporal regularities. We provide extensive experimental results on a data set of live concert videos to demonstrate the advantage of the proposed CC-Net over existing popular fusion approaches for shot classification.
Jen-Chun Lin, Wen-Li Wei, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao
IEEE Trans. Multim.1
2017 Deep-net fusion to classify shots in concert videos
abstract
Varying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director to convey the emotion, ideas, and art. To classify such types of shots from images, we present a new framework that facilitates the intriguing task by addressing two key issues. We first focus on learning more effective features by fusing the layer-wise outputs extracted from a deep convolutional neural network (CNN), pre-trained on a large-scale dataset for object recognition. We then introduce a probabilistic fusion model, termed as error weighted deep cross-correlation model (EW-Deep-CCM), to boost the classification accuracy. Specifically, the deep neural network-based cross-correlation model (Deep-CCM) is constructed to not only model the extracted feature hierarchies of CNN independently but also relate the statistical dependencies of paired features from different layers. Then, a Bayesian error weighting scheme for classifier combination is adopted to explore the contributions from individual Deep-CCM classifiers to enhance the accuracy of shot classification. We provide extensive experimental results on a dataset of live concert videos to demonstrate the advantage of the proposed EW-Deep-CCM over existing popular fusion approaches. The video demos can be found at https://sites.google.com/site/ewdeepccm2/demo.
Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao
ICASSP2
2017 Automatic Music Video Generation Based on Simultaneous Soundtrack Recommendation and Video Editing
abstract
An automated process that can suggest a soundtrack to a user-generated video (UGV) and make the UGV a music-compliant professional-like video is challenging but desirable. To this end, this paper presents an automatic music video (MV) generation system that conducts soundtrack recommendation and video editing simultaneously. Given a long UGV, it is first divided into a sequence of fixed-length short (e.g., 2 seconds) segments, and then a multi-task deep neural network (MDNN) is applied to predict the pseudo acoustic (music) features (or called the pseudo song) from the visual (video) features of each video segment. In this way, the distance between any pair of video and music segments of same length can be computed in the music feature space. Second, the sequence of pseudo acoustic (music) features of the UGV and the sequence of the acoustic (music) features of each music track in the music collection are temporarily aligned by the dynamic time warping (DTW) algorithm with a pseudo-song-based deep similarity matching (PDSM) metric. Third, for each music track, the video editing module selects and concatenates the segments of the UGV based on the target and concatenation costs given by a pseudo-song-based deep concatenation cost (PDCC) metric according to the DTW-aligned result to generate a music-compliant professional-like video. Finally, all the generated MVs are ranked, and the best MV is recommended to the user. The MDNN for pseudo song prediction and the PDSM and PDCC metrics are trained by an annotated official music video (OMV) corpus. The results of objective and subjective experiments demonstrate that the proposed system performs well and can generate appealing MVs with better viewing and listening experiences.
Jen-Chun Lin, Wen-Li Wei, Hsin-Min Wang, Hong-Yuan Mark Liao
ACM Multimedia1
2017 Interaction Style Recognition Based on Multi-Layer Multi-View Profile Representation
abstract
Interaction Style (IS) refers to patterns of interaction containing highly contextual and innate information. Awareness of our IS can help us discover interpersonal conflicts and guide us how to interact with others. Recently, automatic IS recognition is becoming increasingly important in the design of a dialogue system for harmonious interaction. With the goal to select appropriate responses, four IS types proposed by Berens are selected as the basis for our study. In this study, multiple views (multi-views) of the utterances during interaction, including emotions and dialogue topics, are recognized first. Inspired by the emotion profile theory, the IS profiles are then extracted using the multi-view features to better characterize the IS of the interactional utterances. Similar to the multilayer architectures in deep neural networks, a multi-layer multi-view IS profile representation method, structured layer by layer through embedding the multi-views, is proposed to better interpret intermediate representations in the feature space of the interactional utterances based on a probabilistic fusion model. The IS is finally recognized by using the Support Vector Machine (SVM) based on the obtained IS profiles. Experimental results demonstrate that the proposed method achieved an encouraging IS recognition accuracy and outperformed the previous method.
Wen-Li Wei, Jen-Chun Lin, Chung-Hsien Wu 0001
IEEE Trans. Affect. Comput.2
2016 DEMV-matchmaker: Emotional temporal course representation and deep similarity matching for automatic music video generation
abstract
This paper presents a deep similarity matching-based emotion-oriented music video (MV) generation system, called DEMV-matchmaker, which utilizes an emotion-oriented deep similarity matching (EDSM) metric as a bridge to connect music and video. Specifically, we adopt an emotional temporal course model (ETCM) to respectively learn the relationship between music and its emotional temporal phase sequence and the relationship between video and its emotional temporal phase sequence from an emotion-annotated MV corpus. An emotional temporal structure preserved histogram (ETPH) representation is proposed to keep the recognized emotional temporal phase sequence information for EDSM metric construction. A deep neural network (DNN) is then applied to learn an EDSM metric based on the ETPHs for the given positive (official) and negative (artificial) MV examples. For MV generation, the EDSM metric is applied to measure the similarity between ETPHs of video and music. The results of objective and subjective experiments demonstrate that DEMV-matchmaker performs well and can generate appealing music videos that can enhance the viewing and listening experience.
Jen-Chun Lin, Wen-Li Wei, Hsin-Min Wang
ICASSP1
2016 Automatic Music Video Generation Based on Emotion-Oriented Pseudo Song Prediction and Matching
abstract
The main difficulty in automatic music video (MV) generation lies in how to match two different media (i.e., video and music). This paper proposes a novel content-based MV generation system based on emotion-oriented pseudo song prediction and matching. We use a multi-task deep neural network (MDNN) to jointly learn the relationship among music, video, and emotion from an emotion-annotated MV corpus. Given a queried video, the MDNN is applied to predict the acoustic (music) features from the visual (video) features, i.e., the pseudo song corresponding to the video. Then, the pseudo acoustic (music) features are matched with the acoustic (music) features of each music track in the music collection according to a pseudo-song-based deep similarity matching (PDSM) metric given by another deep neural network (DNN) trained on the acoustic and pseudo acoustic features of the positive (official), less-positive (artificial), and negative (artificial) MV examples. The results of objective and subjective experiments demonstrate that the proposed pseudo-song-based framework performs well and can generate appealing MVs with better viewing and listening experiences.
Jen-Chun Lin, Wen-Li Wei, Hsin-Min Wang
ACM Multimedia1
2015 Hierarchical modeling of temporal course in emotional expression for speech emotion recognition
abstract
This paper presents an approach to hierarchical modeling of temporal course in emotional expression for speech emotion recognition. In the proposed approach, a segmentation algorithm is employed to hierarchically chunk an input utterance into three-level temporal units, including low-level descriptors (LLDs)-based sub-utterance level, emotion profile (EP)-based sub-utterance level and utterance level. An emotion-oriented hierarchical structure is constructed based on the three-level units to describe the temporal emotion expression in an utterance. A hierarchical correlation model is also proposed to fuse the three-level outputs from the corresponding emotion recognizers and further model the correlation among them to determine the emotional state of the utterance. The EMO-DB corpus was used to evaluate the performance on speech emotion recognition. Experimental results show that the proposed method considering the temporal course in emotional expression provides the potential to improve the speech emotion recognition performance.
Chung-Hsien Wu 0001, Wei-Bin Liang, Kuan-Chun Cheng, Jen-Chun Lin
ACII4
2015 EMV-matchmaker: Emotional Temporal Course Modeling and Matching for Automatic Music Video Generation
abstract
This paper presents a novel content-based emotion-oriented music video (MV) generation system, called EMV-matchmaker, which utilizes the emotional temporal phase sequence of the multimedia content as a bridge to connect music and video. Specifically, we adopt an emotional temporal course model (ETCM) to respectively learn the relationship between music and its emotional temporal phase sequence and the relationship between video and its emotional temporal phase sequence from an emotion-annotated MV corpus. Then, given a video clip (or a music clip), the visual (or acoustic) ETCM is applied to predict its emotional temporal phase sequence in a valence-arousal (VA) emotional space from the corresponding low-level visual (or acoustic) features. For MV generation, string matching is applied to measure the similarity between the emotional temporal phase sequences of video and music. The results of objective and subjective experiments demonstrate that EMV-matchmaker performs well and can generate appealing music videos that can enhance the viewing and listening experience.
Jen-Chun Lin, Wen-Li Wei, Hsin-Min Wang
ACM Multimedia1
2014 Exploiting Psychological Factors for Interaction Style Recognition in Spoken Conversation
abstract
Determining how a speaker is engaged in a conversation is crucial for achieving harmonious interaction between computers and humans. In this study, a fusion approach was developed based on psychological factors to recognize Interaction Style ($IS$) in spoken conversation, which plays a key role in creating natural dialogue agents. The proposed Fused Cross-Correlation Model (FCCM) provides a unified probabilistic framework to model the relationships among the psychological factors of emotion, personality trait ($PT$), transient$IS$, and$IS$history, for recognizing$IS$. An emotional arousal-dependent speech recognizer was used to obtain the recognized spoken text for extracting linguistic features to estimate transient$IS$likelihood and recognize$PT$. A temporal course modeling approach and an emotional sub-state language model, based on the temporal phases of an emotional expression, were employed to obtain a better emotion recognition result. The experimental results indicate that the proposed FCCM yields satisfactory results in$IS$recognition and also demonstrate that combining psychological factors effectively improves$IS$recognition accuracy.
Wen-Li Wei, Chung-Hsien Wu 0001, Jen-Chun Lin
IEEE ACM Trans. Audio Speech Lang. Process.3
2013 Facial action unit prediction under partial occlusion based on Error Weighted Cross-Correlation Model
abstract
Occlusive effect is a crucial issue that may dramatically degrade performance on facial expression recognition. As emotion recognition from facial expression is based on the entire facial feature, occlusive effect remains a challenging problem to be solved. To manage this problem, an Error Weighted Cross-Correlation Model (EWCCM) is proposed to effectively predict the facial Action Unit (AU) under partial facial occlusion from non-occluded facial regions for providing the correct AU information for emotion recognition. The Gaussian Mixture Model (GMM)-based Cross-Correlation Model (CCM) in EWCCM is first proposed not only modeling the extracted facial features but also constructing the statistical dependency among features from paired facial regions for AU prediction. The Bayesian classifier weighting scheme is then adopted to explore the contributions of the GMM-based CCMs to enhance the prediction accuracy. Experiments show that a promising result of the proposed approach can be obtained.
Jen-Chun Lin, Chung-Hsien Wu 0001, Wen-Li Wei
ICASSP1
2013 Interaction style detection based on Fused Cross-Correlation Model in spoken conversation
abstract
In recent years, much attention has been given to dialogue strategy design to achieve intelligent speech-based human-computer interaction. Since speakers generally express their intents in different Interaction Styles (ISs), the responses of a spoken dialogue system should be versatile instead of invariable and planned. This paper presents an approach to automatic detection of a user's IS using a Fused Cross-Correlation Model (FCCM). As IS generally involves high level psychological meaning, cross-correlation among various psychological factors including emotion, personality trait, and IS is thus considered for IS detection modeling. The Bayes' theorem is then used to integrate the cross-correlation into the IS detector for enhancing the IS detection accuracy. Experiments show a promising result of the proposed approach.
Wen-Li Wei, Chung-Hsien Wu 0001, Jen-Chun Lin
ICASSP3
2013 Emotion recognition of conversational affective speech using temporal course modeling
Jen-Chun Lin, Chung-Hsien Wu 0001, Wen-Li Wei
INTERSPEECH1
2013 Two-Level Hierarchical Alignment for Semi-Coupled HMM-Based Audiovisual Emotion Recognition With Temporal Course
abstract
A complete emotional expression typically contains a complex temporal course in face-to-face natural conversation. To address this problem, a bimodal hidden Markov model (HMM)-based emotion recognition scheme, constructed in terms of sub-emotional states, which are defined to represent temporal phases of onset, apex, and offset, is adopted to model the temporal course of an emotional expression for audio and visual signal streams. A two-level hierarchical alignment mechanism is proposed to align the relationship within and between the temporal phases in the audio and visual HMM sequences at the model and state levels in a proposed semi-coupled hidden Markov model (SC-HMM). Furthermore, by integrating a sub-emotion language model, which considers the temporal transition between sub-emotional states, the proposed two-level hierarchical alignment-based SC-HMM (2H-SC-HMM) can provide a constraint on allowable temporal structures to determine an optimal emotional state. Experimental results show that the proposed approach can yield satisfactory results in both the posed MHMC and the naturalistic SEMAINE databases, and shows that modeling the complex temporal structure is useful to improve the emotion recognition performance, especially for the naturalistic database (i.e., natural conversation). The experimental results also confirm that the proposed 2H-SC-HMM can achieve an acceptable performance for the systems with sparse training data or noisy conditions.
Chung-Hsien Wu 0001, Jen-Chun Lin, Wen-Li Wei
IEEE Trans. Multim.2
2013 Speaking Effect Removal on Emotion Recognition From Facial Expressions Based on Eigenface Conversion
abstract
Speaking effect is a crucial issue that may dramatically degrade performance in emotion recognition from facial expressions. To manage this problem, an eigenface conversion-based approach is proposed to remove speaking effect on facial expressions for improving accuracy of emotion recognition. In the proposed approach, a context-dependent linear conversion function modeled by a statistical Gaussian Mixture Model (GMM) is constructed with parallel data from speaking and non-speaking facial expressions with emotions. To model the speaking effect in more detail, the conversion functions are categorized using a decision tree considering the visual temporal context of the Articulatory Attribute (AA) classes of the corresponding input speech segments. For verification of the identified quadrant of emotional expression on the Arousal-Valence (A-V) emotion plane, which is commonly used to dimensionally define the emotion classes, from the reconstructed facial feature points, an expression template is constructed to represent the feature points of the non-speaking facial expressions for each quadrant. With the verified quadrant, a regression scheme is further employed to estimate the A-V values of the facial expression as a precise point in the A-V emotion plane. Experimental results show that the proposed method outperforms current approaches and demonstrates that removing the speaking effect on facial expression is useful for improving the performance of emotion recognition.
Chung-Hsien Wu 0001, Wen-Li Wei, Jen-Chun Lin, Wei-Yu Lee
IEEE Trans. Multim.3
2012 Error Weighted Semi-Coupled Hidden Markov Model for Audio-Visual Emotion Recognition
abstract
This paper presents an approach to the automatic recognition of human emotions from audio-visual bimodal signals using an error weighted semi-coupled hidden Markov model (EWSC-HMM). The proposed approach combines an SC-HMM with a state-based bimodal alignment strategy and a Bayesian classifier weighting scheme to obtain the optimal emotion recognition result based on audio-visual bimodal fusion. The state-based bimodal alignment strategy in SC-HMM is proposed to align the temporal relation between audio and visual streams. The Bayesian classifier weighting scheme is then adopted to explore the contributions of the SC-HMM-based classifiers for different audio-visual feature pairs in order to obtain the emotion recognition output. For performance evaluation, two databases are considered: the MHMC posed database and the SEMAINE naturalistic database. Experimental results show that the proposed approach not only outperforms other fusion-based bimodal emotion recognition methods for posed expressions but also provides satisfactory results for naturalistic expressions.
Jen-Chun Lin, Chung-Hsien Wu 0001, Wen-Li Wei
IEEE Trans. Multim.1
2011 Semi-Coupled Hidden Markov Model with State-Based Alignment Strategy for Audio-Visual Emotion Recognition
Jen-Chun Lin, Chung-Hsien Wu 0001, Wen-Li Wei
ACII (1)1
2011 A Regression Approach to Affective Rating of Chinese Words from ANEW
Wen-Li Wei, Chung-Hsien Wu 0001, Jen-Chun Lin
ACII (2)3