Yaohua Bu

dblp:182/4193 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
2since 2021 · last 2023
0009-0005-3024-0704ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-authorArtificial intelligence and machine learning · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Speech recognition and synthesis · 41% Graph learning · 14% Generative modeling · 14%
Human-computer interaction and pervasive computing
4 papers
Learning and educational technologies · 61% Haptics and multimodal interaction · 16% Accessibility and assistive technology · 11%
Computer graphics and multimedia
3 papers
Computer animation and physical simulation · 31% Visual content generation and editing · 26% Computational photography and imaging · 26%
Databases, data mining, and information retrieval
2 papers
Recommender systems · 70% Web and social media mining · 30%

Topics — the 21 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Learning and educational technologies › language learning
computer-assisted language learning
0.512021
PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback · CHI 2021
Learning and educational technologies › language learning
pronunciation training
0.512021
PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback · CHI 2021
Machine learning › Generative modeling
motion generation
0.412020
ChoreoNet: Towards Music to Dance Synthesis with Choreographic Action Unit · ACM Multimedia 2020
Machine learning › Graph learning
social network analysis
0.412020
Mining Unfollow Behavior in Large-Scale Online Social Networks via Spatial-Temporal Interaction · AAAI 2020
Natural language and speech › Speech recognition and synthesis
speech-driven animation
0.412020
Visual-speech Synthesis of Exaggerated Corrective Feedback · ACM Multimedia 2020
Natural language and speech › Speech recognition and synthesis
speech synthesis
0.412020
Visual-speech Synthesis of Exaggerated Corrective Feedback · ACM Multimedia 2020
Recommender systems
music recommendation
0.412020
PEIA: Personality and Emotion Integrated Attentive Model for Music Recommendation on Social Media Platforms · AAAI 2020
Recommender systems › social recommendation
social media recommendation
0.412020
PEIA: Personality and Emotion Integrated Attentive Model for Music Recommendation on Social Media Platforms · AAAI 2020
Visual content generation and editing › style transfer
image style transfer
0.412020
Aesthetic-Aware Image Style Transfer · ACM Multimedia 2020
Learning and educational technologies › language learning
computer-aided pronunciation training
0.412020
Visual-speech Synthesis of Exaggerated Corrective Feedback · ACM Multimedia 2020
Machine learning › Representation and self-supervised learning › representation learning › feature extraction
discriminative feature learning
0.412019
Towards Discriminative Representation Learning for Speech Emotion Recognition · IJCAI 2019
Natural language and speech › Speech recognition and synthesis › paralinguistic analysis
speech emotion recognition
0.412019
Towards Discriminative Representation Learning for Speech Emotion Recognition · IJCAI 2019
Computer vision › Face, body and person analysis › facial expression analysis
facial expression recognition
0.312018
Lookine: Let the Blind Hear a Smile · AAAI 2018
Human-robot interaction
affective interaction
0.312018
IcooBook: When the Picture Book for Children Encounters Aesthetics of Interaction · ACM Multimedia 2018
Accessibility and assistive technology
assistive technology for visual impairment
0.312018
Lookine: Let the Blind Hear a Smile · AAAI 2018
Haptics and multimodal interaction
multisensory interaction
0.312018
IcooBook: When the Picture Book for Children Encounters Aesthetics of Interaction · ACM Multimedia 2018
Haptics and multimodal interaction › multimodal feedback
audiovisual feedback
0.112021
PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback · CHI 2021
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatio-temporal interaction modeling
0.112020
Mining Unfollow Behavior in Large-Scale Online Social Networks via Spatial-Temporal Interaction · AAAI 2020
Recommender systems
user profiling
0.112020
PEIA: Personality and Emotion Integrated Attentive Model for Music Recommendation on Social Media Platforms · AAAI 2020
Computer vision › Face, body and person analysis
head pose estimation
0.112018
Lookine: Let the Blind Hear a Smile · AAAI 2018
Computer animation and physical simulation › motion synthesis › human motion synthesis
dance generation
0.112017
AniDraw: When Music and Dance Meet Harmoniously · AAAI 2017

Methods — techniques the papers use, named apart from their topics

viseme blending · 0.9tacotron · 0.9spatial-temporal inpainting · 0.9heterogeneous information model · 0.9choreographic action unit prediction · 0.9MLP layers · 0.9facial action recognition · 0.7user study · 0.5interviews · 0.5multi-reference style transfer · 0.4model optimization · 0.4hierarchical attention · 0.4attentive model · 0.4multi-head self-attention · 0.4global context-aware attention LSTM · 0.4emotion recognition · 0.3
YearPublicationVenuePosition
2023 "I Never Envy Anyone, for I Have Already Built a Kingdom With My Fingertips": Exploring Teenagers' Experience in Chat-based Cosplay Community
abstract
This paper reports an interview study about the practice of teenagers’ chat-based cosplay in China. Findings reveal the four primary motivations of the participants and their main practice in chat-based cosplay. We found that adolescents perceived character presentation and portrayal as a central aspect of chat-based cosplay and they devoted significant effort to refine their characters to achieve higher character consistency. We highlighted the positive feedback loop between social relationships and story creation in chat-based cosplay community. In addition, we identified the influence and negative experiences on adolescents in the chat-based cosplay community.
Yaohua Bu, Suqi Lou, Shi Chen 0005, Lingyun Sun, Chang-yuan Yang
IDC2
2021 PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback
abstract
Second language (L2) English learners often find it difficult to improve their pronunciations due to the lack of expressive and personalized corrective feedback. In this paper, we present Pronunciation Teacher (PTeacher), a Computer-Aided Pronunciation Training (CAPT) system that provides personalized exaggerated audio-visual corrective feedback for mispronunciations. Though the effectiveness of exaggerated feedback has been demonstrated, it is still unclear how to define the appropriate degrees of exaggeration when interacting with individual learners. To fill in this gap, we interview 100 L2 English learners and 22 professional native teachers to understand their needs and experiences. Three critical metrics are proposed for both learners and teachers to identify the best exaggeration levels in both audio and visual modalities. Additionally, we incorporate the personalized dynamic feedback mechanism given the English proficiency of learners. Based on the obtained insights, a comprehensive interactive pronunciation training course is designed to help L2 learners rectify mispronunciations in a more perceptible, understandable, and discriminative manner. Extensive user studies demonstrate that our system significantly promotes the learners’ learning efficiency.
Yaohua Bu, Hang Zhou 0009, Jia Jia 0001, Shengqi Chen 0001, Dachuan Shi, Haozhe Wu, Kun Li 0003, Zhiyong Wu 0001, Yuanchun Shi, Xiaobo Lu, Ziwei Liu 0002
CHI1
2020 PEIA: Personality and Emotion Integrated Attentive Model for Music Recommendation on Social Media Platforms
abstract
With the rapid expansion of digital music formats, it's indispensable to recommend users with their favorite music. For music recommendation, users' personality and emotion greatly affect their music preference, respectively in a long-term and short-term manner, while rich social media data provides effective feedback on these information. In this paper, aiming at music recommendation on social media platforms, we propose a Personality and Emotion Integrated Attentive model (PEIA), which fully utilizes social media data to comprehensively model users' long-term taste (personality) and short-term preference (emotion). Specifically, it takes full advantage of personality-oriented user features, emotion-oriented user features and music features of multi-faceted attributes. Hierarchical attention is employed to distinguish the important factors when incorporating the latent representations of users' personality and emotion. Extensive experiments on a large real-world dataset of 171,254 users demonstrate the effectiveness of our PEIA model which achieves an NDCG of 0.5369, outperforming the state-of-the-art methods. We also perform detailed parameter analysis and feature contribution analysis, which further verify our scheme and demonstrate the significance of co-modeling of user personality and emotion in music recommendation.
Tiancheng Shen, Jia Jia 0001, Yan Li 0068, Yihui Ma, Yaohua Bu, Hanjie Wang, Tat-Seng Chua, Wendy Hall 0001
AAAI5
2020 Mining Unfollow Behavior in Large-Scale Online Social Networks via Spatial-Temporal Interaction
abstract
Online Social Networks (OSNs) evolve through two pervasive behaviors: follow and unfollow, which respectively signify relationship creation and relationship dissolution. Researches on social network evolution mainly focus on the follow behavior, while the unfollow behavior has largely been ignored. Mining unfollow behavior is challenging because user's decision on unfollow is not only affected by the simple combination of user's attributes like informativeness and reciprocity, but also affected by the complex interaction among them. Meanwhile, prior datasets seldom contain sufficient records for inferring such complex interaction. To address these issues, we first construct a large-scale real-world Weibo1 dataset, which records detailed post content and relationship dynamics of 1.8 million Chinese users. Next, we define user's attributes as two categories: spatial attributes (e.g., social role of user) and temporal attributes (e.g., post content of user). Leveraging the constructed dataset, we systematically study how the interaction effects between user's spatial and temporal attributes contribute to the unfollow behavior. Afterwards, we propose a novel unified model with heterogeneous information (UMHI) for unfollow prediction. Specifically, our UMHI model: 1) captures user's spatial attributes through social network structure; 2) infers user's temporal attributes through user-posted content and unfollow history; and 3) models the interaction between spatial and temporal attributes by the nonlinear MLP layers. Comprehensive evaluations on the constructed dataset demonstrate that the proposed UMHI model outperforms baseline methods by 16.44 on average in terms of precision. In addition, factor analyses verify that both spatial attributes and temporal attributes are essential for mining unfollow behavior.
Haozhe Wu, Jia Jia 0001, Yaohua Bu, Xiangnan He 0001, Tat-Seng Chua
AAAI4
2020 Visual-speech Synthesis of Exaggerated Corrective Feedback
abstract
To provide more discriminative feedback for the second language (L2) learners to better identify their mispronunciation, we propose a method for exaggerated visual-speech feedback in computer-assisted pronunciation training (CAPT). The speech exaggeration is realized by an emphatic speech generation neural network based on Tacotron, while the visual exaggeration is accomplished by ADC Viseme Blending, namely increasing Amplitude of movement, extending the phone's Duration and enhancing the color Contrast. User studies show that exaggerated feedback outperforms non-exaggerated version on helping learners with pronunciation identification and pronunciation improvement.
Yaohua Bu, Shengqi Chen 0001, Jia Jia 0001, Kun Li 0003, Xiaobo Lu
ACM Multimedia1
2020 Aesthetic-Aware Image Style Transfer
abstract
Style transfer aims to synthesize an image which inherits the content of one image while preserving a similar style of the other one. The "style'' of an image usually refers to its unique feeling conveyed from visual features, which is highly related to the aesthetic effect of the image. Aesthetic effect can be mainly decomposed as two factors: colour and texture. Previous methods like Neural Style Transfer and Colour Transfer have shown strong abilities in transferring colour and texture features. However, such approaches neglect to further disentangle colour and texture, which makes some of unique aesthetic effects designed by human artists hard to express. In this paper, we propose a novel problem called Aesthetic-Aware Image Style Transfer task, which aims to transfer colour and texture separately and independently to manipulate the aesthetic effect of an image. We propose a novel Aesthetic-Aware Model-Optimisation-Based Style Transfer (AAMOBST) model to solve this problem. Specifically, AAMOBST is a multi-reference, two-path model. It uses different reference images to decide desired colour and texture features. It can segregate colour and texture into two distinct paths and transfer them independently. Qualitative and quantitative experiments show that our model can decide colour and texture features separately and is able to keep one of them fixed while changing the other one, which is not applicable for previous methods. Furthermore, on tasks that are applicable for previous methods (such as style transfer, colour-preserved transfer and colour-only transfer), our model shows comparable abilities with other baseline methods.
Jia Jia 0001, Bei Liu 0001, Yaohua Bu, Jianlong Fu
ACM Multimedia4
2020 ChoreoNet: Towards Music to Dance Synthesis with Choreographic Action Unit
abstract
Dance and music are two highly correlated artistic forms. Synthesizing dance motions has attracted much attention recently. Most previous works conduct music-to-dance synthesis via directly music to human skeleton keypoints mapping. Meanwhile, human choreographers design dance motions from music in a two-stage manner: they firstly devise multiple choreographic dance units (CAUs), each with a series of dance motions, and then arrange the CAU sequence according to the rhythm, melody and emotion of the music. Inspired by these, we systematically study such two-stage choreography approach and construct a dataset to incorporate such choreography knowledge. Based on the constructed dataset, we design a two-stage music-to-dance synthesis framework ChoreoNet to imitate human choreography procedure. Our framework firstly devises a CAU prediction model to learn the mapping relationship between music and CAU sequences. Afterwards, we devise a spatial-temporal inpainting model to convert the CAU sequence into continuous dance motions. Experimental results demonstrate that the proposed ChoreoNet outperforms baseline methods (0.622 in terms of CAU BLEU score and 1.59 in terms of user study score).
Zijie Ye, Haozhe Wu, Jia Jia 0001, Yaohua Bu, Wei Chen 0071
ACM Multimedia4
2019 Towards Discriminative Representation Learning for Speech Emotion Recognition
abstract
In intelligent speech interaction, automatic speech emotion recognition (SER) plays an important role in understanding user intention. While sentimental speech has different speaker characteristics but similar acoustic attributes, one vital challenge in SER is how to learn robust and discriminative representations for emotion inferring. In this paper, inspired by human emotion perception, we propose a novel representation learning component (RLC) for SER system, which is constructed with Multi-head Self-attention and Global Context-aware Attention Long Short-Term Memory Recurrent Neutral Network (GCA-LSTM). With the ability of Multi-head Self-attention mechanism in modeling the element-wise correlative dependencies, RLC can exploit the common patterns of sentimental speech features to enhance emotion-salient information importing in representation learning. By employing GCA-LSTM, RLC can selectively focus on emotion-salient factors with the consideration of entire utterance context, and gradually produce discriminative representation for emotion inferring. Experiments on public emotional benchmark database IEMOCAP and a tremendous realistic interaction database demonstrate the outperformance of the proposed SER framework, with 6.6% to 26.7% relative improvement on unweighted accuracy compared to state-of-the-art techniques.
Runnan Li, Zhiyong Wu 0001, Jia Jia 0001, Yaohua Bu, Sheng Zhao 0002, Helen M. Meng
IJCAI4
2018 Lookine: Let the Blind Hear a Smile
abstract
It is believed that nonverbal visual information including facial expressions, facial micro-actions and head movements plays a significant role in fundamental social communication. Unfortunately it is regretful that the blind can not achieve such necessary information. Therefore, we propose a social assistant system, Lookine, to help them to go beyond this limitation. For Lookine, we apply the novel techniques including facial expression recognition, facial action recognition and head pose estimation, and obey barrier-free principles in our design. In experiments, the algorithm evaluation and user study prove that our system has promising accuracy, good real-time performance, and great user experience.
Yaohua Bu, Jia Jia 0001, Yuhan Tang, Xuan Zang
AAAI1
2018 Understanding The Aesthetic Styles of Social Images
abstract
Aesthetic perception is nearly the most direct impact people could receive from images. Recent research on image understanding is mainly focused on image analysis, recognition and classification, regardless of the aesthetic meanings embedded in images. In this paper, we systematically study the problem of understanding the aesthetic styles of social images. First, we build a two-dimensional Image Aesthetic Space (IAS) to describe image aesthetic styles quantitatively and universally. Then, we propose a Bimodal Deep Autoen-coder with Cross Edges (BDA-CE) to deeply fuse the social image related features (i.e. images' visual features, tags' textual features). Connecting BDA-CE with a regression model, we are able to map the features to the IAS. The experimental results on the benchmark dataset we build with 120 thousand Flickr images show that our model outperforms (+5.5% in terms of MSE) alternative baselines. Furthermore, we conduct an interesting case study to demonstrate the advantages of our methods.
Yihui Ma, Jia Jia 0001, Yufan Hou, Yaohua Bu
ICASSP4
2018 IcooBook: When the Picture Book for Children Encounters Aesthetics of Interaction
abstract
In this work, we propose a novel PCA (Perception & Cognition & Affection) model from the prospective of aesthetics in interaction. Based on PCA, we establish a new electronic interactive picture book for children, named IcooBook. At the first level of perception, the proposed IcooBook provides interfaces of multi-sensory interaction; at the second level of cognition, IcooBook builds immersive interactive scenes; at the third level of affection, IcooBook creates high-level interaction modes based on automatic emotion recognition. The research on user study had proved the effectiveness of IcooBook in helping children being focusing on reading, getting better understanding about the context, and further encouraging children to appreciate the beauty of deep affective interaction.
Yaohua Bu, Jia Jia 0001, Xiang Li 0105, Suping Zhou, Xiaobo Lu
ACM Multimedia1
2017 AniDraw: When Music and Dance Meet Harmoniously
Yaohua Bu, Taoran Tang, Jia Jia 0001, Songyao Wu, Yuming You
AAAI1
2017 Better deep visual attention with reinforcement learning in action recognition
abstract
Deep visual attention in computer vision has attracted much attention over the past years, which achieves great contributions especially in image classification, image caption and action recognition. However, due to taking BP training wholly or partially, they can not show the true power of attention in computational efficiency and focusing accuracy. Our intuition is that attention mechanism should be similar to the process in which human draw attention and select the next location to focus, by observing, analyzing and jumping instead of existing describing continuous features. Based on this insight, we formulate our model as a recurrent neural network-based agent that chooses attention region by reinforcement learning at each timestep. In experiments, our model explicitly outperforms baselines not only in focusing and recognizing accuracy, but also consumes much less computational resources, which can be honored as better deep visual attention.
Wenmin Wang 0001, Jingzhuo Wang, Yaohua Bu
ISCAS4