Suping Zhou

dblp:195/8208 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
1since 2021 · last 2021
0000-0002-3472-1434ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Speech recognition and synthesis · 41% Learning paradigms · 26% Information extraction and text analysis · 22%
Human-computer interaction and pervasive computing
2 papers
Learning and educational technologies · 52% Haptics and multimodal interaction · 24% Human-robot interaction · 24%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 8 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis › paralinguistic analysis
speech emotion recognition
1.232021
Inferring Emotion from Large-scale Internet Voice Data: A Semi-supervised Curriculum Augmentation based Deep Learning Approach · AAAI 2021
Inferring Emotions From Large-Scale Internet Voice Data · IEEE Trans. Multim. 2019
Inferring Emotion from Conversational Voice Data: A Semi-Supervised Multi-Path Generative Neural Network Approach · AAAI 2018
Machine learning › Learning paradigms
semi-supervised learning
0.822021
Inferring Emotion from Large-scale Internet Voice Data: A Semi-supervised Curriculum Augmentation based Deep Learning Approach · AAAI 2021
Inferring Emotion from Conversational Voice Data: A Semi-Supervised Multi-Path Generative Neural Network Approach · AAAI 2018
Natural language and speech › Information extraction and text analysis › emotion recognition
emotion recognition in conversation
0.312018
Inferring Emotion from Conversational Voice Data: A Semi-Supervised Multi-Path Generative Neural Network Approach · AAAI 2018
Machine learning › Generative modeling
variational autoencoder
0.312018
Inferring Emotion from Conversational Voice Data: A Semi-Supervised Multi-Path Generative Neural Network Approach · AAAI 2018
Human-robot interaction
affective interaction
0.312018
IcooBook: When the Picture Book for Children Encounters Aesthetics of Interaction · ACM Multimedia 2018
Haptics and multimodal interaction
multisensory interaction
0.312018
IcooBook: When the Picture Book for Children Encounters Aesthetics of Interaction · ACM Multimedia 2018
Multimedia analysis and retrieval › image analysis
fashion analysis
0.312017
Towards Better Understanding the Clothing Fashion Styles: A Multimodal Deep Learning Approach · AAAI 2017
Multimedia analysis and retrieval
multimodal learning
0.312017
Towards Better Understanding the Clothing Fashion Styles: A Multimodal Deep Learning Approach · AAAI 2017

Methods — techniques the papers use, named apart from their topics

semi-supervised learning · 0.5multi-path mix-match multimodal neural network · 0.5curriculum learning · 0.5multi-task deep learning · 0.4long short-term memory · 0.4latent dirichlet allocation · 0.4deep sparse neural network · 0.4cross-media feature fusion · 0.4attention mechanism · 0.4semi-supervised variational autoencoder · 0.3multi-path neural network · 0.3emotion recognition · 0.3fashion semantic space · 0.3bimodal correlative deep autoencoder · 0.3
YearPublicationVenuePosition
2021 Inferring Emotion from Large-scale Internet Voice Data: A Semi-supervised Curriculum Augmentation based Deep Learning Approach
abstract
Effective emotion inference from user queries helps to give a more personified response for Voice Dialogue Applications(VDAs). The tremendous amounts of VDA users bring in diverse emotion expressions. How to achieve a high emotion inferring performance from large-scale Internet Voice Data in VDAs? Traditionally, researches on speech emotion recognition are based on acted voice datasets, which have limited speakers but strong and clear emotion expressions. Inspired by this, in this paper, we propose a novel approach to leverage acted voice data with strong emotion expressions to enhance large-scale unlabeled internet voice data with diverse emotion expressions for emotion inferring. Specifically, we propose a novel semi-supervised multi-modal curriculum augmentation deep learning framework. First, to learn more general emotion cues, we adopt a curriculum learning based epoch-wise training strategy, which trains our model guided by strong and balanced emotion samples from acted voice data and sub-sequently leverages weak and unbalanced emotion samples from internet voice data.Second, to employ more diverse emotion expressions, we design a Multi-path Mix-match Multimodal Deep Neural Network(MMMD), which effectively learns feature representations for multiple modalities and trains labeled and unlabeled data in hybrid semi-supervised methods for superior generalization and robustness. Experiments on an internet voice dataset with 500,000 utterances show our method outperforms (+10.09% in terms of F1) several alternative baselines, while an acted corpus with 2,397 utterances contributes 4.35%. To further compare our method with state-of-the-art techniques in traditionally acted voice datasets, we also conduct experiments on public dataset IEMOCAP. The results reveal the effectiveness of the proposed approach.
Suping Zhou, Jia Jia 0001, Zhiyong Wu 0001, Wei Chen 0071, Shuo Huang 0005, Jialie Shen 0001
AAAI1
2020 Inferring Emphasis for Real Voice Data: An Attentive Multimodal Neural Network Approach
Suping Zhou, Jia Jia 0001, Wei Chen 0071, Jialie Shen 0001
MMM (2)1
2019 Understanding the Teaching Styles by an Attention based Multi-task Cross-media Dimensional Modeling
abstract
Teaching style plays an influential role in helping students to achieve academic success. In this paper, we explore a new problem of effectively understanding teachers' teaching styles. Specifically, we study 1) how to quantitatively characterize various teachers' teaching styles for various teachers and 2) how to model the subtle relationship between cross-media teaching related data (speech, facial expressions and body motions, content et al.) and teaching styles. Using the adjectives selected from more than 10,000 feedback questionnaires provided by an educational enterprise, a novel concept called Teaching Style Semantic Space (TSSS) is developed based on the pleasure-arousal dimensional theory to describe teaching styles quantitatively and comprehensively. Then a multi-task deep learning based model, Attention-based Multi-path Multi-task Deep Neural Network (AMMDNN), is proposed to accurately and robustly capture the internal correlations between cross-media features and TSSS. Based on the benchmark dataset, we further develop a comprehensive data set including 4,541 full-annotated cross-modality teaching classes. Our experimental results demonstrate that the proposed AMMDNN outperforms (+0.0842% in terms of the concordance correlation coefficient (CCC) on average) baseline methods. To further demonstrate the advantages of the proposed TSSS and our model, several interesting case studies are carried out, such as teaching styles comparison among different teachers and courses, and leveraging the proposed method for teaching quality analysis.
Suping Zhou, Jia Jia 0001, Yufeng Yin 0002, Xiang Li 0105, Zeyang Ye, Kehua Lei, Jialie Shen 0001
ACM Multimedia1
2019 Inferring Emotions From Large-Scale Internet Voice Data
abstract
As voice dialog applications (VDAs, e.g., Siri,11http://www.apple.com/ios/siri/. Cortana,22http://www.microsoft.com/en-us/mobile/campaign-cortana/. Google Now33http://www.google.com/landing/now/.) are increasing in popularity, inferring emotions from the large-scale internet voice data generated from VDAs can help give a more reasonable and humane response. However, the tremendous amounts of users in large-scale internet voice data lead to a great diversity of users accents and expression patterns. Therefore, the traditional speech emotion recognition methods, which mainly target acted corpora, cannot effectively handle the massive and diverse amount of internet voice data. To address this issue, we carry out a series of observations, find suitable emotion categories for large-scale internet voice data, and verify the indicators of the social attributes (query time, query topic, and users location) and emotion inferring. Based on our observations, two different strategies are employed to solve the problem. First, a deep sparse neural network model that uses acoustic information, textual information, and three indicators (a temporal indicator, descriptive indicator, and geo-social indicator) as the input is proposed. Then, to capture the contextual information, we propose a hybrid emotion inference model that includes long short-term memory to capture the acoustic features and a latent dirichlet allocation to extract text features. Experiments on 93 000 utterances collected from the Sogou Voice Assistant44http://yy.sogou.com. (Chinese Siri) validate the effectiveness of the proposed methodologies. Furthermore, we compare the two methodologies and give their advantages and disadvantages.
Jia Jia 0001, Suping Zhou, Yufeng Yin 0002, Boya Wu, Wei Chen 0071
IEEE Trans. Multim.2
2018 Inferring Emotion from Conversational Voice Data: A Semi-Supervised Multi-Path Generative Neural Network Approach
abstract
To give a more humanized response in Voice Dialogue Applications (VDAs), inferring emotion states from users’ queries may play an important role. However, in VDAs, we have tremendous amount of VDA users and massive scale of unlabeled data with high dimension features from multimodal information, which challenge the traditional speech emotion recognition methods. In this paper, to better infer emotion from conversational voice data, we proposed a semi-supervised multi-path generative neural network. Specifically, first, we build a novel supervised multi-path deep neural network framework. To avoid high dimensional input, raw features are trained by groups in local classifiers. Then high-level features of each local classifiers are concatenated as input of a global classifier. These two kinds classifiers are trained simultaneously through a single objective function to achieve a more effective and discriminative emotion inferring. To further solve the labeled-data-scarcity problem, we extend the multi-path deep neural network to a generative model based on semi-supervised variational autoencoder (semi-VAE), which is able to train the labeled and unlabeled data simultaneously. Experiment based on a 24,000 real-world dataset collected from Sogou Voice Assistant (SVAD13) and a benchmark dataset IEMOCAP show that our method significantly outperforms the existing state-of-the-art results.
Suping Zhou, Jia Jia 0001, Yufei Dong, Yufeng Yin 0002, Kehua Lei
AAAI1
2018 IcooBook: When the Picture Book for Children Encounters Aesthetics of Interaction
abstract
In this work, we propose a novel PCA (Perception & Cognition & Affection) model from the prospective of aesthetics in interaction. Based on PCA, we establish a new electronic interactive picture book for children, named IcooBook. At the first level of perception, the proposed IcooBook provides interfaces of multi-sensory interaction; at the second level of cognition, IcooBook builds immersive interactive scenes; at the third level of affection, IcooBook creates high-level interaction modes based on automatic emotion recognition. The research on user study had proved the effectiveness of IcooBook in helping children being focusing on reading, getting better understanding about the context, and further encouraging children to appreciate the beauty of deep affective interaction.
Yaohua Bu, Jia Jia 0001, Xiang Li 0105, Suping Zhou, Xiaobo Lu
ACM Multimedia4
2017 Towards Better Understanding the Clothing Fashion Styles: A Multimodal Deep Learning Approach
abstract
In this paper, we aim to better understand the clothing fashion styles. There remain two challenges for us: 1) how to quantitatively describe the fashion styles of various clothing, 2) how to model the subtle relationship between visual features and fashion styles, especially considering the clothing collocations. Using the words that people usually use to describe clothing fashion styles on shopping websites, we build a Fashion Semantic Space (FSS) based on Kobayashi's aesthetics theory to describe clothing fashion styles quantitatively and universally. Then we propose a novel fashion-oriented multimodal deep learning based model, Bimodal Correlative Deep Autoencoder (BCDA), to capture the internal correlation in clothing collocations. Employing the benchmark dataset we build with 32133 full-body fashion show images, we use BCDA to map the visual features to the FSS. The experiment results indicate that our model outperforms (+13% in terms of MSE) several alternative baselines, confirming that our model can better understand the clothing fashion styles. To further demonstrate the advantages of our model, we conduct some interesting case studies, including fashion trends analyses of brands, clothing collocation recommendation, etc.
Yihui Ma, Jia Jia 0001, Suping Zhou, Jingtian Fu, Yejun Liu, Zijian Tong
AAAI3