Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yongjin Wang

dblp:65/5570 · DBLP profile ↗
← Back
27ranked-venue papers
11as first author
2since 2021 · last 2025
0000-0001-8109-4640ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 first-authorComputer networks · 5Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
5 papers
Physical-layer communications · 72% Wireless sensing and localization · 14% Edge and fog computing · 4%
Artificial intelligence
3 papers
Face, body and person analysis · 77% Representation and self-supervised learning · 23%

Topics — the 20 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Physical-layer communications
full-duplex communication
0.912025
Full-duplex ultraviolet light communication network for space-chip interconnection · Sci. China Inf. Sci. 2025
Physical-layer communications
optical wireless communication
0.912025
Mobile wireless light communication network · Sci. China Inf. Sci. 2025
Physical-layer communications › optical wireless communication
ultraviolet communication
0.912025
Full-duplex ultraviolet light communication network for space-chip interconnection · Sci. China Inf. Sci. 2025
Physical-layer communications › diversity combining
combining schemes
0.312018
A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading Channels · IEEE Trans. Commun. 2018
Physical-layer communications › fading channels › correlated fading
correlated lognormal fading
0.312018
A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading Channels · IEEE Trans. Commun. 2018
Physical-layer communications
diversity combining
0.312018
A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading Channels · IEEE Trans. Commun. 2018
Physical-layer communications
fading channels
0.312018
A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading Channels · IEEE Trans. Commun. 2018
Wireless sensing and localization
indoor localization
0.312018
Three-Dimensional VLC Positioning Based on Angle Difference of Arrival With Arbitrary Tilting Angle of Receiver · IEEE J. Sel. Areas Commun. 2018
Physical-layer communications › channel estimation
least-squares estimation
0.312018
Three-Dimensional VLC Positioning Based on Angle Difference of Arrival With Arbitrary Tilting Angle of Receiver · IEEE J. Sel. Areas Commun. 2018
Wireless sensing and localization
localization algorithms
0.312018
Three-Dimensional VLC Positioning Based on Angle Difference of Arrival With Arbitrary Tilting Angle of Receiver · IEEE J. Sel. Areas Commun. 2018
Physical-layer communications › fading channels › fading models
lognormal fading
0.312018
A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading Channels · IEEE Trans. Commun. 2018
Physical-layer communications
outage probability
0.312018
A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading Channels · IEEE Trans. Commun. 2018
Physical-layer communications
signal processing for communications
0.312018
Three-Dimensional VLC Positioning Based on Angle Difference of Arrival With Arbitrary Tilting Angle of Receiver · IEEE J. Sel. Areas Commun. 2018
Wireless sensing and localization › indoor localization
visible light positioning
0.312018
Three-Dimensional VLC Positioning Based on Angle Difference of Arrival With Arbitrary Tilting Angle of Receiver · IEEE J. Sel. Areas Commun. 2018
Computer vision › Face, body and person analysis
affective computing
0.332012
Kernel Cross-Modal Factor Analysis for Information Fusion With Application to Bimodal Emotion Recognition · IEEE Trans. Multim. 2012
Recognizing Human Emotional State From Audiovisual Signals · IEEE Trans. Multim. 2008
Recognizing Human Emotional State From Audiovisual Signals · IEEE Trans. Multim. 2008
Content delivery and video streaming
adaptive video streaming
0.312017
Interactive Screen Video Streaming-Based Pervasive Mobile Workstyle · IEEE Trans. Multim. 2017
Cellular and mobile networks
mobile networks
0.312025
Mobile wireless light communication network · Sci. China Inf. Sci. 2025
Audio and music processing › emotion recognition
speech emotion recognition
0.232012
Recognizing Human Emotional State From Audiovisual Signals · IEEE Trans. Multim. 2008
Recognizing Human Emotional State From Audiovisual Signals · IEEE Trans. Multim. 2008
Kernel Cross-Modal Factor Analysis for Information Fusion With Application to Bimodal Emotion Recognition · IEEE Trans. Multim. 2012
Computer vision › Face, body and person analysis › affect recognition
audiovisual emotion recognition
0.222008
Recognizing Human Emotional State From Audiovisual Signals · IEEE Trans. Multim. 2008
Recognizing Human Emotional State From Audiovisual Signals · IEEE Trans. Multim. 2008
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.112012
Kernel Cross-Modal Factor Analysis for Information Fusion With Application to Bimodal Emotion Recognition · IEEE Trans. Multim. 2012

Methods — techniques the papers use, named apart from their topics

screen content coding · 0.6network estimation · 0.6dual decomposition · 0.6probability density function · 0.3multiclassifier ensemble · 0.3moment generating function · 0.3method of exhaustion · 0.3mel-frequency cepstral coefficients · 0.3mahalanobis distance feature selection · 0.3least squares method · 0.3gabor wavelet · 0.3asymptotic analysis · 0.3kernel matrix fusion · 0.3kernel canonical correlation analysis · 0.3hidden markov model · 0.3
YearPublicationVenuePosition
2025 Mobile wireless light communication network
Ziqian Qi, Linning Wang, Yingze Liang, Jiayao Zhou, Pengzhan Liu, Yongjin Wang
Sci. China Inf. Sci.7
2025 Full-duplex ultraviolet light communication network for space-chip interconnection
Ziqian Qi, Mingyuan Xie, Linning Wang, Jiayao Zhou, Xinjie Mo, Jiabin Yan, Yingze Liang, Xianwu Tang, Pengzhan Liu, Jiahao Gou, Yongjin Wang
Sci. China Inf. Sci.12
2018 Asymptotic Outage Probability of Dual-Branch Equal-Gain Combining over Correlated, Non-Identically Distributed Lognormal Fading Channels
abstract
Exact outage probability analysis for equal-gain combining (EGC) over lognormal fading channels results in multi-fold integrals, and traditional asymptotic analysis techniques cannot work for lognormal fading channels due to the channels' infinite diversity order. In this work, we derive a closed-form asymptotic outage probability expression for dual-branch EGC over correlated lognormal fading channels with different statistics. The new result reveals insights into the long-standing problem of asymptotic analysis for correlated lognormal fading channels, paving the way for analysis on more general channel models. The result also provides an efficient way to evaluate the performance of dual-branch EGC over correlated and nonidentically distributed lognormal channels at high signal-to-noise ratio.
Bingcheng Zhu, Julian Cheng 0001, Yongjin Wang, Jin-Yuan Wang
ICC3
2018 Three-Dimensional VLC Positioning Based on Angle Difference of Arrival With Arbitrary Tilting Angle of Receiver
abstract
Existing indoor positioning methods for visible-light communication systems require large database, powerful signal processing units, additional sensors, such as gyroscopes, or the receiver placed toward a certain angle. These assumptions limit the applications of the indoor positioning systems in low-cost scenarios. In this paper, we propose a novel positioning framework based on the angle differences of arrival (ADOA) in 3-D coordinate systems, which can be used in receivers with image sensors or photodiodes. The proposed ADOA positioning does not require the receiver to be placed toward a certain angle and no additional sensor is required. Two positioning algorithms are proposed: one is based on the method of exhaustion (MEX), and the other is based on the least squares method (LSM). The MEX algorithm is analytically proved to be the optimal, while the LSM algorithm has much lower complexity. Both the upper and lower bounds are derived for the average discrepancy between the exact position and the estimated position. These performance bounds can facilitate the design of a light-emitting diode array. Experimental results show that the MEX algorithm can achieve an average error of 3.20 cm with a time cost of 0.36 s, and the LSM algorithm can achieve an average error of 14.66 cm with a time cost of 0.001 s.
Bingcheng Zhu, Julian Cheng 0001, Yongjin Wang, Jun Yan 0006, Jin-Yuan Wang
IEEE J. Sel. Areas Commun.3
2018 A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading Channels
abstract
Prior asymptotic performance analyses are based on the series expansion of the moment-generating function (MGF) or the probability density function (PDF) of channel coefficients. However, these techniques fail for lognormal fading channels because the Taylor series of the PDF of a lognormal random variable is zero at the origin and the MGF does not have an explicit form. Although lognormal fading model has been widely applied in wireless communications and free-space optical communications, few analytical tools are available to provide elegant performance expressions for correlated lognormal channels. In this paper, we propose a novel framework to analyze the asymptotic outage probabilities of selection combining (SC), equal-gain combining (EGC), and maximum-ratio combining (MRC) over equally correlated lognormal fading channels. Based on these closed-form results, we show: 1) the outage probability of EGC or MRC becomes an infinitely small quantity compared to that of SC at high signal-to-noise ratio (SNR); 2) channel correlation can cause an infinite performance loss at high SNR; and 3) negatively correlated lognormal channels can outperform the independent lognormal channels. The analyses reveal insights into the long-standing problem of asymptotic performance analyses over correlated lognormal channels, and circumvent the time-consuming Monte Carlo simulation and numerical integration.
Bingcheng Zhu, Julian Cheng 0001, Jun Yan 0006, Jin-Yuan Wang, Lenan Wu, Yongjin Wang
IEEE Trans. Commun.6
2017 VLC Positioning Using Cameras with Unknown Tilting Angles
abstract
Prior visible-light communication (VLC) positioning systems require the receiver to be placed towards a certain angle, or extra sensors like gyroscopes must be equipped to estimate the angle of the receiver. In this work, we propose a positioning scheme using cameras based on angle difference of arrival (ADOA). ADOA-based positioning does not need the receiver to know its tilting angle, and the position estimation is modeled as a three-variate optimization problem in a finite search region, which is tractable in most smart devices. Our experimental and numerical results show that the ADOA-based VLC positioning scheme achieves accuracy of centimeters and the average cost time is as low as 0.2 seconds.
Bingcheng Zhu, Julian Cheng 0001, Jun Yan 0006, Jin-Yuan Wang, Yongjin Wang
GLOBECOM5
2017 Improvement of BER performance by tilting receiver plane for indoor visible light communications with input-dependent noise
abstract
In this paper, an indoor visible light communication (VLC) system with the input-dependent noise is considered. In the system, the main noise is caused by Gaussian noise, however, with a noise variance depending on the current input signal strength. In the presence of the input-dependent noise, the theoretical expression of the bit error rate (BER) for the VLC using on-off keying is derived. Based on the derived BER, an optimization problem is formulated to improve the BER performance by tilting the receiver plane. The proposed optimization problem is proven to be a convex optimization problem, which can be efficiently solved by using the specialized solver such as the CVX toolbox for MATLAB. To verify the accuracy of the derived expression of the BER, all theoretical results are thoroughly confirmed by using the Monte-Carlo simulations. Moreover, simulation results show that the larger the variance of the input-dependent noise is, the worse the BER performance becomes. Additionally, the BER performance can be dramatically improved by tilting the receiver plane properly.
Jin-Yuan Wang, Jun-Bo Wang 0001, Bingcheng Zhu, Min Lin 0001, Yongpeng Wu 0001, Yongjin Wang, Ming Chen 0001
ICC6
2017 A Practical System Towards the Secure, Robust and Pervasive Mobile Workstyle
abstract
We develop an innovative PC2PC (personal computer to pervasive computing) system to enable the secure, robust and pervasive mobile workstyle. PC2PC server compresses the desktop screens of any virtualized system, and delivers the stream through any popular networks to PC2PC client remotely for stream decoding, rendering and end-user interaction (such as keyboard/mouse commands). We have implemented the overall system from the scratch, where the emerging screen content coding (SCC) extension of the High-Efficiency Video Coding (HEVC) is implemented to compress and stream the desktop screens in real-time, and three core asset channels (i.e., system, display, inputs, etc) are defined to enable systematic end-to-end communication. Compared with the commercial Red Hat SPICE virtual desktop infrastructure (VDI) scheme, our PC2PC could save the network bandwidth by a factor of 2, 7 and 4 respectively for typical video streaming, web browsing and stationary office applications at same visual quality. Meanwhile, we have also measured the delays in the system and presented the preliminary study on the user experience impact. A simple network estimation is applied to optimize the quality-bandwidth adaptation for both single user and multiuser scenarios to combat the network dynamics.
Zhan Ma 0001, Tao Yue 0003, Xun Cao, Yiling Xu, Xin Li 0106, Yongjin Wang
VTC Spring6
2017 Co-Time Co-Frequency Full-Duplex Visible Light On-Chip Communication Using a Pair of InGaN/GaN Quantum-Well Diodes
abstract
Co-time co-frequency full-duplex (CCFD) communication is challenging for radio frequency communications due to the overwhelming self- interference. For visible light communications, it was reported that two light-emitting diodes could construct a half-duplex communication link, but no CCFD communication using two diodes have been realized. In this work, we introduce an on-chip CCFD system using a pair of micrometer-scale InGaN/GaN multiple quantum-well diodes, which can detect and emit light at the same time. Maximum-likelihood estimator is derived and applied to extract the useful signal from the received mixed signal. The proposed CCFD system provides a promising way to decrease the size of on-chip optical communication systems and portable optical communication devices, and it can also reduce the number of optical devices in optical fiber communications and wireless optical communications without introducing much extra cost.
Bingcheng Zhu, Yong-chao Yang, Xuemin Gao, Jia-lei Yuan, Yongjin Wang
VTC Fall6
2017 On-chip optical interconnect using visible light
abstract
We propose and fabricate a monolithic optical interconnect on a GaN-on-silicon platform using a wafer-level technique. Because the InGaN/GaN multiple-quantum-well diodes (MQWDs) can achieve light emission and detection simultaneously, the emitter and collector sharing identical MQW structure are produced using the same process. Suspended waveguides interconnect the emitter with the collector to form in-plane light coupling. Monolithic optical interconnect chip integrates the emitter, waveguide, base, and collector into a multi-component system with a common base. Output states superposition and 1×2 in-plane light communication are experimentally demonstrated. The proposed monolithic optical interconnect opens a promising way toward the diverse applications from in-plane visible light communication to light-induced artificial synaptic devices, intelligent display, on-chip imaging, and optical sensing.
Bingcheng Zhu, Xuemin Gao, Yong-chao Yang, Jia-lei Yuan, Gui-xia Zhu, Yongjin Wang, Peter Grünberg
Frontiers Inf. Technol. Electron. Eng.7
2017 Interactive Screen Video Streaming-Based Pervasive Mobile Workstyle
abstract
In this paper, we develop an interactive screen video streaming-based system to enable the ubiquitous mobile workstyle, which is referred to as personal computer to pervasive computing (PC2PC). The desktop screens of virtualized systems are compressed in the PC2PC servers and delivered to remote end users for stream decoding, rendering, and interactions. We have implemented a system from the scratch, where the emerging screen content coding extension of high-efficiency video coding is implemented to compress and stream the desktop screens of the virtualized system in real time. Three core asset channels, system, display, and inputs, are defined to enable systematic end-to-end communication. Compared with Red Hat SPICE virtual desktop infrastructure scheme, the proposed PC2PC could save network bandwidth consumption by a factor of 2, 7, and 4, respectively, in terms of typical video streaming, web browsing, and stationary office applications at the same visual quality. Meanwhile, we have also measured the delays of the system and presented preliminary results on the user experience aspect. A simple network estimation is applied to optimize the quality bandwidth adaptation for both single user and multiuser scenarios to consider the network dynamics.
Zhan Ma 0001, Tao Yue 0003, Xun Cao, Yiling Xu, Xin Li 0106, Yongjin Wang
IEEE Trans. Multim.6
2016 A Context-Aware Method for Top-k Recommendation in Smart TV
Jun Ma 0001, Yongjin Wang, Lintao Ma, Shanshan Huang 0003
APWeb (2)3
2012 2D-FRFT Based Rotation Invariant Digital Image Watermarking
abstract
The extraction of rotation invariant representation is important for many signal processing problems such as image analysis, computer vision, and pattern recognition. In this paper, we present a systematic analysis of the Two-Dimensional Fractional Fourier Transform (2D-FRFT), and show that under certain conditions, the 2D-FRFT technique possesses the attractive property of rotation invariance. Based on our analysis, we proposed a novel digital image watermarking method which combines 2D chirp signal with the addition and rotation invariant properties of 2D-FRFT to achieve improved robustness and security. The effectiveness of the proposed solution is demonstrated through experiments.
Lei Gao 0001, Lin Qi 0001, Shouyi Yang, Yongjin Wang, Tie Yun, Ling Guan
ISM4
2012 Kernel Cross-Modal Factor Analysis for Information Fusion With Application to Bimodal Emotion Recognition
abstract
In this paper, we investigate kernel based methods for multimodal information analysis and fusion. We introduce a novel approach, kernel cross-modal factor analysis, which identifies the optimal transformations that are capable of representing the coupled patterns between two different subsets of features by minimizing the Frobenius norm in the transformed domain. The kernel trick is utilized for modeling the nonlinear relationship between two multidimensional variables. We examine and compare with kernel canonical correlation analysis which finds projection directions that maximize the correlation between two modalities, and kernel matrix fusion which integrates the kernel matrices of respective modalities through algebraic operations. The performance of the introduced method is evaluated on an audiovisual based bimodal emotion recognition problem. We first perform feature extraction from the audio and visual channels respectively. The presented approaches are then utilized to analyze the cross-modal relationship between audio and visual features. A hidden Markov model is subsequently applied for characterizing the statistical dependence across successive time segments, and identifying the inherent temporal structure of the features in the transformed domain. The effectiveness of the proposed solution is demonstrated through extensive experimentation.
Yongjin Wang, Ling Guan, Anastasios N. Venetsanopoulos
IEEE Trans. Multim.1
2011 Kernel cross-modal factor analysis for multimodal information fusion
abstract
This paper presents a novel approach for multimodal information fusion. The proposed method is based on kernel cross-modal factor analysis (KCFA), in which the optimal transformations that represent the coupled patterns between two different subsets of features are identified by minimizing the Frobenius norm in the transformed domain. It generalizes the linear cross-modal factor analysis (CFA) method via the kernel trick to model the nonlinear relationship between two multidimensional variables. The effectiveness of the introduced solution is demonstrated through experimentation on an audiovisual based emotion recognition problem. Experimental results show that the proposed approach outperforms the concatenation based feature level fusion, the linear CFA, as well as the canonical correlation analysis (CCA) and kernel CCA methods.
Yongjin Wang, Ling Guan, Anastasios N. Venetsanopoulos
ICASSP1
2011 Audiovisual emotion recognition via cross-modal association in kernel space
abstract
In this paper, we introduce a new method for audiovisual based multimodal emotion recognition. The proposed method identifies the optimal transformations that are capable of representing the coupled patterns between audio and visual information through cross-modal association. Specifically, kernel machine technique is utilized for capturing the nonlinear relationship between two different subsets of features. A hidden Markov model is subsequently applied for characterizing the statistical dependence across successive time segments, and identifying the inherent temporal structure of the features in the transformed domain. Information fusion at the feature and score levels are examined and compared. The effectiveness of the introduced solution is demonstrated through extensive experimentation.
Yongjin Wang, Ling Guan, Anastasios N. Venetsanopoulos
ICME1
2011 On Random Transformations for Changeable Face Verification
abstract
The generation of changeable and privacy-preserving biometric templates is important for the pervasive deployment of biometric technology in a wide variety of applications. This paper presents a systematic analysis of random transformation-based methods for addressing the changeability and privacy problems in biometrics-based verification systems. The proposed methods transform the original biometric feature vectors using random transformations, and the sorted index numbers (SIN) of the resulting vectors in the transformed domain are stored as the biometric templates. Three types of random transformations, namely, random additive transform, random multiplicative transform, and random projection, are discussed and analyzed. The random transformations, in combination with the SIN approach, constitute repeatable and noninvertible transformations; hence, the generated templates are changeable and provide privacy protection. The effectiveness of the proposed methods is well supported by both detailed analysis and extensive experimentation on a face verification problem.
Yongjin Wang, Dimitrios Hatzinakos
IEEE Trans. Syst. Man Cybern. Part B1
2010 Cancelable Face Recognition Using Random Multiplicative Transform
abstract
The generation of cancelable and privacy preserving biometric templates is important for the pervasive deployment of biometric technology in a wide variety of applications. This paper presents a novel approach for cancelable biometric authentication using random multiplicative transform. The proposed method transforms the original biometric feature vector through element-wise multiplication with a random vector, and the sorted index numbers of the resulting vector in the transformed domain are stored as the biometric template. The changeability and privacy protecting properties of the generated biometric template are analyzed in detail. The effectiveness of the proposed method is well supported by extensive experimentation on a face verification problem.
Yongjin Wang, Dimitrios Hatzinakos
ICPR1
2010 An Analysis of Random Projection for Changeable and Privacy-Preserving Biometric Verification
abstract
Changeability and privacy protection are important factors for widespread deployment of biometrics-based verification systems. This paper presents a systematic analysis of a random-projection (RP)-based method for addressing these problems. The employed method transforms biometric data using a random matrix with each entry an independent and identically distributed Gaussian random variable. The similarity- and privacy-preserving properties, as well as the changeability of the biometric information in the transformed domain, are analyzed in detail. Specifically, RP on both high-dimensional image vectors and dimensionality-reduced feature vectors is discussed and compared. A vector translation method is proposed to improve the changeability of the generated templates. The feasibility of the introduced solution is well supported by detailed theoretical analyses. Extensive experimentation on a face-based biometric verification problem shows the effectiveness of the proposed method.
Yongjin Wang, Konstantinos N. Plataniotis
IEEE Trans. Syst. Man Cybern. Part B1
2009 Face recognition with enhanced privacy protection
abstract
This paper presents a novel approach for face based biometric recognition. The proposed method is based on the sorted index numbers (SIN) of appearance based facial features. A new algorithm is introduced to measure the similarity between SIN vectors. Due to the non-invertibility of the transformation from the original features to the SIN vectors, the proposed method can preserve the privacy of the users. The effectiveness of the proposed method is tested on a large generic data set, which contains images from several well known face databases. Experimental results demonstrate that the proposed solution may improve the recognition accuracy in both identification and verification scenarios.
Yongjin Wang, Dimitrios Hatzinakos
ICASSP1
2009 Multimedia multimodal methodologies
abstract
This paper outlines several multimedia systems that utilize a multimodal approach. These systems include audiovisual based emotion recognition, image and video retrieval, and face and head tracking. Data collected from diverse sources/sensors are employed to improve the accuracy of correctly detecting, classifying, identifying, and tracking of a desired object or target. It is shown that the integration of multimodality data will be more efficient and potentially more accurate than if the data was acquired from a single source. A number of cutting-edge applications for multimodal systems will be discussed. An advanced assistance robot using the multimodal systems will be presented.
Ling Guan, Paisarn Muneesawang, Yongjin Wang, Rui Zhang 0010, Tie Yun, Adrian Bulzacki, Muhammad Talal Ibrahim
ICME3
2009 Toward natural and efficient human computer interaction
abstract
Effective detection, recognition, interpretation, and analysis of human physiological and behavioral characteristics are of fundamental importance in the design and development of intelligent human computer interaction (HCI) systems. This paper illustrates the issues and challenges in the design of such systems through two real examples, emotion recognition and face detection. In particular, we focus on audiovisual based bimodal emotion recognition, face detection in crowded scene, and facial fiducial points detection. The integration of these systems is expected to produce more robust and stronger performance, and provide more natural and friendly man-machine interaction.
Ling Guan, Yongjin Wang, Tie Yun
ICME2
2008 Recognizing Human Emotional State From Audiovisual Signals
abstract
Machine recognition of human emotional state is an important component for efficient human-computer interaction. The majority of existing works address this problem by utilizing audio signals alone, or visual information only. In this paper, we explore a systematic approach for recognition of human emotional state from audiovisual signals. The audio characteristics of emotional speech are represented by the extracted prosodic, Mel-frequency Cepstral Coefficient (MFCC), and formant frequency features. A face detection scheme based on HSV color model is used to detect the face from the background. The visual information is represented by Gabor wavelet features. We perform feature selection by using a stepwise method based on Mahalanobis distance. The selected audiovisual features are used to classify the data into their corresponding emotions. Based on a comparative study of different classification algorithms and specific characteristics of individual emotion, a novel multiclassifier scheme is proposed to boost the recognition performance. The feasibility of the proposed system is tested over a database that incorporates human subjects from different languages and cultural backgrounds. Experimental results demonstrate the effectiveness of the proposed system. The multiclassifier scheme achieves the best overall recognition rate of 82.14%.
Yongjin Wang, Ling Guan
IEEE Trans. Multim.1
2008 Recognizing Human Emotional State From Audiovisual Signals
abstract
Machine recognition of human emotional state is an important component for efficient human-computer interaction. The majority of existing works address this problem by utilizing audio signals alone, or visual information only. In this paper, we explore a systematic approach for recognition of human emotional state from audiovisual signals. The audio characteristics of emotional speech are represented by the extracted prosodic, Mel-frequency Cepstral Coefficient (MFCC), and formant frequency features. A face detection scheme based on HSV color model is used to detect the face from the background. The visual information is represented by Gabor wavelet features. We perform feature selection by using a stepwise method based on Mahalanobis distance. The selected audiovisual features are used to classify the data into their corresponding emotions. Based on a comparative study of different classification algorithms and specific characteristics of individual emotion, a novel multiclassifier scheme is proposed to boost the recognition performance. The feasibility of the proposed system is tested over a database that incorporates human subjects from different languages and cultural backgrounds. Experimental results demonstrate the effectiveness of the proposed system. The multiclassifier scheme achieves the best overall recognition rate of 82.14%.
Yongjin Wang, Ling Guan
IEEE Trans. Multim.1
2006 Special Effects in Film/Video Making: A New Media Initiative Project
abstract
We present a system and a set of tools for producing special effects in film/video making by applying image processing and human centered computing techniques. A combination of shot detection, object segmentation, background generation, and image warping techniques are used. The user selects the image frame or the object of interest, and the image warp transformation to be used from the GUI. Wrapping can be performed either on a whole image sequence or on an object of interest in the sequence. In the latter, the object is first segmented and its motion tracked. Object segmentation is achieved either by snakes or graph cuts. Steerable pyramid background generation is then used to fill in the portion cut from the foreground
Chun-Hao Wang, Yongjin Wang, Meifeng Lian, Bruce Elder, Xiaoou Tang, Ling Guan
ICME2
2005 Recognizing human emotion from audiovisual information
abstract
In this paper, we present an emotion recognition system to classify human emotional state from audiovisual signals. We extract prosodic, mel-frequency cepstral coefficient (MFCC), and formant frequency features to represent the audio characteristics of the emotional speech. A face detection scheme, based on the HSV color model, is used to detect the face from the background. The facial expressions are represented by Gabor wavelet features. We perform feature selection by using a stepwise method based on Mahalanobis distance. A classification scheme involving the analysis of individual class and combinations of different classes is proposed. Our emotion recognition system is tested over a language and race independent database, and an overall recognition accuracy of 82.14% is achieved.
Yongjin Wang, Ling Guan
ICASSP (2)1
2004 An investigation of speech-based human emotion recognition
abstract
This paper presents our recent work on recognizing human emotion from the speech signal. The proposed recognition system was tested over a language, speaker, and context independent emotional speech database. Prosodic, Mel-frequency cepstral coefficient (MFCC), and formant frequency features are extracted from the speech utterances. We perform feature selection by using the stepwise method based on Mahalanobis distance. The selected features are used to classify the speeches into their corresponding emotional classes. Different classification algorithms including maximum likelihood classifier (MLC), Gaussian mixture model (GMM), neural network (NN), K-nearest neighbors (K-NN), and Fisher's linear discriminant analysis (FLDA) are compared in this study. The recognition results show that FLDA gives the best recognition accuracy by using the selected features.
Yongjin Wang, Ling Guan
MMSP1