EDBT 2026 Demo / reviewers in the wild / expert
Yongjin Wang
dblp:65/5570
· DBLP profile ↗
27ranked-venue papers
11as first author
2since 2021 · last 2025
0000-0001-8109-4640ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 first-authorComputer networks · 5Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
5 papers |
Physical-layer communications · 72% Wireless sensing and localization · 14% Edge and fog computing · 4% | |
| Artificial intelligence
3 papers |
Face, body and person analysis · 77% Representation and self-supervised learning · 23% |
Topics — the 20 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Physical-layer communications
full-duplex communication |
0.9 | 1 | 2025 | Full-duplex ultraviolet light communication network for space-chip interconnection · Sci. China Inf. Sci. 2025 |
Physical-layer communications
optical wireless communication |
0.9 | 1 | 2025 | Mobile wireless light communication network · Sci. China Inf. Sci. 2025 |
Physical-layer communications › optical wireless communication
ultraviolet communication |
0.9 | 1 | 2025 | Full-duplex ultraviolet light communication network for space-chip interconnection · Sci. China Inf. Sci. 2025 |
Physical-layer communications › diversity combining
combining schemes |
0.3 | 1 | 2018 | A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading Channels · IEEE Trans. Commun. 2018 |
Physical-layer communications › fading channels › correlated fading
correlated lognormal fading |
0.3 | 1 | 2018 | A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading Channels · IEEE Trans. Commun. 2018 |
Physical-layer communications
diversity combining |
0.3 | 1 | 2018 | A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading Channels · IEEE Trans. Commun. 2018 |
Physical-layer communications
fading channels |
0.3 | 1 | 2018 | A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading Channels · IEEE Trans. Commun. 2018 |
Wireless sensing and localization
indoor localization |
0.3 | 1 | 2018 | Three-Dimensional VLC Positioning Based on Angle Difference of Arrival With Arbitrary Tilting Angle of Receiver · IEEE J. Sel. Areas Commun. 2018 |
Physical-layer communications › channel estimation
least-squares estimation |
0.3 | 1 | 2018 | Three-Dimensional VLC Positioning Based on Angle Difference of Arrival With Arbitrary Tilting Angle of Receiver · IEEE J. Sel. Areas Commun. 2018 |
Wireless sensing and localization
localization algorithms |
0.3 | 1 | 2018 | Three-Dimensional VLC Positioning Based on Angle Difference of Arrival With Arbitrary Tilting Angle of Receiver · IEEE J. Sel. Areas Commun. 2018 |
Physical-layer communications › fading channels › fading models
lognormal fading |
0.3 | 1 | 2018 | A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading Channels · IEEE Trans. Commun. 2018 |
Physical-layer communications
outage probability |
0.3 | 1 | 2018 | A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading Channels · IEEE Trans. Commun. 2018 |
Physical-layer communications
signal processing for communications |
0.3 | 1 | 2018 | Three-Dimensional VLC Positioning Based on Angle Difference of Arrival With Arbitrary Tilting Angle of Receiver · IEEE J. Sel. Areas Commun. 2018 |
Wireless sensing and localization › indoor localization
visible light positioning |
0.3 | 1 | 2018 | Three-Dimensional VLC Positioning Based on Angle Difference of Arrival With Arbitrary Tilting Angle of Receiver · IEEE J. Sel. Areas Commun. 2018 |
Computer vision › Face, body and person analysis
affective computing |
0.3 | 3 | 2012 | Kernel Cross-Modal Factor Analysis for Information Fusion With Application to Bimodal Emotion Recognition · IEEE Trans. Multim. 2012 Recognizing Human Emotional State From Audiovisual Signals · IEEE Trans. Multim. 2008 Recognizing Human Emotional State From Audiovisual Signals · IEEE Trans. Multim. 2008 |
Content delivery and video streaming
adaptive video streaming |
0.3 | 1 | 2017 | Interactive Screen Video Streaming-Based Pervasive Mobile Workstyle · IEEE Trans. Multim. 2017 |
Cellular and mobile networks
mobile networks |
0.3 | 1 | 2025 | Mobile wireless light communication network · Sci. China Inf. Sci. 2025 |
Audio and music processing › emotion recognition
speech emotion recognition |
0.2 | 3 | 2012 | Recognizing Human Emotional State From Audiovisual Signals · IEEE Trans. Multim. 2008 Recognizing Human Emotional State From Audiovisual Signals · IEEE Trans. Multim. 2008 Kernel Cross-Modal Factor Analysis for Information Fusion With Application to Bimodal Emotion Recognition · IEEE Trans. Multim. 2012 |
Computer vision › Face, body and person analysis › affect recognition
audiovisual emotion recognition |
0.2 | 2 | 2008 | Recognizing Human Emotional State From Audiovisual Signals · IEEE Trans. Multim. 2008 Recognizing Human Emotional State From Audiovisual Signals · IEEE Trans. Multim. 2008 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.1 | 1 | 2012 | Kernel Cross-Modal Factor Analysis for Information Fusion With Application to Bimodal Emotion Recognition · IEEE Trans. Multim. 2012 |
Methods — techniques the papers use, named apart from their topics
screen content coding · 0.6network estimation · 0.6dual decomposition · 0.6probability density function · 0.3multiclassifier ensemble · 0.3moment generating function · 0.3method of exhaustion · 0.3mel-frequency cepstral coefficients · 0.3mahalanobis distance feature selection · 0.3least squares method · 0.3gabor wavelet · 0.3asymptotic analysis · 0.3kernel matrix fusion · 0.3kernel canonical correlation analysis · 0.3hidden markov model · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mobile wireless light communication network
Ziqian Qi, Linning Wang, Yingze Liang, Jiayao Zhou, Pengzhan Liu, Yongjin Wang |
Sci. China Inf. Sci. | 7 |
| 2025 | Full-duplex ultraviolet light communication network for space-chip interconnection
Ziqian Qi, Mingyuan Xie, Linning Wang, Jiayao Zhou, Xinjie Mo, Jiabin Yan, Yingze Liang, Xianwu Tang, Pengzhan Liu, Jiahao Gou, Yongjin Wang |
Sci. China Inf. Sci. | 12 |
| 2018 | Asymptotic Outage Probability of Dual-Branch Equal-Gain Combining over Correlated, Non-Identically Distributed Lognormal Fading ChannelsabstractExact outage probability analysis for equal-gain combining (EGC) over lognormal fading channels results in multi-fold integrals, and traditional asymptotic analysis techniques cannot work for lognormal fading channels due to the channels' infinite diversity order. In this work, we derive a closed-form asymptotic outage probability expression for dual-branch EGC over correlated lognormal fading channels with different statistics. The new result reveals insights into the long-standing problem of asymptotic analysis for correlated lognormal fading channels, paving the way for analysis on more general channel models. The result also provides an efficient way to evaluate the performance of dual-branch EGC over correlated and nonidentically distributed lognormal channels at high signal-to-noise ratio. Bingcheng Zhu, Julian Cheng 0001, Yongjin Wang, Jin-Yuan Wang |
ICC | 3 |
| 2018 | Three-Dimensional VLC Positioning Based on Angle Difference of Arrival With Arbitrary Tilting Angle of ReceiverabstractExisting indoor positioning methods for visible-light communication systems require large database, powerful signal processing units, additional sensors, such as gyroscopes, or the receiver placed toward a certain angle. These assumptions limit the applications of the indoor positioning systems in low-cost scenarios. In this paper, we propose a novel positioning framework based on the angle differences of arrival (ADOA) in 3-D coordinate systems, which can be used in receivers with image sensors or photodiodes. The proposed ADOA positioning does not require the receiver to be placed toward a certain angle and no additional sensor is required. Two positioning algorithms are proposed: one is based on the method of exhaustion (MEX), and the other is based on the least squares method (LSM). The MEX algorithm is analytically proved to be the optimal, while the LSM algorithm has much lower complexity. Both the upper and lower bounds are derived for the average discrepancy between the exact position and the estimated position. These performance bounds can facilitate the design of a light-emitting diode array. Experimental results show that the MEX algorithm can achieve an average error of 3.20 cm with a time cost of 0.36 s, and the LSM algorithm can achieve an average error of 14.66 cm with a time cost of 0.001 s. Bingcheng Zhu, Julian Cheng 0001, Yongjin Wang, Jun Yan 0006, Jin-Yuan Wang |
IEEE J. Sel. Areas Commun. | 3 |
| 2018 | A New Asymptotic Analysis Technique for Diversity Receptions Over Correlated Lognormal Fading ChannelsabstractPrior asymptotic performance analyses are based on the series expansion of the moment-generating function (MGF) or the probability density function (PDF) of channel coefficients. However, these techniques fail for lognormal fading channels because the Taylor series of the PDF of a lognormal random variable is zero at the origin and the MGF does not have an explicit form. Although lognormal fading model has been widely applied in wireless communications and free-space optical communications, few analytical tools are available to provide elegant performance expressions for correlated lognormal channels. In this paper, we propose a novel framework to analyze the asymptotic outage probabilities of selection combining (SC), equal-gain combining (EGC), and maximum-ratio combining (MRC) over equally correlated lognormal fading channels. Based on these closed-form results, we show: 1) the outage probability of EGC or MRC becomes an infinitely small quantity compared to that of SC at high signal-to-noise ratio (SNR); 2) channel correlation can cause an infinite performance loss at high SNR; and 3) negatively correlated lognormal channels can outperform the independent lognormal channels. The analyses reveal insights into the long-standing problem of asymptotic performance analyses over correlated lognormal channels, and circumvent the time-consuming Monte Carlo simulation and numerical integration. Bingcheng Zhu, Julian Cheng 0001, Jun Yan 0006, Jin-Yuan Wang, Lenan Wu, Yongjin Wang |
IEEE Trans. Commun. | 6 |
| 2017 | VLC Positioning Using Cameras with Unknown Tilting AnglesabstractPrior visible-light communication (VLC) positioning systems require the receiver to be placed towards a certain angle, or extra sensors like gyroscopes must be equipped to estimate the angle of the receiver. In this work, we propose a positioning scheme using cameras based on angle difference of arrival (ADOA). ADOA-based positioning does not need the receiver to know its tilting angle, and the position estimation is modeled as a three-variate optimization problem in a finite search region, which is tractable in most smart devices. Our experimental and numerical results show that the ADOA-based VLC positioning scheme achieves accuracy of centimeters and the average cost time is as low as 0.2 seconds. Bingcheng Zhu, Julian Cheng 0001, Jun Yan 0006, Jin-Yuan Wang, Yongjin Wang |
GLOBECOM | 5 |
| 2017 | Improvement of BER performance by tilting receiver plane for indoor visible light communications with input-dependent noiseabstractIn this paper, an indoor visible light communication (VLC) system with the input-dependent noise is considered. In the system, the main noise is caused by Gaussian noise, however, with a noise variance depending on the current input signal strength. In the presence of the input-dependent noise, the theoretical expression of the bit error rate (BER) for the VLC using on-off keying is derived. Based on the derived BER, an optimization problem is formulated to improve the BER performance by tilting the receiver plane. The proposed optimization problem is proven to be a convex optimization problem, which can be efficiently solved by using the specialized solver such as the CVX toolbox for MATLAB. To verify the accuracy of the derived expression of the BER, all theoretical results are thoroughly confirmed by using the Monte-Carlo simulations. Moreover, simulation results show that the larger the variance of the input-dependent noise is, the worse the BER performance becomes. Additionally, the BER performance can be dramatically improved by tilting the receiver plane properly. Jin-Yuan Wang, Jun-Bo Wang 0001, Bingcheng Zhu, Min Lin 0001, Yongpeng Wu 0001, Yongjin Wang, Ming Chen 0001 |
ICC | 6 |
| 2017 | A Practical System Towards the Secure, Robust and Pervasive Mobile WorkstyleabstractWe develop an innovative PC2PC (personal computer to pervasive computing) system to enable the secure, robust and pervasive mobile workstyle. PC2PC server compresses the desktop screens of any virtualized system, and delivers the stream through any popular networks to PC2PC client remotely for stream decoding, rendering and end-user interaction (such as keyboard/mouse commands). We have implemented the overall system from the scratch, where the emerging screen content coding (SCC) extension of the High-Efficiency Video Coding (HEVC) is implemented to compress and stream the desktop screens in real-time, and three core asset channels (i.e., system, display, inputs, etc) are defined to enable systematic end-to-end communication. Compared with the commercial Red Hat SPICE virtual desktop infrastructure (VDI) scheme, our PC2PC could save the network bandwidth by a factor of 2, 7 and 4 respectively for typical video streaming, web browsing and stationary office applications at same visual quality. Meanwhile, we have also measured the delays in the system and presented the preliminary study on the user experience impact. A simple network estimation is applied to optimize the quality-bandwidth adaptation for both single user and multiuser scenarios to combat the network dynamics. Zhan Ma 0001, Tao Yue 0003, Xun Cao, Yiling Xu, Xin Li 0106, Yongjin Wang |
VTC Spring | 6 |
| 2017 | Co-Time Co-Frequency Full-Duplex Visible Light On-Chip Communication Using a Pair of InGaN/GaN Quantum-Well DiodesabstractCo-time co-frequency full-duplex (CCFD) communication is challenging for radio frequency communications due to the overwhelming self- interference. For visible light communications, it was reported that two light-emitting diodes could construct a half-duplex communication link, but no CCFD communication using two diodes have been realized. In this work, we introduce an on-chip CCFD system using a pair of micrometer-scale InGaN/GaN multiple quantum-well diodes, which can detect and emit light at the same time. Maximum-likelihood estimator is derived and applied to extract the useful signal from the received mixed signal. The proposed CCFD system provides a promising way to decrease the size of on-chip optical communication systems and portable optical communication devices, and it can also reduce the number of optical devices in optical fiber communications and wireless optical communications without introducing much extra cost. Bingcheng Zhu, Yong-chao Yang, Xuemin Gao, Jia-lei Yuan, Yongjin Wang |
VTC Fall | 6 |
| 2017 | On-chip optical interconnect using visible lightabstractWe propose and fabricate a monolithic optical interconnect on a GaN-on-silicon platform using a wafer-level technique. Because the InGaN/GaN multiple-quantum-well diodes (MQWDs) can achieve light emission and detection simultaneously, the emitter and collector sharing identical MQW structure are produced using the same process. Suspended waveguides interconnect the emitter with the collector to form in-plane light coupling. Monolithic optical interconnect chip integrates the emitter, waveguide, base, and collector into a multi-component system with a common base. Output states superposition and 1×2 in-plane light communication are experimentally demonstrated. The proposed monolithic optical interconnect opens a promising way toward the diverse applications from in-plane visible light communication to light-induced artificial synaptic devices, intelligent display, on-chip imaging, and optical sensing. Bingcheng Zhu, Xuemin Gao, Yong-chao Yang, Jia-lei Yuan, Gui-xia Zhu, Yongjin Wang, Peter Grünberg |
Frontiers Inf. Technol. Electron. Eng. | 7 |
| 2017 | Interactive Screen Video Streaming-Based Pervasive Mobile WorkstyleabstractIn this paper, we develop an interactive screen video streaming-based system to enable the ubiquitous mobile workstyle, which is referred to as personal computer to pervasive computing (PC2PC). The desktop screens of virtualized systems are compressed in the PC2PC servers and delivered to remote end users for stream decoding, rendering, and interactions. We have implemented a system from the scratch, where the emerging screen content coding extension of high-efficiency video coding is implemented to compress and stream the desktop screens of the virtualized system in real time. Three core asset channels, system, display, and inputs, are defined to enable systematic end-to-end communication. Compared with Red Hat SPICE virtual desktop infrastructure scheme, the proposed PC2PC could save network bandwidth consumption by a factor of 2, 7, and 4, respectively, in terms of typical video streaming, web browsing, and stationary office applications at the same visual quality. Meanwhile, we have also measured the delays of the system and presented preliminary results on the user experience aspect. A simple network estimation is applied to optimize the quality bandwidth adaptation for both single user and multiuser scenarios to consider the network dynamics. Zhan Ma 0001, Tao Yue 0003, Xun Cao, Yiling Xu, Xin Li 0106, Yongjin Wang |
IEEE Trans. Multim. | 6 |
| 2016 | A Context-Aware Method for Top-k Recommendation in Smart TV
Jun Ma 0001, Yongjin Wang, Lintao Ma, Shanshan Huang 0003 |
APWeb (2) | 3 |
| 2012 | 2D-FRFT Based Rotation Invariant Digital Image WatermarkingabstractThe extraction of rotation invariant representation is important for many signal processing problems such as image analysis, computer vision, and pattern recognition. In this paper, we present a systematic analysis of the Two-Dimensional Fractional Fourier Transform (2D-FRFT), and show that under certain conditions, the 2D-FRFT technique possesses the attractive property of rotation invariance. Based on our analysis, we proposed a novel digital image watermarking method which combines 2D chirp signal with the addition and rotation invariant properties of 2D-FRFT to achieve improved robustness and security. The effectiveness of the proposed solution is demonstrated through experiments. Lei Gao 0001, Lin Qi 0001, Shouyi Yang, Yongjin Wang, Tie Yun, Ling Guan |
ISM | 4 |
| 2012 | Kernel Cross-Modal Factor Analysis for Information Fusion With Application to Bimodal Emotion RecognitionabstractIn this paper, we investigate kernel based methods for multimodal information analysis and fusion. We introduce a novel approach, kernel cross-modal factor analysis, which identifies the optimal transformations that are capable of representing the coupled patterns between two different subsets of features by minimizing the Frobenius norm in the transformed domain. The kernel trick is utilized for modeling the nonlinear relationship between two multidimensional variables. We examine and compare with kernel canonical correlation analysis which finds projection directions that maximize the correlation between two modalities, and kernel matrix fusion which integrates the kernel matrices of respective modalities through algebraic operations. The performance of the introduced method is evaluated on an audiovisual based bimodal emotion recognition problem. We first perform feature extraction from the audio and visual channels respectively. The presented approaches are then utilized to analyze the cross-modal relationship between audio and visual features. A hidden Markov model is subsequently applied for characterizing the statistical dependence across successive time segments, and identifying the inherent temporal structure of the features in the transformed domain. The effectiveness of the proposed solution is demonstrated through extensive experimentation. Yongjin Wang, Ling Guan, Anastasios N. Venetsanopoulos |
IEEE Trans. Multim. | 1 |
| 2011 | Kernel cross-modal factor analysis for multimodal information fusionabstractThis paper presents a novel approach for multimodal information fusion. The proposed method is based on kernel cross-modal factor analysis (KCFA), in which the optimal transformations that represent the coupled patterns between two different subsets of features are identified by minimizing the Frobenius norm in the transformed domain. It generalizes the linear cross-modal factor analysis (CFA) method via the kernel trick to model the nonlinear relationship between two multidimensional variables. The effectiveness of the introduced solution is demonstrated through experimentation on an audiovisual based emotion recognition problem. Experimental results show that the proposed approach outperforms the concatenation based feature level fusion, the linear CFA, as well as the canonical correlation analysis (CCA) and kernel CCA methods. Yongjin Wang, Ling Guan, Anastasios N. Venetsanopoulos |
ICASSP | 1 |
| 2011 | Audiovisual emotion recognition via cross-modal association in kernel spaceabstractIn this paper, we introduce a new method for audiovisual based multimodal emotion recognition. The proposed method identifies the optimal transformations that are capable of representing the coupled patterns between audio and visual information through cross-modal association. Specifically, kernel machine technique is utilized for capturing the nonlinear relationship between two different subsets of features. A hidden Markov model is subsequently applied for characterizing the statistical dependence across successive time segments, and identifying the inherent temporal structure of the features in the transformed domain. Information fusion at the feature and score levels are examined and compared. The effectiveness of the introduced solution is demonstrated through extensive experimentation. Yongjin Wang, Ling Guan, Anastasios N. Venetsanopoulos |
ICME | 1 |
| 2011 | On Random Transformations for Changeable Face VerificationabstractThe generation of changeable and privacy-preserving biometric templates is important for the pervasive deployment of biometric technology in a wide variety of applications. This paper presents a systematic analysis of random transformation-based methods for addressing the changeability and privacy problems in biometrics-based verification systems. The proposed methods transform the original biometric feature vectors using random transformations, and the sorted index numbers (SIN) of the resulting vectors in the transformed domain are stored as the biometric templates. Three types of random transformations, namely, random additive transform, random multiplicative transform, and random projection, are discussed and analyzed. The random transformations, in combination with the SIN approach, constitute repeatable and noninvertible transformations; hence, the generated templates are changeable and provide privacy protection. The effectiveness of the proposed methods is well supported by both detailed analysis and extensive experimentation on a face verification problem. Yongjin Wang, Dimitrios Hatzinakos |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2010 | Cancelable Face Recognition Using Random Multiplicative TransformabstractThe generation of cancelable and privacy preserving biometric templates is important for the pervasive deployment of biometric technology in a wide variety of applications. This paper presents a novel approach for cancelable biometric authentication using random multiplicative transform. The proposed method transforms the original biometric feature vector through element-wise multiplication with a random vector, and the sorted index numbers of the resulting vector in the transformed domain are stored as the biometric template. The changeability and privacy protecting properties of the generated biometric template are analyzed in detail. The effectiveness of the proposed method is well supported by extensive experimentation on a face verification problem. Yongjin Wang, Dimitrios Hatzinakos |
ICPR | 1 |
| 2010 | An Analysis of Random Projection for Changeable and Privacy-Preserving Biometric VerificationabstractChangeability and privacy protection are important factors for widespread deployment of biometrics-based verification systems. This paper presents a systematic analysis of a random-projection (RP)-based method for addressing these problems. The employed method transforms biometric data using a random matrix with each entry an independent and identically distributed Gaussian random variable. The similarity- and privacy-preserving properties, as well as the changeability of the biometric information in the transformed domain, are analyzed in detail. Specifically, RP on both high-dimensional image vectors and dimensionality-reduced feature vectors is discussed and compared. A vector translation method is proposed to improve the changeability of the generated templates. The feasibility of the introduced solution is well supported by detailed theoretical analyses. Extensive experimentation on a face-based biometric verification problem shows the effectiveness of the proposed method. Yongjin Wang, Konstantinos N. Plataniotis |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2009 | Face recognition with enhanced privacy protectionabstractThis paper presents a novel approach for face based biometric recognition. The proposed method is based on the sorted index numbers (SIN) of appearance based facial features. A new algorithm is introduced to measure the similarity between SIN vectors. Due to the non-invertibility of the transformation from the original features to the SIN vectors, the proposed method can preserve the privacy of the users. The effectiveness of the proposed method is tested on a large generic data set, which contains images from several well known face databases. Experimental results demonstrate that the proposed solution may improve the recognition accuracy in both identification and verification scenarios. Yongjin Wang, Dimitrios Hatzinakos |
ICASSP | 1 |
| 2009 | Multimedia multimodal methodologiesabstractThis paper outlines several multimedia systems that utilize a multimodal approach. These systems include audiovisual based emotion recognition, image and video retrieval, and face and head tracking. Data collected from diverse sources/sensors are employed to improve the accuracy of correctly detecting, classifying, identifying, and tracking of a desired object or target. It is shown that the integration of multimodality data will be more efficient and potentially more accurate than if the data was acquired from a single source. A number of cutting-edge applications for multimodal systems will be discussed. An advanced assistance robot using the multimodal systems will be presented. Ling Guan, Paisarn Muneesawang, Yongjin Wang, Rui Zhang 0010, Tie Yun, Adrian Bulzacki, Muhammad Talal Ibrahim |
ICME | 3 |
| 2009 | Toward natural and efficient human computer interactionabstractEffective detection, recognition, interpretation, and analysis of human physiological and behavioral characteristics are of fundamental importance in the design and development of intelligent human computer interaction (HCI) systems. This paper illustrates the issues and challenges in the design of such systems through two real examples, emotion recognition and face detection. In particular, we focus on audiovisual based bimodal emotion recognition, face detection in crowded scene, and facial fiducial points detection. The integration of these systems is expected to produce more robust and stronger performance, and provide more natural and friendly man-machine interaction. Ling Guan, Yongjin Wang, Tie Yun |
ICME | 2 |
| 2008 | Recognizing Human Emotional State From Audiovisual SignalsabstractMachine recognition of human emotional state is an important component for efficient human-computer interaction. The majority of existing works address this problem by utilizing audio signals alone, or visual information only. In this paper, we explore a systematic approach for recognition of human emotional state from audiovisual signals. The audio characteristics of emotional speech are represented by the extracted prosodic, Mel-frequency Cepstral Coefficient (MFCC), and formant frequency features. A face detection scheme based on HSV color model is used to detect the face from the background. The visual information is represented by Gabor wavelet features. We perform feature selection by using a stepwise method based on Mahalanobis distance. The selected audiovisual features are used to classify the data into their corresponding emotions. Based on a comparative study of different classification algorithms and specific characteristics of individual emotion, a novel multiclassifier scheme is proposed to boost the recognition performance. The feasibility of the proposed system is tested over a database that incorporates human subjects from different languages and cultural backgrounds. Experimental results demonstrate the effectiveness of the proposed system. The multiclassifier scheme achieves the best overall recognition rate of 82.14%. Yongjin Wang, Ling Guan |
IEEE Trans. Multim. | 1 |
| 2008 | Recognizing Human Emotional State From Audiovisual SignalsabstractMachine recognition of human emotional state is an important component for efficient human-computer interaction. The majority of existing works address this problem by utilizing audio signals alone, or visual information only. In this paper, we explore a systematic approach for recognition of human emotional state from audiovisual signals. The audio characteristics of emotional speech are represented by the extracted prosodic, Mel-frequency Cepstral Coefficient (MFCC), and formant frequency features. A face detection scheme based on HSV color model is used to detect the face from the background. The visual information is represented by Gabor wavelet features. We perform feature selection by using a stepwise method based on Mahalanobis distance. The selected audiovisual features are used to classify the data into their corresponding emotions. Based on a comparative study of different classification algorithms and specific characteristics of individual emotion, a novel multiclassifier scheme is proposed to boost the recognition performance. The feasibility of the proposed system is tested over a database that incorporates human subjects from different languages and cultural backgrounds. Experimental results demonstrate the effectiveness of the proposed system. The multiclassifier scheme achieves the best overall recognition rate of 82.14%. Yongjin Wang, Ling Guan |
IEEE Trans. Multim. | 1 |
| 2006 | Special Effects in Film/Video Making: A New Media Initiative ProjectabstractWe present a system and a set of tools for producing special effects in film/video making by applying image processing and human centered computing techniques. A combination of shot detection, object segmentation, background generation, and image warping techniques are used. The user selects the image frame or the object of interest, and the image warp transformation to be used from the GUI. Wrapping can be performed either on a whole image sequence or on an object of interest in the sequence. In the latter, the object is first segmented and its motion tracked. Object segmentation is achieved either by snakes or graph cuts. Steerable pyramid background generation is then used to fill in the portion cut from the foreground Chun-Hao Wang, Yongjin Wang, Meifeng Lian, Bruce Elder, Xiaoou Tang, Ling Guan |
ICME | 2 |
| 2005 | Recognizing human emotion from audiovisual informationabstractIn this paper, we present an emotion recognition system to classify human emotional state from audiovisual signals. We extract prosodic, mel-frequency cepstral coefficient (MFCC), and formant frequency features to represent the audio characteristics of the emotional speech. A face detection scheme, based on the HSV color model, is used to detect the face from the background. The facial expressions are represented by Gabor wavelet features. We perform feature selection by using a stepwise method based on Mahalanobis distance. A classification scheme involving the analysis of individual class and combinations of different classes is proposed. Our emotion recognition system is tested over a language and race independent database, and an overall recognition accuracy of 82.14% is achieved. Yongjin Wang, Ling Guan |
ICASSP (2) | 1 |
| 2004 | An investigation of speech-based human emotion recognitionabstractThis paper presents our recent work on recognizing human emotion from the speech signal. The proposed recognition system was tested over a language, speaker, and context independent emotional speech database. Prosodic, Mel-frequency cepstral coefficient (MFCC), and formant frequency features are extracted from the speech utterances. We perform feature selection by using the stepwise method based on Mahalanobis distance. The selected features are used to classify the speeches into their corresponding emotional classes. Different classification algorithms including maximum likelihood classifier (MLC), Gaussian mixture model (GMM), neural network (NN), K-nearest neighbors (K-NN), and Fisher's linear discriminant analysis (FLDA) are compared in this study. The recognition results show that FLDA gives the best recognition accuracy by using the selected features. Yongjin Wang, Ling Guan |
MMSP | 1 |