VLDB 2026 Research / reviewers in the wild / expert
Xu Shao
dblp:27/6319
· DBLP profile ↗
34ranked-venue papers
13as first author
3since 2021 · last 2023
0000-0003-1960-1343ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-authorArtificial intelligence and machine learning · 11 · 3 first-author · 1 since 2021Computer networks · 8 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Learning paradigms · 46% Efficient and distributed learning · 46% Speech recognition and synthesis · 8% | |
| Computer networks
1 paper |
Optical networks · 56% Network optimization and economics · 44% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 88% Image and video processing · 12% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning paradigms
continual learning |
0.7 | 1 | 2023 | Rehearsal-free Continual Language Learning via Efficient Parameter Isolation · ACL (1) 2023 |
Machine learning › Efficient and distributed learning
parameter-efficient learning |
0.7 | 1 | 2023 | Rehearsal-free Continual Language Learning via Efficient Parameter Isolation · ACL (1) 2023 |
Network optimization and economics
dynamic resource allocation |
0.2 | 1 | 2013 | A Passive Optical Network with Shared Transceivers for Dynamical Resource Allocation · IEEE Trans. Commun. 2013 |
Optical networks › optical access network
passive optical network |
0.2 | 1 | 2013 | A Passive Optical Network with Shared Transceivers for Dynamical Resource Allocation · IEEE Trans. Commun. 2013 |
Audio and music processing › speech recognition
audio-visual speech recognition |
0.1 | 1 | 2009 | Energetic and Informational Masking Effects in an Audiovisual Speech Recognition System · IEEE Trans. Speech Audio Process. 2009 |
Audio and music processing
speech recognition |
0.1 | 1 | 2009 | Energetic and Informational Masking Effects in an Audiovisual Speech Recognition System · IEEE Trans. Speech Audio Process. 2009 |
Natural language and speech › Speech recognition and synthesis › acoustic modeling
acoustic feature prediction |
0.1 | 1 | 2007 | Prediction of Fundamental Frequency and Voicing From Mel-Frequency Cepstral Coefficients for Unconstrained Speech Reconstruction · IEEE Trans. Speech Audio Process. 2007 |
Optical networks
wavelength-division multiplexing |
0.0 | 1 | 2013 | A Passive Optical Network with Shared Transceivers for Dynamical Resource Allocation · IEEE Trans. Commun. 2013 |
Image and video processing › perceptual modeling › visual perception modeling
masking effects |
0.0 | 1 | 2009 | Energetic and Informational Masking Effects in an Audiovisual Speech Recognition System · IEEE Trans. Speech Audio Process. 2009 |
Audio and music processing › auditory processing
speech perception |
0.0 | 1 | 2009 | Energetic and Informational Masking Effects in an Audiovisual Speech Recognition System · IEEE Trans. Speech Audio Process. 2009 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
distributed speech recognition |
0.0 | 1 | 2007 | Prediction of Fundamental Frequency and Voicing From Mel-Frequency Cepstral Coefficients for Unconstrained Speech Reconstruction · IEEE Trans. Speech Audio Process. 2007 |
Natural language and speech › Speech recognition and synthesis
speech reconstruction |
0.0 | 1 | 2007 | Prediction of Fundamental Frequency and Voicing From Mel-Frequency Cepstral Coefficients for Unconstrained Speech Reconstruction · IEEE Trans. Speech Audio Process. 2007 |
Methods — techniques the papers use, named apart from their topics
rehearsal-free learning · 0.7parameter isolation · 0.7optical multicast switch · 0.2spectro-temporal fragment decomposition · 0.1audiovisual speech models · 0.1hidden markov model · 0.1gaussian mixture model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Rehearsal-free Continual Language Learning via Efficient Parameter IsolationabstractZhicheng Wang, Yufang Liu, Tao Ji, Xiaoling Wang, Yuanbin Wu, Congcong Jiang, Ye Chao, Zhencong Han, Ling Wang, Xu Shao, Wenqiu Zeng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yufang Liu, Yuanbin Wu, Congcong Jiang, Ye Chao, Zhencong Han, Xu Shao, Wenqiu Zeng |
ACL (1) | 10 |
| 2023 | An Empirical Study of the Role of Big Data Analytics in Corporate Decision MakingabstractBounded rationality prevents firms from achieving their full potential. However, intelligent solutions can help eliminate bias in decision making. This study examines whether biases diminish or disappear when novel and powerful digital resources, such as big data analytics, are applied in management practice. The authors use a massive matched database of 1,942 large Chinese firms to find significant and positive effects of data processing frequency on high-level metrics of rational decision-making outcomes, such as productivity and profitability. Moreover, the increase in between-firm variance is the result of both differences in firm characteristics and a widening gap in their workflows and coordinating mechanisms. Heterogeneity can effectively explain 13.18% of the marginal effect of big data analytics on firm metrics, such as productivity and profitability. The results also indicate that the human capital of the C-suite partially mediates the link between big data analytics and firm performance. Xu Shao |
J. Glob. Inf. Manag. | 1 |
| 2021 | Digital Divide or Digital Welfare?: The Role of the Internet in Shaping the Sustainable Employability of Chinese AdultsabstractWith the widespread use of the internet, exploring how it will influence the labor market is of great significance. Based on the 2010-2018 China Family Panel Studies dataset, this paper investigates the effect of the internet on sustainable employability among Chinese aged 16-60. The empirical results of the panel double-hurdle model show that the internet can significantly enhance an individual's competitiveness in the labor market. Moreover, the heterogeneity tests show that the middle aged and older adults, freelancers, and those living in disadvantaged regions can benefit more on employability brought about by the internet. The authors define this phenomenon as the information welfare of the internet, which has narrowed the digital gap caused by the uneven development of technology among different social groups. In addition, the positive coefficient associated with internet use is driven by higher skill requirements in specific workplaces. The authors further explored the role workplace computerization has had in this process. Xu Shao, Yanlin Yang |
J. Glob. Inf. Manag. | 1 |
| 2018 | Nodes contact probability estimation approach based on Bayesian network for DTNabstractDelay tolerant network (DTN) known as suffering from frequent disruption, high latency and heterogeneous, resulting in low network availability. To improve DTN availability, routing protocols typically need to predict the probability of encountering the nodes. In this paper, we use the Bayesian Network (BN) to construct the knowledge base, which is an unique tool for creating a representation of the dependence relationships among DTN parameters. Then developed a Bayesian network- based approach to estimate the contact probability among nodes of DTN. We conducted an experiment to compare our approach against its counterparts in PROPHET routing protocol and power law distribution-based method. The experiment shows our approach is superior to other methods in both recall ratio and precision in all four datasets, including HAGGLE, NUS, REALITY and SASSY. Yuebin Bai, Xu Shao, Wentao Yang 0001, Peng Feng 0003, Rui Wang 0014 |
NOMS | 2 |
| 2016 | Model-Based Parametric Prosody Synthesis with Deep Neural NetworkabstractConventional statistical parametric speech synthesis (SPSS) captures only frame-wise acoustic observations and computes probability densities at HMM state level to obtain statistical acoustic models combined with decision trees, which is therefore a purely statistical data-driven approach without explicit integration of any articulatory mechanisms found in speech production research. The present study explores an alternative paradigm, namely, model-based parametric prosody synthesis (MPPS), which integrates dynamic mechanisms of human speech production as a core component of F0 generation. In this paradigm, contextual variations in prosody are processed in two separate yet integrated stages: linguistic to motor, and motor to acoustic. Here the motor model is target approximation (TA), which generates syllable-sized F0 contours with only three motor parameters that are associated to linguistic functions. In this study, we simulate this two-stage process by linking the TA model to a deep neural network (DNN), which learns the “linguistic-motor” mapping given the “motor-acoustic” mapping provided by TA-based syllable-wise F0 production. The proposed prosody modeling system outperforms the HMM-based baseline system in both objective and subjective evaluations. Xu Shao |
INTERSPEECH | 3 |
| 2015 | Pruning redundant synthesis units based on static and delta unit appearance frequency
Wei Zhang 0189, Xu Shao, Wenhui Lei, Hongbin Zhou, Andrew P. Breen |
INTERSPEECH | 3 |
| 2013 | A Passive Optical Network with Shared Transceivers for Dynamical Resource AllocationabstractThis paper presents a new passive optical network (PON) architecture that applies an optical multicast capable switch (OMS) to direct wavelengths to a set of PON branches and enables them to share the transceivers in the central office (CO) on demand. The OMS is composed of fiber splitters and switching gates, and by controlling the on-off gates the wavelengths at the inputs can reach any outputs of the OMS. Through the OMS configuration, logical PONs are dynamically formed according to changing traffic patterns, and traffic is further scheduled in the time domain within each logical PON. The proposed network achieves network resource sharing without requiring tunable transceivers or wavelength filters, and without changing deployed network devices in the field or at the user end. Analysis and experimental study results show the efficiency of resource sharing and demonstrate the concept of resource sharing among PON branches. Luying Zhou, Zhaowen Xu, Xiaofei Cheng, Yong-Kee Yeo, Xu Shao |
IEEE Trans. Commun. | 6 |
| 2012 | A dynamic wavelength resource allocation capable passive optical network with shared transceiversabstractThe paper proposes a new PON architecture that applies an optical broadcast capable router (OBR) to route a wavelength to a set of selected PON branches and thus enables the transceivers in the central office (CO) to be shared among them. The OBR is built on fiber splitters and switching gates, and by controlling the gates the wavelengths at the inputs can reach any outputs of the OBR. Logical PONs are dynamically configured based on users' changing traffic patterns, and the traffic is further scheduled in the time domain in the logical PONs. The network achieves network resource sharing without requiring tunable transceivers or wavelength filters, and furthermore without changing deployed network devices in the field or at the user end. Experiment study and analysis results demonstrate the concept of resource sharing among PON branches and show the efficiency of resource sharing. Luying Zhou, Zhaowen Xu, Xiaofei Cheng, Yong-Kee Yeo, Xu Shao |
ICC | 5 |
| 2012 | Semantic IPTV Service Discovery SystemabstractIPTV deals with multimedia services delivered over IP based networks managed to provide the required level of quality of service and experience, security, interactivity and reliability. Service discovery is an important component of IPTV system, but it is challenging due to the unique domain specifics of IPTV and complexities of IPTV ecosystem. In this paper, we formulate the IPTV service discovery problem under an SOA framework, and propose a novel standards-based semantic IPTV service discovery system. The proposed system is based on unifying the traditionally step-wise iterative process of IPTV service discovery, and a semantic integration of service provider and IPTV service descriptions. For demonstration, we consider a specific discovery scenario involving service providers that offer a combination of standard TV services such as linear TV and VoD content, and also advanced interactive multimedia-based IPTV services. With detailed illustration, we demonstrate that our proposed system provides simple but effective discovery of IPTV services, and possesses several benefits that can potentially enhance end-user experience. Fon Lin Lai, Xu Shao, Kanagasabai Rajaraman |
SERVICES | 2 |
| 2012 | FAST: Fuzzy Decision-Based Resource Admission Control Mechanism for MANETs
Yuebin Bai, Xu Shao, Wentao Yang 0001 |
Mob. Networks Appl. | 3 |
| 2010 | A Case Study on the Value Models of Ocean Transportation ServiceabstractVASEM is a methodology to be aware of service values among participants and to analyze values' relationship to direct service system design. Value model is an important part in VASEM, and is foundation of value analysis. Ocean Transportation Service is a typical service system. Take it as a case, we verify the usability of value model in a certain field. We analyze the values in Ocean Transportation Service System with 4-layer value model, introduce the forms and specifications of each model, and present a whole view of the service by VPM, POVN and VDM. On the last layer - VAM, Value Model is combined with SMDA model by value annotation. And we show a segment of the service process to form primary methods for value annotation. Xu Shao, Zhongjie Wang 0003, Xiaofei Xu 0001, Chao Ma 0017 |
ICSS | 1 |
| 2010 | Value Annotation for Service Model AnalysisabstractService modeling is a critical step in describing customer requirements and designing service systems. The quality of service models determines the quality of service systems to a great extent. To analyze the service models in a way that would show deficiencies in delivering service values, we propose the idea of "value annotation" by which expected values are labeled on the functional elements of service models, and their implementation degrees are connected with QoS of these functional elements. Value annotation's elementary principles, process, and model forms are presented. Then, the preliminary principles of value annotation-based service model analysis are briefly introduced. To complement the above discussion, a small case study from ocean transportation service is applied for demonstration. Zhongjie Wang 0003, Xiaofei Xu 0001, Chao Ma 0017, Xu Shao |
ICSS | 5 |
| 2009 | An IMS-based testbed for real-time services integration and orchestrationabstractIn this paper, we introduce an IP multimedia sub-system (IMS) based testbed which provides a platform for the study of real-time services integration and orchestration. This open-source based testbed is built on the principle of service oriented architecture (SOA), with an emphasis for real-time network services. We further developed service-oriented system functionalities such as optical network connection and bandwidth management, mobile client authentication protocols, and video streaming services. They are packaged as interoperable service components, so that they can be integrated and orchestrated through their respective standard interfaces. Finally, we elaborated our proof-of-concept environment via a use-case scenario of having various client-server interactions over a heterogeneous network environment. Teck Kiong Lee, Teck Yoong Chai, Lek Heng Ngoh, Xu Shao, Joseph Chee Ming Teo, Luying Zhou |
APSCC | 4 |
| 2009 | Multipath cross-layer service discovery (MCSD) for mobile ad hoc networksabstractIn this paper, we propose and study how to provide multipath cross-layer service discovery (MCSD) for mobile ad hoc networks (MANETs). Cross-layer service discovery integrates service discovery into route discovery by taking advantage of network-layer topology information and routing message exchange. Multipath service discovery differs from multipath routing in that multipath service discovery can be either multipath from a client to a server or multipath from a client to multiple servers, depending on the policy used in service selection and invocation. Compared with the traditional unipath cross-layer service discovery (UCSD), MCSD performs better in terms of improving throughput, enhancing service availability and optimizing network-layer resources. However, as a new service discovery scheme, MCSD posts a lot of challenges, particularly in the process of service discovery and how to find the optimal multipath. Due to the complexity of the problem, we focus on double-path cross-layer service discovery (DCSD), a special and most important case of MCSD. We present a heuristic, called iDCSD, which can intelligently find the optimal routes out of the candidate paths from a client to a server and from a client to two servers by minimizing hop count in network layer. Extensive simulation results show that compared with UCSD, (i) MCSD delivers better service availability; and (ii) the proposed iDCSD leads to a better network-layer performance with reasonable computational complexity. Xu Shao, Lek Heng Ngoh, Teck Kiong Lee, Teck Yoong Chai, Luying Zhou, Joseph Chee Ming Teo |
APSCC | 1 |
| 2009 | On minimum data replication for delay-bounded query in wireless ad hoc networksabstractIn this paper, we study the problem of minimizing the number of data replicas for delay-bounded queries in wireless ad hoc networks. We focus our attention on step-by-step expanding ring search, which provides an upper bound on query delay to any expanding ring based search strategies. We analyze the probabilistic behavior of query delay, and develop an analytical approach to approximate the minimum number of data replicas for delay bounded data query in wireless ad hoc networks. We validate our analysis through extensive simulations. Jun Huang 0001, Yuebin Bai, Xu Shao |
WCNC | 3 |
| 2009 | Energetic and Informational Masking Effects in an Audiovisual Speech Recognition SystemabstractThe paper presents a robust audiovisual speech recognition technique called audiovisual speech fragment decoding. The technique addresses the challenge of recognizing speech in the presence of competing nonstationary noise sources. It employs two stages. First, an acoustic analysis decomposes the acoustic signal into a number of spectro-temporall fragments. Second, audiovisual speech models are used to select fragments belonging to the target speech source. The approach is evaluated on a small vocabulary simultaneous speech recognition task in conditions that promote two contrasting types of masking:energeticmaskingcaused by the energy of the masker utterance swamping that of the target, andinformationalmasking, caused by similarity between the target and masker making it difficult to selectively attend to the correct source. Results show that the system is able to use the visual cues to reduce the effects of both types of masking. Further, whereas recovery fromenergeticmaskingmay require detailed visual information (i.e., sufficient to carry phonetic content), release frominformationalmaskingcan be achieved using very crude visual representations that encode little more than the timing of mouth opening and closure. Jon Barker, Xu Shao |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | An Integrated Telecom and IT Service Delivery PlatformabstractIP Multimedia Subsystem (IMS) and Web services (WS) are service-oriented architectures developed separately for service delivery in the next generation telecommunications, and IT-centric computing environment, respectively. In order to harness services in both of these platforms and to facilitate combining and blending of services, we propose an integrated telecom and IT Service Delivery Platform for interworking between IMS and WS. We propose to use SIP-based Micro Service Orchestration and Web service bus to seamlessly integrate IMS, WS and the underlying services. We further present an example of managing multimedia services over an Ethernet Passive Optical Networks (EPON) infrastructure to illustrate that the proposed Service Delivery Platform has the benefits of supporting rapid development and deployment of new converged multimedia services. Xu Shao, Teck Yoong Chai, Teck Kiong Lee, Lek Heng Ngoh, Luying Zhou, Markus Kirchberg |
APSCC | 1 |
| 2008 | Best Effort Shared Risk Link Group (SRLG) Failure Protection in WDM NetworksabstractWith the increase of size and number of shared risk link groups (SRLGs), capacity efficiency of shared-path protection becomes much poorer due to SRLG-disjoint constraints and blocking probability becomes much higher due to severe traps. As a result, 100% SRLG failure protection is no longer a practical protection scheme. To solve this problem, we present a new protection scheme called best effort SRLG failure protection, in which we try to provide SRLG-disjoint backup path by choosing the backup path sharing the least number of SRLGs with the working path, so as to make the impact of SRLG failures as low as possible and accept as many as possible connection requests. 100% SRLG failure protection becomes a special case of best effort SRLG failure protection when the working path and backup path share zero SRLG. We propose a heuristic to find the best effort SRLG-disjoint backup path under dynamic traffic. The best effort SRLG failure protection scheme tries to make a trade-off between blocking probability and survivability. Analytical and simulation results show, compared with 100% SRLG failure protection, the proposed scheme offers much better capacity efficiency and much lower blocking probability while keeping survivability as high as possible. Xu Shao, Luying Zhou, Xiaofei Cheng, Weiguo Zheng, Yixin Wang 0005 |
ICC | 1 |
| 2008 | Stream weight estimation for multistream audio-visual speech recognition in a multispeaker environment
Xu Shao, Jon Barker |
Speech Commun. | 1 |
| 2007 | Providing Differentiated Quality-of-Protection for Surviving Double-Link Failures in WDM Mesh NetworksabstractProviding differentiated quality-of-protection (QoP) for surviving single-link failures in WDM mesh networks has been extensively studied in recent years. This paper investigates the problem of providing differentiated QoP for surviving arbitrary double-link failures by allowing a connection request to choose from several QoP classes. In this paper, we propose to use three classes, i.e., single shared-path protection (SSPP), single dedicated-path protection (SDPP), and double shared-path protection (DSPP) to provide differentiated QoP. We present two differentiated QoP schemes. Scheme 1 (conventional differentiated QoP) is a natural extension of conventional differentiated QoP for surviving single-link failures, which uses SDPP, SSPP, and SDPP separately to satisfy different QoP requirements. Scheme 2 (shared differentiated QoP) tries to share backup resources between SSPP and SDPP. Simulation results show that our proposed architecture of QoP can satisfy different QoP requirements for surviving double-link failures by making a balance between blocking probability and average QoP. Analytical and numerical results indicate that the proposed differentiated QoP scheme 2 are more efficient in improving not only blocking probability but also average QoP. Xu Shao, Luying Zhou, Weiguo Zheng, Yixin Wang 0005 |
ICC | 1 |
| 2007 | iOPEN Network: Operation Mechanisms and Experimental StudyabstractIntegrated OPtical EtherNet (iOPEN) network was earlier proposed as an integrated Ethernet and reconfigurable WDM network, which supports Ethernet transport services in campus area and metro area networks. In this paper we propose and study network operation mechanisms for iOPEN network to achieve cost-effective network operation. VLAN approach is extended to operate new lightpaths, to support QoS services and traffic engineering, and to retain the Ethernet simplicity features. The proposed operation mechanisms are evaluated over an iOPEN testbed network. Luying Zhou, Xu Shao, Teck Yoong Chai, Chava Vijaya Saradhi, Yixin Wang 0005 |
ICC | 2 |
| 2007 | Prediction of Fundamental Frequency and Voicing From Mel-Frequency Cepstral Coefficients for Unconstrained Speech ReconstructionabstractThis work proposes a method for predicting the fundamental frequency and voicing of a frame of speech from its mel-frequency cepstral coefficient (MFCC) vector representation. This information is subsequently used to enable a speech signal to be reconstructed solely from a stream of MFCC vectors and has particular application in distributed speech recognition systems. Prediction is achieved by modeling the joint density of fundamental frequency and MFCCs. This is first modeled using a Gaussian mixture model (GMM) and then extended by using a set of hidden Markov models to link together a series of state-dependent GMMs. Prediction accuracy is measured on unconstrained speech input for both a speaker-dependent system and a speaker-independent system. A fundamental frequency prediction error of 3.06% is obtained on the speaker-dependent system in comparison to 8.27% on the speaker-independent system. On the speaker-dependent system 5.22% of frames have voicing errors compared to 8.82% on the speaker-independent system. Spectrogram analysis of reconstructed speech shows that highly intelligible speech is produced with the quality of the speaker-dependent speech being slightly higher owing to the more accurate fundamental frequency and voicing predictions Ben P. Milner, Xu Shao |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Complementary Protection Under Double-Link Failure for Survivable Optical NetworksabstractIn previous protection schemes against random double-link failure scenarios, two link-disjoint backup paths are provided for each connection. However, the protection domains (protected double-link failure scenarios) of the two backup paths are not optimized, i.e., they overlap, such that the two backup paths of a connection will protect the same double-link failure scenarios. This will waste a large amount of backup bandwidth. In this paper, we present a novel complementary protection scheme. The two backup paths of each connection will complementarily protect the connection so that the protection domains of each connection are optimized. The path selection and the bandwidth allocation are thus optimized. The mathematical model of complementary protection and the resource allocation/release heuristic algorithm are given. Simulation results show that our algorithm has better performance in terms of total bandwidth consumption and blocking probability under double-link failure scenarios, meanwhile, our algorithm has fast on-line characteristic for dynamic connections. Xiaofei Cheng, Teck Yoong Chai, Xu Shao, Yixin Wang 0005 |
GLOBECOM | 3 |
| 2006 | Audio-visual speech recognition in the presence of a competing speakerabstractThis paper examines the problem of estimating stream weights for a multistream audio-visual speech recogniser in the context of a simultaneous speaker task. The task is challenging because signalto-noise ratio (SNR) cannot be readily inferred from the acoustics alone. The method proposed employs artificial neural networks (ANNs) to estimate the SNR from HMM state-likelihoods. SNR is converted to stream weight using a mapping optimised on development data. The method produces an audio-visual recognition performance better than that of both the audio-only and the videoonly baselines across a wide range of SNRs. The performance using SNR estimates based on audio state-likelihoods is compared to that obtained using both audio and visual likelihoods. Although the audio-visual SNR estimator outperforms the audio-only SNR estimator, the recognition performance benefit is small. Ideas for making fuller use of the visual information are discussed. Index Terms: audio-visual speech recognition, multistream, stream weighting, SNR estimation, artificial neural networks Xu Shao, Jon Barker |
INTERSPEECH | 1 |
| 2006 | Clean speech reconstruction from MFCC vectors and fundamental frequency using an integrated front-end
Ben P. Milner, Xu Shao |
Speech Commun. | 2 |
| 2005 | Predicting Formant Frequencies from MFCC VectorsabstractThis work proposes a novel method of predicting formant frequencies from a stream of mel-frequency cepstral coefficients (MFCC) feature vectors. Prediction is based on modelling the joint density of MFCCs and formant frequencies using a Gaussian mixture model (GMM). Using this GMM and an input MFCC vector, two maximum a posteriori (MAP) prediction methods are developed. The first method predicts formants from the closest, in some sense, cluster to the input MFCC vector, while the second method takes a weighted contribution of formants predicted from all clusters. Experimental results are presented using the ETSI Aurora connected digit database and show that predicted formant frequencies are within 3.2% of reference formant frequencies. Jonathan Darch, Ben P. Milner, Xu Shao, Saeed Vaseghi, Qin Yan |
ICASSP (1) | 3 |
| 2005 | Fundamental frequency and voicing prediction from MFCCs for speech reconstruction from unconstrained speechabstractThis work proposes a method to predict the fundamental frequency and voicing of a frame of speech from its MFCC representation. This has particular use in distributed speech recognition systems where the ability to predict fundamental frequency and voicing allows a time-domain speech signal to be reconstructed solely from the MFCC vectors. Prediction is achieved by modeling the joint density of MFCCs and fundamental frequency with a combined hidden Markov model-Gaussian mixture model (HMM-GMM) framework. Prediction results are presented on unconstrained speech using both a speaker-dependent database and a speaker-independent database. Spectrogram comparisons of the reconstructed and original speech are also made. The results show for the speaker-dependent task a percentage fundamental frequency prediction error of 3.1% is made while for the speaker-independent task this rises to 8.3%. Ben P. Milner, Xu Shao, Jonathan Darch |
INTERSPEECH | 2 |
| 2004 | Pitch prediction from MFCC vectors for speech reconstructionabstractThe paper proposes a technique for reconstructing an acoustic speech signal solely from a stream of Mel-frequency cepstral coefficients (MFCCs). Previous speech reconstruction methods have required an additional pitch element, but this work proposes two maximum a posteriori (MAP) methods for predicting pitch from the MFCC vectors themselves. The first method is based on a Gaussian mixture model (GMM) while the second scheme utilises the temporal correlation available from a hidden Markov model (HMM) framework. A formal measurement of both frame classification accuracy and RMS pitch error shows that an HMM-based scheme with 5 clusters per state is able to classify correctly over 94% of frames and has an RMS pitch error of 3.1 Hz in comparison to a reference pitch. Informal listening tests and analysis of spectrograms reveals that speech reconstructed solely from the MFCC vectors is almost indistinguishable from that using the reference pitch. Xu Shao, Ben P. Milner |
ICASSP (1) | 1 |
| 2004 | MAP prediction of pitch from MFCC vectors for speech reconstructionabstractThis work proposes a method of predicting pitch and voicing from mel-frequency cepstral coefficient (MFCC) vectors. Two maximum a posteriori (MAP) methods are considered. The first models the joint distribution of the MFCC vector and pitch using a Gaussian mixture model (GMM) while the second method also models the temporal correlation of the pitch contour using a combined hidden Markov model (HMM)-GMM framework. Monophone-based HMMs are connected together in the form of an unconstrained monophone grammar which enables pitch to be predicted from unconstrained speech input. Evaluation on 130,000 MFCC vectors reveals a voicing classification accuracy of over 92% and an RMS pitch error of 10Hz. The predicted pitch contour is also applied to MFCC-based speech reconstruction with the resultant speech almost indistinguishable from that reconstructed using a reference pitch. 1. Xu Shao, Ben P. Milner |
INTERSPEECH | 1 |
| 2003 | Low bit-rate feature vector compression using transform coding and non-uniform bit allocationabstractThe paper presents a novel method for the low bit-rate compression of a feature vector stream with particular application to distributed speech recognition. The scheme operates by grouping feature vectors into non-overlapping blocks and applying a transformation to give a more compact matrix representation. Both Karhunen-Loeve and discrete cosine transforms are considered. Following transformation, higher-order columns of the matrix can be removed without loss in recognition performance. The number of bits allocated to the remaining elements in the matrix is determined automatically using a measure of their relative information content. Analysis of the amplitude distribution of the elements indicates that non-linear quantisation is more appropriate than linear quantisation. Comparative results, based on both spectral distortion and speech recognition accuracy, confirm this. Speech recognition tests using the ETSI Aurora database demonstrate that compression to bits rates of 2400 bps, 1200 bps and 800 bps has very little effect on recognition accuracy. For example at a bit rate of 1200 bps, recognition accuracy is 98.0% compared to 98.6% with no compression. Ben P. Milner, Xu Shao |
ICASSP (2) | 2 |
| 2003 | Clean speech reconstruction from noisy mel-frequency cepstral coefficients using a sinusoidal modelabstractThis paper extends the technique of speech reconstruction from MFCC by considering the effect of noisy speech. To reconstruct a clean speech signal from noise contaminated MFCC an estimate of the clean mel-filterbank vector is required together with a robust estimate of the pitch. This work applies spectral subtraction to the mel-filterbank vector (derived from noisy MFCC) to provide a clean speech spectral estimate. To obtain a reliable estimate of pitch a robust extraction technique is used. Spectrograms and informal listening tests reveal that a clean speech signal can be successfully reconstructed from the noisy MFCC. Pitch errors are shown to manifest themselves as artificial sounding bursts in the reconstructed speech signal. Incorrect estimates of the spectral envelope introduce periods of noise into the reconstructed speech. Xu Shao, Ben P. Milner |
ICASSP (1) | 1 |
| 2003 | Integrated pitch and MFCC extraction for speech reconstruction and speech recognition applicationsabstractThis paper proposes an integrated speech front-end for both speech recognition and speech reconstruction applications. Speech is first decomposed into a set of frequency bands by an auditory model. The output of this is then used to extract both robust pitch estimates and MFCC vectors. Initial tests used a 128 channel auditory model, but results show that this can be reduced significantly to between 23 and 32 channels. A detailed analysis of the pitch classification accuracy and the RMS pitch error shows the system to be more robust than both comb function and LPC-based pitch extraction. Speech recognition results show that the auditory-based cepstral coefficients give very similar performance to conventional MFCCs. Spectrograms and informal listening tests also reveal that speech reconstructed from the auditory-based cepstral coefficients and pitch has similar quality to that reconstructed from conventional MFCCs and pitch. 1. Xu Shao, Ben P. Milner, Stephen J. Cox |
INTERSPEECH | 1 |
| 2002 | Speech reconstruction from mel-frequency cepstral coefficients using a source-filter modelabstractThis work presents a method of reconstructing a speech signal from a stream of MFCC vectors using a source-filter model of speech production. The MFCC vectors are used to provide an estimate of the vocal tract filter. This is achieved by inverting the MFCC vector back to a smoothed estimate of the magnitude spectrum. The Wiener- Khintchine theorem and linear predictive analysis transform this into an estimate of the vocal tract filter coefficients. The excitation signal is produced from a series of pitch pulses or white noise, depending on whether the speech is voiced or unvoiced. This pitch estimate forms an extra element of the feature vector. Listening tests reveal that the reconstructed speech is intelligible and of similar quality to a system based on LPC analysis of the original speech. Spectrograms of the MFCC-derived speech and the real speech are included which confirm the similarity. Ben P. Milner, Xu Shao |
INTERSPEECH | 2 |
| 2002 | Transform-based feature vector compression for distributed speech recognitionabstractThe technique of distributed speech recognition (DSR) has recently become an interesting area of research. One of the main issues with DSR is the need to compress the feature vector stream, produced on the terminal device, into a sufficiently low bit-rate such that it can be sent across low bandwidth channels. This work proposes a compression technique based upon first transforming a block of feature vectors into a more compact matrix representation. Columns of the resulting matrix that correspond to faster temporal variation can be removed without loss in recognition performance. The number of bits allocated to the remaining coefficients in the matrix is determined automatically, based on a measure of the information present. Experiments show that the transform-based compression gives good recognition accuracy at bit rates of 4800, 2400 and 1200bps. For example at 1200bps the recognition performance is 98.03% compared to 98.57% with uncompressed speech. Ben P. Milner, Xu Shao |
INTERSPEECH | 2 |