EDBT 2026 Demo / reviewers in the wild / expert
Feipeng Li
dblp:71/7663
· DBLP profile ↗
22ranked-venue papers
9as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-authorComputer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | xWitch: Towards Fast and Accurate Performance Evaluation for Hierarchical QoSabstractHierarchical Quality of Service (HQoS) has been designed to meet the diverse needs of different users and applications, and it is widely applied in commercial routers. However, the large number of users and applications results in numerous HQoS queue parameters that need to be configured. Fast and accurate performance evaluation of these configurations is crucial. Traditional discrete event network simulators experience significant slowdowns as the traffic increases. Existing machine learning-based methods for performance evaluation also face challenges, such as low accuracy and poor generalization, due to the long input sequences caused by large network traffic.In this paper, we propose xWitch, a packet-level performance evaluation scheme that supports HQoS. xWitch uses a sequence-to-sequence model to achieve short sequence performance prediction. A long sequence parallel prediction scheme based on dependency prediction is proposed to support fast and accurate prediction of long sequences. Experimental results show that xWitch outperforms all baselines. It achieves an average error of under 3% in latency prediction for short sequences. For long sequence prediction, it can reduce the error by more than 20% and consistently maintains latency prediction errors below 10% across different traffic distributions. Gang Yi, Mowei Wang, Chuxuan Zeng, Feipeng Li, Yong Cui 0001 |
IWQoS | 5 |
| 2025 | LSTM Network Assisted Construction of the Angle-Dependent Point Spread Function and Its Applications in Seismic ImagingabstractMigration is the core link in reflection seismic exploration. Seismic images are often extended into angle domain for interpretations. However, affected by limited acquisition aperture and complex overburden, the generated images are far from ideal. The illumination is unbalanced, causing unreliable amplitude variation in angle gathers. Band-limited seismic data and wavelet stretch in large angles, lead to low-resolution angle gathers. Image-domain least-squares migration (IDLSM) implemented by point spread function (PSF) deconvolution is a promising solution. Extending the concept of IDLSM to the angle domain, we develop a new method to construct angle-dependent PSFs and optimize angle gathers. The essential element to construct PSFs is the Green’s function. The proposed method reconstructs Green’s functions using a bidirectional long short-term memory (LSTM) network. We use a ray tracing method to efficiently obtain wave propagation directions (travel-time gradients). And wave-equation forward modeling is used to accurately calculate wavefront amplitudes. The LSTM network is trained by labels composed of travel-time gradients and amplitudes to surrogate the solver of Green’s functions. Angle-dependent PSFs are constructed according to the mathematical model of the angular local Hessian. And inversions with PSFs are performed to optimize angle gathers. Numerical tests on a 3-D synthetic model demonstrate that the proposed method is able to improve the image quality of prestack angle gathers and poststack seismic images. The proposed method compensates illumination and improves the resolution of angle-dependent seismic images. Both vertical and lateral resolution are enhanced. Amplitude versus angle (AVA) responses can be corrected for further analysis and interpretations. Feipeng Li, Jinghuai Gao, Zhiguo Wang 0002, Chuang Li 0003, Zhaoqi Gao, Zongben Xu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | True Amplitude Seismic Imaging With Wave Equation-Based Illumination Compensation in the Dip and Reflection Angle DomainabstractSeismic interpretation and reservoir characterization require the seismic data having faithful amplitudes that relate to subsurface physical parameters. Nowadays, the amplitude fidelity of seismic imaging becomes more important than ever. Although reverse time migration (RTM) adopts the full wave equation as true amplitude seismic wave propagator, it is still not sufficient for true amplitude seismic imaging since migration is only the adjoint operator corresponding to the forward modeling process. The complex overburden and limited migration aperture lead to unbalanced illumination of subsurface structures. Least-squares migration was proposed to correct amplitudes of seismic images, but it is computationally expensive and sometimes unstable. The illumination compensation is an available alternative which only considers the amplitude correction regardless of the resolution issue. In this article, we propose a true amplitude seismic imaging method with illumination compensation performed on both RTM stacked images and angle gathers. We derive the angle-dependent illumination intensity from the Hessian of least-squares migration in which Green’s functions are essential components. We propose a new method to estimate the Green’s function and its corresponding wave propagation direction based on wavefields excitation amplitudes and Poynting vectors at excitation times. Then, the illumination intensity is constructed as a function of dip and reflection angles to correct both angle gathers and stacked images. The proposed method is tested using two synthetic models and a real marine dataset. Numerical results demonstrate that the proposed method can effectively correct amplitudes of seismic images. Deep events beneath complex structures are enhanced with more balance illumination. Feipeng Li, Jinghuai Gao, Zhiguo Wang 0002, Chuang Li 0003, Zhaoqi Gao, Zongben Xu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Diffraction Separation and Imaging Using Multidirectional Wavefield Low-Rank ApproximationabstractLow-rank approximation (LRA) is a powerful technique for seismic diffraction separation and imaging, providing higher-resolution images of subsurface discontinuities compared to traditional reflection imaging. However, in complex wavefields where reflections lack distinct low-rank characteristics, diffractions and reflections can overlap within the same eigenimages, making traditional LRA less effective for separation. To address this limitation, we propose a diffraction separation and imaging method based on multi-directional wavefield low-rank approximation (MDWLRA). The MDWLRA method employs multi-directional wavefield decomposition (MDWD) to divide complex wavefields into angular slices with similar dip angles. These slices are classified as either diffraction slices or reflection slices, with the latter containing mostly reflections and some residual diffractions. LRA is then applied to the reflection slices to separate the remaining diffractions from reflections. By reducing wavefield complexity using MDWD, reflections in the reflection slices exhibit clearer low-rank characteristics than those in the full wavefields, allowing for more effective separation using LRA. Numerical tests on synthetic data from the modified Sigsbee2A model and field data demonstrate that the MDWLRA method outperforms traditional methods, achieving more accurate separation with fewer leakages than traditional LRA, while also improving diffraction fidelity compared to the Curvelet-transform-based method. Chuang Li 0003, Yibo Hou, Shixuan Jia, Zhaoqi Gao, Feipeng Li, Zhen Li 0016, Jinghuai Gao, Zhiguo Huang, Ling Qian |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Hessian-Assisted Iterative Self-Training Learning for Seismic MigrationabstractSeismic migration produces the migrated images of subsurface media using seismic data, which is important for geophysical exploration. However, the adjoint-based migration methods may produce a blurry image, convolved by a Hessian matrix. To address this problem, we propose a Hessian-assisted iterative self-training learning (HAISTL) method aimed at approximating the inverse Hessian matrix and deblurring the migrated image. First, we train a long short-term (LSTM) network using labeled images and use it as a teacher network to generate pseudolabels for the unlabeled images. Subsequently, we integrate the demigration and migration operators to identify the pseudolabels with high confidence levels and construct a dataset containing both the true and pseudolabels. The dataset is then used to train a student network with the injection of model noise into the network. Finally, we regard the student network as a new teacher and repeat the process in an iterative STL framework. We demonstrate the effectiveness of our proposed method using two synthetic datasets and field data. Compared with the supervised learning (SL) method, the proposed method exhibits superior generalization capabilities. This advantage stems from the incorporation of the demigration and migration operators, providing a valuable prior for the inverse Hessian matrix in training the model. In contrast to the model-driven least-squares migration (LSM) methods, the proposed method yields high-resolution images with significantly reduced computational costs. However, it may be less effective in recovering small-scale structures when confronted with an extremely limited number of labels. Chuang Li 0003, Bingbing Wu, Zhaoqi Gao, Wei Zhang 0212, Feipeng Li, Jincheng Xu, Jinghuai Gao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Generating Azimuth-Reflection Angle Gathers From Reverse Time Migration Using the High-Dimensional Local Phase Space Approximation of Seismic WavefieldsabstractAmplitude-preserving angle gathers are ideal inputs for seismic prestack inversion. However, due to the limitation of computational efficiency, generating subsurface azimuth–reflection angle gathers from 3-D seismic imaging is still a very difficult task. In this article, we propose a new method to generate azimuth–reflection angle gathers from 3-D reverse time migration (RTM). The proposed method approximately reconstructs the source wavefield using high-dimensional wavelets and the excitation information. After using directional vectors to calculate the subsurface observation angles and applying the cross correlation imaging condition, we can generate azimuth–reflection angle gathers by angle binning. Without storing source wavefields or reconstructing source wavefields using boundary conditions, the proposed method has high computational efficiency. Numerical experiments on a synthetic model and a real marine seismic dataset demonstrate that compared with the excitation amplitude imaging condition, the proposed method can generate azimuth–reflection angle gathers with continuous complete events and high signal-to-noise ratio. The image quality and resolution of angle gathers are significantly improved. At the same time, the computational complexity does not increase much. Feipeng Li, Jinghuai Gao, Zhaoqi Gao, Chuang Li 0003, Qingzhen Wang, Zongben Xu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Diffraction Separation and Least-Squares Imaging Based on Multiscale and Multidirectional Wavefield and Image DecompositionabstractDiffraction separation and imaging are important for subsurface discontinuities characterization. However, conventional diffraction separation methods may loss validity when the diffractions and reflections do not have discernible differences in data domain. Moreover, due to limited acquisition geometry and narrow frequency band of seismic data, the diffraction imaging methods that use conventional ray-based or wave-equation-based migration operators may produce images with low resolution. We propose a diffraction separation and least-squares imaging method based on multi-scale and multi-directional wave-field and image decomposition. First, by using the multi-scale and multi-directional properties of the generalized curvelet transform, we reproduce the diffractions from the plane-wave sections according to the differences between the diffractions and reflections in terms of scale and angle. When the diffractions and reflections do not have discernible differences in data domain, their migrated images generally have different dip angles. Therefore, we propose a plane-wave least-squares diffraction imaging method with a curvelet-domain regularization which suppresses the images of residual reflections with small dip angles. Finally, we obtain high-resolution images of the subsurface discontinuities by using a regularized conjugate gradient method. Synthetic and field data examples verify the superiority of the proposed diffraction separation method over the plane-wave destruction filter in terms of better suppression of the reflections and better recovery of the diffractions. Compared with plane-wave reverse time migration, the proposed diffraction imaging method effectively suppresses the residual reflectors and produces images with higher resolution and signal-to-noise ratio. Chuang Li 0003, Shixuan Jia, Zhen Li 0016, Zhaoqi Gao, Feipeng Li, Jinghuai Gao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Self-Supervised Deep Learning for Nonlinear Seismic Full Waveform InversionabstractSeismic full waveform inversion (FWI) is able to build high-resolution velocity model based on the full information carried by seismic wave. However, FWI requires an accurate enough initial model to ensure convergence. In this paper, we propose a new nonlinear FWI method to mitigate the initial model dependence problem. Specifically, we firstly propose a nonlinear operator within the hybrid model- and data-driven framework based on the frequency controllable envelope operator (FCEO) and a deep learning architecture U-Net. FCEO is used to obtain the envelope of a band-limited data and U-Net realizes the mapping from this envelope to that corresponding to a lower frequency band. The U-Net is trained in a self-supervised manner that avoids the reliance on labeled data and benefits the generalization ability. Based on the nonlinear operator, a nonlinear FWI method is proposed by defining a new misfit function. In addition, the calculation of gradient is derived using the adjoint-state method. Using numerical examples, we investigate the performance of the proposed nonlinear operator and the new nonlinear FWI method. The results clearly demonstrate that the proposed nonlinear operator is effective in obtaining low-frequency envelope data, and the new nonlinear FWI method has advantages over common method in mitigating cycle-skipping and building an initial model for conventional FWI. Zhaoqi Gao, Chuang Li 0003, Feipeng Li, Qingzhen Wang, Jicai Ding, Jinghuai Gao, Zongben Xu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Least-Squares Reverse Time Migration With Curvelet-Domain Preconditioning OperatorsabstractLeast-squares reverse time migration (LSRTM) is an amplitude-preserving seismic imaging technique that aims at finding the subsurface reflectivity model. It is often performed iteratively using an inversion algorithm, such as the conjugate gradient method. Such an implementation requires a huge amount of calculation as it may converge slowly. Preconditioning plays a crucial role in seismic inverse problems. In this study, we propose a novel preconditioning method for LSRTM. The proposed method estimates a new guided curvelet-domain deblurring filter for one-step LSRTM and preconditioned LSRTM. Then, the filter is applied to migrated images and gradients in LSRTM. Such a deblurring filter acts as a curvelet-domain local linear approximation of the least-squares functional inverse Hessian, which can improve the image quality and accelerate the convergence. Numerical tests on the synthetic model and a field data example demonstrate that the preconditioning operator can effectively accelerate the convergence of LSRTM. One-step LSRTM can obtain comparable image quality to that of conventional iterative LSRTM with only a single iteration. The comparison of the convergence curves demonstrates that the curvelet-domain preconditioning operators accelerate the convergence of LSRTM. Furthermore, the preconditioned LSRTM achieves better image quality than the conventional LSRTM. We compare preconditioning operators based on diagonal and local linear approximations. The preconditioning operator based on the local linear approximation has more robust performance and more stable convergence curves than the diagonal-based approximation. Feipeng Li, Jinghuai Gao, Zhaoqi Gao, Chuang Li 0003, Wei Zhang 0212 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Least-Squares Reverse Time Migration for Reflection-Angle-Dependent ReflectivityabstractLeast-squares reverse time migration (LSRTM) can estimate high-quality reflectivity of subsurface medium from seismic data. However, the subsurface reflectivity depends on reflection angles, and its variations over reflection angles are extremely important because they can be used to estimate sub-surface physical properties for seismic interpretation. We present a new formulation of the LSRTM method that can estimate reflection-angle-dependent reflectivity from seismic data. We derive a forward modeling operator which predicts the reflection data without calculating the reflection angles, and verify that it approximately equals to the reflection-angle-dependent wave-equation-based Kirchhoff modeling operator under the assumption that the velocity perturbation is small and the reflection angle is smaller than the critical angle. Based on the proposed modeling operator associated with the adjoint of the angle-dependent wave-equation-based Kirchhoff modeling operator, we reformulate LSRTM as an inverse problem to invert for reflection-angle-dependent reflectivity using a preconditioned conjugate gradient algorithm. The algorithm uses a low-rank filter as the preconditioner to attenuate migration artifacts. Imaging tests on synthetic and field seismic data are used to verify validity and superiority of the proposed method. The tests illustrate that the proposed method can produce the reflection-angle-dependent reflectivity with much higher signal-to-noise ratio, resolution and amplitude fidelity than reverse time migration. Compared with conventional LSRTM, it can produce more focused stacked image when the migration velocity contains errors. Moreover, conventional LSRTM only produces the angle-independent reflectivity, whereas the proposed method has the feasibility to produce the reflection-angle-dependent reflectivity. Chuang Li 0003, Zhaoqi Gao, Feipeng Li, Zhen Li 0016, Jinghuai Gao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Multi-hour and multi-site air quality index forecasting in Beijing using CNN, LSTM, CNN-LSTM, and spatiotemporal clustering
Jiaqiang Liao, Wei Sun 0053, Mingyue Nong, Feipeng Li |
Expert Syst. Appl. | 6 |
| 2020 | Data-Driven Analysis for RFID-Enabled Smart Factory: A Case StudyabstractThe emergence of Internet of Things (IoT) and new manufacturing paradigms have brought greater complexity of massive datasets. Radio frequency identification (RFID), as one of the key IoT technologies, has been used to collect real-time production data to support the manufacturing decision-making in smart factories. The adoption of these technologies results in a large amount of data collection. To extract useful information from this data, this paper utilizes a big data approach to figure out useful insights from RFID-enabled data regarding possible bottlenecks or inefficiencies on the shop floor so as to improve the quality management. Time and quality are the main metrics measured in this paper, where the longest process times, part accuracy percentage, and failure rate are determined for each of the workers (UserIDs) and process types (ProcCodes). Key findings and observations are significant to make advanced decisions in the smart factory by making full use of the RFID captured data. Jiqiang Feng, Feipeng Li, Chen Xu 0004, Ray Y. Zhong |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2019 | A LPSO-SGD algorithm for the Optimization of Convolutional Neural NetworkabstractIn recent years, Convolutional Neural Networks (CNN) perform very well in many complex tasks. When we train CNN, the Stochastic Gradient Descent (SGD) algorithm is widely used to optimize the loss function of CNN. However, SGD algorithm has some disadvantages such as being easy to fall into local optimum and vanishing gradient problems that need to be solved. In this paper, we propose a new hybrid algorithm that aims to tackle the disadvantage mentioned above by combining the advantages of the Lclose Particle Swarm Optimization (LPSO) and SGD algorithm. Particle Swarm Optimization (PSO) is a Global optimization algorithm, but it does not perform very well in optimizing the loss function of the neural network because of the neural network's high dimensional weight parameters and the infinite search area. To take advantage of the excellent global search capability of LPSO and the rapid convergence capability, we design the LPSO-SGD algorithm. In the experimental part, we construct the LeNet-5 deep CNN to classify the MNIST data set and the experimental results demonstrate that the proposed algorithm perform better than standard SGD algorithm. Guixiang Lai, Feipeng Li, Jiqiang Feng |
CEC | 2 |
| 2014 | Subband hybrid feature for multi-stream speech recognitionabstractA subband hybrid (SBH) feature is developed for multi-stream (MS) speech recognition. The fullband speech signal is decomposed into multiple subbands, each covers about 3 Bark along the frequency. Speech signal is analyzed by a high-resolution filterbank of 4 filters/Bark and a low-resolution filterbank of 2 filters/Bark to facilitate the representation of both short-term spectral modulation and long-term temporal modulation within a frequency subband. Experiments on TIMIT corpus for English and RATS corpus for Arabic Levantine show that the SBH feature significantly enhances the amount of information being extracted from individual subbands. The MS system with performance monitor achieves a substantial gain in performance over the single-stream baseline. Feipeng Li |
ICASSP | 1 |
| 2014 | A long, deep and wide artificial neural net for robust speech recognition in unknown noiseabstractA long deep and wide artificial neural net (LDWNN) with multiple ensemble neural nets for individual frequency subbands is proposed for robust speech recognition in unknown noise. It is assumed that the effect of arbitrary additive noise on speech recognition can be approximated by white noise (or speech-shaped noise) of similar level across multiple frequency subbands. The ensemble neural nets are trained in clean and speech-shaped noise at 20, 10, and 5 dB SNR to accommodate noise of different levels, followed by a neural net trained to select the most suitable neural net for optimum information extraction within a frequency subband. The posteriors from multiple frequency subbands are fused by another neural net to give a more reliable estimation. Experimental results show that the subband ensemble net adapts well to unknow noise. Feipeng Li, Phani S. Nidadavolu, Hynek Hermansky |
INTERSPEECH | 1 |
| 2013 | Effect of filter bandwidth and spectral sampling rate of analysis filterbank on automatic phoneme recognitionabstractIn this study we investigate the effect of filter bandwidth and spectral sampling rate of analysis filterbank for speech recognition. Two experiments are conducted to evaluate the performance of an automatic phoneme recognition system on clean speech and speech in noise as the filter bandwidth increases from 0.5 to 3.5 ERB and the spectral resolution changes from 1, 1.5, 2, 3, 4, to 6 samples per Bark. Results indicate that the optimum filter bandwidth varies for different speech sounds at different frequency ranges. A spectral sampling of 4 filters per Bark with the filter bandwidth being ≈ 1 ERB produces the best performance on average. Feipeng Li, Hynek Hermansky |
ICASSP | 1 |
| 2013 | Improvements in language identification on the RATS noisy speech corpus
Jeff Z. Ma, Bing Zhang 0004, Spyridon Matsoukas, Sri Harish Reddy Mallidi, Feipeng Li, Hynek Hermansky |
INTERSPEECH | 5 |
| 2013 | Stream selection and integration in multistream ASR using GMM-based performance monitoringabstractA moderately deep and rather wide artificial neural net is applied in phoneme recognition of noisy speech. The net is formed by first estimating posterior probabilities of phonemes in 21 band-limited streams covering the whole speech spectrum. These 21 band-limited streams are subdivided into three seven band-limited stream subsets, by differently sub-sampling the original 21 band-limited streams. In the second processing stage, all non-empty combinations of seven band-limited streams from each subset are formed as inputs to 127 artificial neural nets that are again trained to yield phoneme posteriors. In this way, 127 × 3 = 381 processing streams are formed. A novel technique for finding the best combination of the resulting 381 parallel processing streams, which uses the likelihood of a single-state Gaussian mixture model of the final classifier output is applied to selecting the most efficient streams. The technique is efficient in phoneme recognition of speech that is corrupted by realistic additive noise. Tetsuji Ogawa, Feipeng Li, Hynek Hermansky |
INTERSPEECH | 2 |
| 2013 | Multi-stream recognition of noisy speech with performance monitoringabstractA prototype multi-stream system with a performance monitor for stream selection is proposed to recognize speech in un-known noise. The speech signal is decomposed into seven band-limited streams. Posterior probabilities of phonemes are estimated by a multi-layer perceptron (MLP) in each of these band-limited streams. Estimated posterior vectors of all 127 combinations (processing streams) of the seven band-limited streams form inputs to a second-stage MLP that esti-mates posterior probabilities of phonemes in each processing stream. A performance monitor is designed to predict the re-liability of individual processing streams based on the outputs from these streams. The top N streams that are least affected by noise are selected and their outputs are averaged to yield the final posterior probability vector used in Viterbi search for the best phoneme sequence. Experimental results show that the proposed technique is effective in dealing with noise. Index Terms: Multi-stream speech recognition, Performance monitoring Ehsan Variani, Feipeng Li, Hynek Hermansky |
INTERSPEECH | 2 |
| 2012 | Phone recognition in critical bands using sub-band temporal modulations
Feipeng Li, Sri Harish Reddy Mallidi, Hynek Hermansky |
INTERSPEECH | 1 |
| 2011 | Manipulation of Consonants in Natural SpeechabstractNatural speech often contains conflicting cues that are characteristic of confusable sounds. For example, the /k/, defined by a mid-frequency burst within 1-2 kHz, may also contain a high-frequency burst above 4 kHz indicative of /ta/, or vice versa. Conflicting cues can cause people to confuse the two sounds in a noisy environment. An efficient way of reducing confusion and improving speech intelligibility in noise is to modify these speech cues. This paper describes a method to manipulate consonant sounds in natural speech, based on our a priori knowledge of perceptual cues of consonants. We demonstrate that: 1) the percept of consonants in natural speech can be controlled through the manipulation of perceptual cues; 2) speech sounds can be made much more robust to noise by removing the conflicting cue and enhancing the target cue. Feipeng Li, Jont B. Allen |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Manipulation of consonants in natural speechabstractSummary form only given - Starting in the 1920s, researchers at AT&T Research characterized speech perception. Until 1950, this work was done by a large group working under Harvey Fletcher, which resulted in the articulation index, an important tool able to predict average speech scores. In the 1950s a dedicated group of researchers at Haskins Labs in NYC attempted to extend these ideas, and then again at MIT under the direction of Ken Stevens, further work was done, on trying to identify the reliable speech cues. Most of this work after 1950 was not successful in finding speech cues, therefore today many consider it impossible. That is, many believe that there is no direct unique mapping from the time-frequency plane to consonant and vowel recognition. For example it has been claimed that context is necessary to successfully identify nonsense consonantvowels. In fact this is not the case. The post 1950 work mostly used synthetic speech. This was a major flaw with all these studies. Also only average results were studied, again a major flaw.In 2007 we carefully measured the consonant error for 20 talkers speaking 16 different consonants, in two types of variable noise. For many consonants, the human performance is well above chance at -20 dB SNR, and at 0 dB SNR, the score is close to 100% for most sounds. The error patterns for individual sounds are quite different from the average. Vowels preform very differently than consonants. The lesson learned is to carefully study token inhomogeneity. The present work is a natural extension of these 1950 studies, but this time we have been successful and have determined the mapping. Using (1) extensive psychoacoustic methods, (2) working with a large data-base (3) of recorded speech sounds, with (4) the newly developed techniques that (5) use a model of the auditory system to (6) predict audible cues in noise, all (7) with a large number of listeners to evaluate the induced confusions, we have precisely identified the acoustic cues for individual utterances and for a large number of consonants. This paper explores the potential use of this new knowledge about perceptual cues of consonant sounds in speech processing. These cues provide deep insight into why Fletcher's articulation index is successful in predicting average "nonsense" speech syllables. Our analysis of a large number of nonsense Consonant-Vowel syllables from the LDC database reveals that natural speech, especially stop consonants, often contain conflicting speech cues that are characteristic of confusable sounds. Through the manipulation of these acoustic cues, one phone (a consonant or vowel sound) is be morphed into another. Meaningful sentences can be morphed into nonsense, or a sentence with a very different meaning. The resulting morphed speech is naturalsounding human speech. These techniques are robust to noise: a weak sound, easily masked by noise, can be converted into a strong one. Results of speech perception experiments on feature-enhanced /ka/ and /ga/ show that any modification of speech cues significantly changes, and can even improve the score in noise, for both normal and hearing-impaired listeners. The implications for ASR will be discussed. Jont B. Allen, Feipeng Li |
ASRU | 2 |