EDBT 2026 Demo / reviewers in the wild / expert
Les E. Atlas
dblp:25/2520 · also Les Atlas
· DBLP profile ↗
87ranked-venue papers
13as first author
1since 2021 · last 2023
0009-0001-9979-5864ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 67 · 8 first-author · 1 since 2021Artificial intelligence and machine learning · 16 · 4 first-authorDatabases, data management, data science and information retrieval · 4Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
6 papers |
Audio and music processing · 97% Image and video coding · 2% Multimedia systems and quality of experience · 1% | |
| Artificial intelligence
7 papers |
Deep learning architectures and training · 96% Question answering and dialogue systems · 2% Learning theory · 1% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 77% Computational science and engineering · 23% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 24 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 2 | 2016 | Full-Capacity Unitary Recurrent Neural Networks · NIPS 2016 Recurrent Networks and NARMA Modeling · NIPS 1991 |
Machine learning › Deep learning architectures and training › recurrent neural network
unitary recurrent neural network |
0.2 | 1 | 2016 | Full-Capacity Unitary Recurrent Neural Networks · NIPS 2016 |
Machine learning › Deep learning architectures and training › training dynamics
vanishing and exploding gradients |
0.2 | 1 | 2016 | Full-Capacity Unitary Recurrent Neural Networks · NIPS 2016 |
Audio and music processing
speech analysis |
0.2 | 2 | 2015 | Speech Analysis With the Strong Uncorrelating Transform · IEEE ACM Trans. Audio Speech Lang. Process. 2015 Applications of positive time-frequency distributions to speech processing · IEEE Trans. Speech Audio Process. 1994 |
Audio and music processing › source separation
blind source separation |
0.2 | 1 | 2015 | Speech Analysis With the Strong Uncorrelating Transform · IEEE ACM Trans. Audio Speech Lang. Process. 2015 |
Audio and music processing › speech processing
voice activity detection |
0.2 | 1 | 2015 | Speech Analysis With the Strong Uncorrelating Transform · IEEE ACM Trans. Audio Speech Lang. Process. 2015 |
Audio and music processing › source separation
single-channel source separation |
0.1 | 1 | 2011 | Single-Channel Source Separation Using Complex Matrix Factorization · IEEE ACM Trans. Audio Speech Lang. Process. 2011 |
Audio and music processing
source separation |
0.1 | 1 | 2011 | Single-Channel Source Separation Using Complex Matrix Factorization · IEEE ACM Trans. Audio Speech Lang. Process. 2011 |
Bioinformatics and computational biology › systems biology › multiscale modeling
multiscale physiological modeling |
0.1 | 1 | 2006 | Strategies and Tactics in Multiscale Modeling of Cell-to-Organ Systems · Proc. IEEE 2006 |
Audio and music processing
speech recognition |
0.0 | 1 | 2011 | Single-Channel Source Separation Using Complex Matrix Factorization · IEEE ACM Trans. Audio Speech Lang. Process. 2011 |
Audio and music processing
time-frequency analysis |
0.0 | 2 | 1996 | Applications of time-frequency analysis to signals from manufacturing and machine monitoring sensors · Proc. IEEE 1996 Applications of positive time-frequency distributions to speech processing · IEEE Trans. Speech Audio Process. 1994 |
Natural language and speech › Question answering and dialogue systems
dialogue modeling |
0.0 | 1 | 1995 | The challenge of spoken language systems: research directions for the nineties · IEEE Trans. Speech Audio Process. 1995 |
Image and video coding
image compression |
0.0 | 1 | 1994 | Index assignment for progressive transmission of full-search vector quantization · IEEE Trans. Image Process. 1994 |
Multimedia systems and quality of experience › image transmission
progressive image transmission |
0.0 | 1 | 1994 | Index assignment for progressive transmission of full-search vector quantization · IEEE Trans. Image Process. 1994 |
Image and video coding › quantization
vector quantization |
0.0 | 1 | 1994 | Index assignment for progressive transmission of full-search vector quantization · IEEE Trans. Image Process. 1994 |
Machine learning › Deep learning architectures and training › feedforward neural network
backpropagation networks |
0.0 | 1 | 1989 | Performance Comparisons Between Backpropagation Networks and Classification Trees on Three Real-World Applications · NIPS 1989 |
Machine learning › Learning theory › query learning
selective sampling |
0.0 | 1 | 1989 | Training Connectionist Networks with Queries and Selective Sampling · NIPS 1989 |
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatio-temporal neural networks |
0.0 | 1 | 1987 | An Artificial Neural Network for Spatio-Temporal Bipolar Patterns: Application to Phoneme Classification · NIPS 1987 |
Natural language and speech › Speech recognition and synthesis
speech synthesis |
0.0 | 1 | 1995 | The challenge of spoken language systems: research directions for the nineties · IEEE Trans. Speech Audio Process. 1995 |
Natural language and speech › Language models and text generation › text generation
text response generation |
0.0 | 1 | 1995 | The challenge of spoken language systems: research directions for the nineties · IEEE Trans. Speech Audio Process. 1995 |
Audio and music processing › time-frequency analysis
spectrogram |
0.0 | 1 | 1994 | Applications of positive time-frequency distributions to speech processing · IEEE Trans. Speech Audio Process. 1994 |
Data mining › predictive modeling
classification |
0.0 | 1 | 1989 | Performance Comparisons Between Backpropagation Networks and Classification Trees on Three Real-World Applications · NIPS 1989 |
Data mining › predictive modeling › classification
decision tree learning |
0.0 | 1 | 1989 | Performance Comparisons Between Backpropagation Networks and Classification Trees on Three Real-World Applications · NIPS 1989 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
phoneme recognition |
0.0 | 1 | 1987 | An Artificial Neural Network for Spatio-Temporal Bipolar Patterns: Application to Phoneme Classification · NIPS 1987 |
Methods — techniques the papers use, named apart from their topics
unitary parameterization · 0.5multiplicative gradient step · 0.5linear gaussian dynamical model · 0.3factor analysis · 0.3expectation-maximization · 0.3strong uncorrelating transform · 0.2independent component analysis · 0.2phase estimation · 0.1nonnegative matrix factorization · 0.1complex matrix factorization · 0.1model decomposition · 0.1mathematical modeling · 0.1error recognition · 0.1time-frequency analysis · 0.0robust speech recognition · 0.0automatic training and adaptation · 0.0iterative time-frequency distribution estimation · 0.0backpropagation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Estimating and Analyzing Neural Information flow using Signal Processing on GraphsabstractCorrelating neural communication in brain networks with behavior and cognition can provide fundamental insights into the functionality of both healthy and diseased brains. We demonstrate how communication in the brain can be estimated from recorded neural activity using concepts from graph signal processing. The communication is modeled as a flow signals on the edges of a graph and naturally arises from a graph diffusion process. We apply the diffusion model to micro-electrocorticography (ECoG) recordings from sensorimo-tor cortex of two non-human primates to estimate the neural communication flow during excitatory optogenetics. Comparisons with a baseline model demonstrate that adding the neural flow can improve ECoG predictions. Finally, we demonstrate how the neural flow can be decomposed into a gradient and rotational component and show that the gradient component depends on the location of stimulation. This technique, for the first time, offers the opportunity to study neural communication on an unprecedented spatiotemporal scale. Felix Schwock, Julien A. Bloch, Les E. Atlas, Shima Abadi, Azadeh Yazdan-Shahmorad |
ICASSP | 3 |
| 2018 | Regression Factor Analysis With an Application to Continuous HRIR MeasurementabstractLinear time-varying (LTV) regression models play a central role in input-output analysis of many real-world dynamical systems. Most existing models for such systems consider either a switching dynamics with a discrete latent variable driving the system or a linear dynamics directly applied to the regression coefficients. These models usually fall short of capturing structural regularities or the total variability in the dynamics of the system. Addressing these issues in this paper, we propose a method to parametrize joint variations of regression coefficients in LTV systems with continuous latent variables based on factor analysis technique, i.e., regression factor analysis . By using a linear Gaussian dynamical model for the time evolution of factor weights, our model constrains the dynamics of the process. Our inference scheme takes advantage of the expectation-maximization algorithm to estimate the model parameters. We show how our proposed algorithm can be utilized in a key application for fast continuous measurement of personalized head-related impulse responses. Majid Mirbagheri, Les E. Atlas, Adrian K. C. Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Building recurrent networks by unfolding iterative thresholding for sequential sparse recoveryabstractHistorically, sparse methods and neural networks, particularly modern deep learning methods, have been relatively disparate areas. Sparse methods are typically used for signal enhancement, compression, and recovery, usually in an unsupervised framework, while neural networks commonly rely on a supervised training set. In this paper, we use the specific problem of sequential sparse recovery, which models a sequence of observations over time using a sequence of sparse coefficients, to show how algorithms for sparse modeling can be combined with supervised deep learning to improve sparse recovery. Specifically, we show that the iterative soft-thresholding algorithm (ISTA) for sequential sparse recovery corresponds to a stacked recurrent neural network (RNN) under specific architecture and parameter constraints. Then we demonstrate the benefit of training this RNN with backpropagation using supervised data for the task of column-wise compressive sensing of images. This training corresponds to adaptation of the original iterative thresholding algorithm and its parameters. Thus, we show by example that sparse modeling can provide a rich source of principled and structured deep network architectures that can be trained to improve performance on specific tasks. Scott Wisdom, Thomas Powers, James W. Pitton, Les E. Atlas |
ICASSP | 4 |
| 2016 | Constrained robust submodular sensor selection with applications to multistatic sonar arrays
Thomas Powers, Jeff A. Bilmes, David W. Krout, Les E. Atlas |
FUSION | 4 |
| 2016 | An alternative approach for auditory attention tracking using single-trial EEGabstractAuditory selective attention plays a central role in the human capacity to reliably process complex sounds in multi-source environments. Stimulus reconstruction has been widely used for the investigation of selective auditory attention using multichannel electroencephalogra-phy (EEG). In particular, the influence of attention on sound representations in the brain has been modeled by linear time-variant filters and have been used to track the attentional state of individuals in multi-source environments. Detection of auditory attention is of interest and is important in the study of attention-related disorders and has potential application in the hearing aid and advertising industries. In analogy with the rake receiver from wireless communications, we propose a new strategy, adapting principles from minimum variance beamforming, to reconstruct stimuli for decoding the at-tentional state of listeners in a competing speaker environment. We show through experiments with real electrophysiological data how decoding accuracies can be improved using our proposed scheme. Bradley Ekin, Les E. Atlas, Majid Mirbagheri, Adrian K. C. Lee |
ICASSP | 2 |
| 2016 | Full-Capacity Unitary Recurrent Neural NetworksabstractRecurrent neural networks are powerful models for processing sequential data, but they are generally plagued by vanishing and exploding gradient problems. Unitary recurrent neural networks (uRNNs), which use unitary recurrence matrices, have recently been proposed as a means to avoid these issues. However, in previous experiments, the recurrence matrices were restricted to be a product of parameterized unitary matrices, and an open question remains: when does such a parameterization fail to represent all unitary matrices, and how does this restricted representational capacity limit what can be learned? To address this question, we propose full-capacity uRNNs that optimize their recurrence matrix over all unitary matrices, leading to significantly improved performance over uRNNs that use a restricted-capacity recurrence matrix. Our contribution consists of two main components. First, we provide a theoretical argument to determine if a unitary parameterization has restricted capacity. Using this argument, we show that a recently proposed unitary parameterization has restricted capacity for hidden state dimension greater than 7. Second,we show how a complete, full-capacity unitary recurrence matrix can be optimized over the differentiable manifold of unitary matrices. The resulting multiplicative gradient step is very simple and does not require gradient clipping or learning rate adaptation. We confirm the utility of our claims by empirically evaluating our new full-capacity uRNNs on both synthetic and natural data, achieving superior performance compared to both LSTMs and the original restricted-capacity uRNNs. Scott Wisdom, Thomas Powers, John R. Hershey, Jonathan Le Roux, Les E. Atlas |
NIPS | 5 |
| 2015 | Sensor selection from independence graphs using submodularity
Thomas Powers, David W. Krout, Les E. Atlas |
FUSION | 3 |
| 2015 | Voice activity detection using subband noncircularityabstractMany voice activity detection (VAD) systems use the magnitude of complex-valued spectral representations. However, using only the magnitude often does not fully characterize the statistical behavior of the complex values. We present two novel methods for performing VAD on single- and dual-channel audio that do completely account for the second-order statistical behavior of complex data. Our methods exploit the second-order noncircularity (also known as impropriety) of complex subbands of speech and noise. Since speech tends to be more improper than noise, higher impropriety suggests speech activity. Our single-channel method is blind in the sense that it is unsupervised and, unlike many VAD systems, does not rely on non-speech periods for noise parameter estimation. Our methods achieve improved performance over other state-of-the-art magnitude-based VADs on the QUT-NOISE-TIMIT corpus, which indicates that impropriety is a compelling new feature for voice activity detection. Scott Wisdom, Greg Okopal, Les E. Atlas, James W. Pitton |
ICASSP | 3 |
| 2015 | Flexible tracking of auditory attention
Majid Mirbagheri, Bradley Ekin, Les E. Atlas, Adrian K. C. Lee |
INTERSPEECH | 3 |
| 2015 | Speech Analysis With the Strong Uncorrelating TransformabstractThe strong uncorrelating transform (SUT) provides estimates of independent components from linear mixtures using only second-order information, provided that the components have unique circularity coefficients. We propose a processing framework for generating complex-valued subbands from real-valued mixtures of speech and noise where the objective is to control the likely values of the sample circularity coefficients of the underlying speech and noise components in each subband. We show how several processing parameters affect the noncircularity of speech-like and noise components in the subband, ultimately informing parameter choices that allow for estimation of each of the components in a subband using the SUT. Additionally, because the speech and noise components will have unique sample circularity coefficients, this statistic can be used to identify time-frequency regions that contain voiced speech. We give an example of the recovery of the circularity coefficients of a real speech signal from a two-channel noisy mixture at -25 dB SNR, which demonstrates how the estimates of noncircularity can reveal the time-frequency structure of a speech signal in very high levels of noise. Finally, we present the results of a voice activity detection (VAD) experiment showing that two new circularity-based statistics, one of which is derived from the SUT processing, can achieve improved performance over state-of-the-art VADs in real-world recordings of noise. Greg Okopal, Scott Wisdom, Les E. Atlas |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Track to track fusion: PACsim data set
Thomas Powers, Les E. Atlas, Evan Hanusa, David W. Krout |
FUSION | 2 |
| 2014 | Extending coherence time for analysis of modulated random processesabstractIn this paper, we relax a commonly-used assumption about a class of nonstationary random processes composed of modulated wide-sense stationary random processes: that the fundamental frequency of the modulator is stationary within the analysis window. To compensate for the relaxation of this assumption, we define the generalized DEMON (“demodulated noise”) spectrum representing modulation frequency, which we use to increase the coherence time of such signals. Increased coherence time means longer analysis windows, which provides higher SNR estimators. We use the example of detection on both synthetic and real-world passive sonar signals to demonstrate this increase. Scott Wisdom, Les E. Atlas, James Pittore |
ICASSP | 2 |
| 2013 | Complementary envelope estimation for frequency-modulated random signalsabstractThe Hilbert envelope is a well-known indicator for amplitude modulation in subband time series. However, frequency modulation estimators typically impose temporal smoothness constraints that limit their general use for stochastic signals. We introduce the complementary envelope as a direct indicator of stochastic phase coherence and frequency modulation. Rooted in the second-order statistics of randomly-phased sinusoids, the complementary envelope is distinct from the Hilbert envelope, and can be estimated easily and separately. We propose a new complementary envelope estimator based on modified multitaper spectral analysis, and use it to reveal previously unseen FM-like behavior in ship propeller noise. Pascal Clark, Ivars P. Kirsteins, Les E. Atlas |
ICASSP | 3 |
| 2012 | Existence and estimation of impropriety in real rhythmic signalsabstractImpropriety in complex signal processing has been studied and used primarily in a communications context, but also in some cases where complex signals are generated by adding real signals in quadrature. We discuss the meaning of impropriety, and the associated use of complementary statistics, when a real-valued random process is improper in the frequency domain. Through the use of modulation signal models, spectral impropriety can be connected explicitly to the frequency and phase of components belonging to a periodic, or more generally rhythmic, modulator waveform. We give theoretical signal models and provide an example of complementary analysis on underwater propeller noise from a merchant ship. Pascal Clark, Ivars P. Kirsteins, Les E. Atlas |
ICASSP | 3 |
| 2012 | Disaggregated water sensing from a single, pressure-based sensor: An extended analysis of HydroSense using staged experiments
Eric C. Larson, Jon Froehlich, Tim Campbell, Conor Haggerty, Les E. Atlas, James Fogarty, Shwetak N. Patel |
Pervasive Mob. Comput. | 5 |
| 2011 | A novel approach using modulation features for multiphone-based speech recognitionabstractRecent advances in coherent and convex demodulation have proven useful for analyzing and modifying the low-frequency envelope structure of speech. This paper reports the application of both methods, referred to here as bandwidth-constrained demodulation, to large-scale speech recognition in the form of new feature representations. Modulation-based features yielded measurable improvement when included as complementary sources of information with a baseline recognizer. Furthermore, both sets of demodulation features showed promise for outperforming the conventional Hilbert envelope method which underlies most modern speech recognition features. These experimental results show the potential for further development in feature representations based on recently-developed bandwidth-constrained modulation signal models. Pascal Clark, Gregory Sell, Les E. Atlas |
ICASSP | 3 |
| 2011 | Speech recognitionwith segmental conditional random fields: A summary of the JHU CLSP 2010 Summer WorkshopabstractThis paper summarizes the 2010 CLSP Summer Workshop on speech recognition at Johns Hopkins University. The key theme of the workshop was to improve on state-of-the-art speech recognition systems by using Segmental Conditional Random Fields (SCRFs) to integrate multiple types of information. This approach uses a state of-the-art baseline as a springboard from which to add a suite of novel features including ones derived from acoustic templates, deep neural net phoneme detections, duration models, modulation features, and whole word point-process models. The SCRF framework is able to appropriately weight these different information sources to produce significant gains on both die Broadcast News and Wall Street Journal tasks. Geoffrey Zweig, Patrick Nguyen, Dirk Van Compernolle, Kris Demuynck, Les E. Atlas, Pascal Clark, Gregory Sell, Meihong Wang, Fei Sha, Hynek Hermansky, Damianos Karakos, Aren Jansen, Samuel Thomas 0001, Sivaram G. S. V. S., Samuel R. Bowman, Justine T. Kao |
ICASSP | 5 |
| 2011 | Single-Channel Source Separation Using Complex Matrix FactorizationabstractNonnegative matrix factorization is gaining popularity in speech and audio processing applications. Performing nonnegative matrix factorization on a complex-valued short-time Fourier transform, however, makes assumptions on the signal, such as additivity in the magnitude domain, potentially degrading the results. One application where these assumptions can cause a problem is in single-channel source separation of overlapping speech. In this paper, we present how this problem can be solved by incorporating phase estimation via complex matrix factorization. Another challenge in source separation is how to select reconstruction bases for optimal separation. In this paper, we compare the most common method with a new, simpler method of finding bases that does not share many of the challenges of the current, established method. The paper will conclude by comparing nonnegative with complex matrix factorization as well as the previous and new methods for finding bases on the task of automatic speech recognition of single-channel two-talker overlapping speech. Brian King, Les E. Atlas |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2010 | Multiband analysis for colored amplitude-modulated ship noiseabstractPropeller radiated cavitation noise is broadband yet audibly rhythmic, taking on the characteristics of amplitude-modulated noise. The study of this process is important because the shaft and blade rates, as well as other identifying features of the ship, can be inferred from the cavitation signal envelope. Unlike the conventional method for estimating the modulation frequency, this paper proposes a multirate subband method applicable to colored signals, and shows the conditions under which it is optimal in the sense of maximum likelihood. Application to actual ship noise data confirms the utility of the new method, yet also reveals previously unobserved signal dynamics, namely acoustic frequency-dependent phase shifts and amplitude variation in the modulation spectrum. Pascal Clark, Ivars P. Kirsteins, Les E. Atlas |
ICASSP | 3 |
| 2010 | Cepstral mean based speech source discriminationabstractThis paper presents and compares methods for discrimination between speech from a broadcast audio device - like a television, radio, or GPS receiver - and live speech in the same acoustic environment. A solution to this discrimination problem has direct application wherever the audio from such a device interferes with voice recognition, verification, or transcription tasks. The methods and theory applied also have potential applications in multimedia and speaker segmentation, as well as in speaker verification. This paper presents a new use of the cepstral mean as an estimator of the linear time-invariant response of a “speaker” - either broadcast or live - over a relatively long time window. The problem is framed in terms of traditional speaker verification, but with two classes of speakers. This method is tested on five different data sets and the results compared for different feature sets, training methods, and window lengths. Adam Greenhall, Les E. Atlas |
ICASSP | 2 |
| 2010 | Single-channel source separation using simplified-training complex matrix factorizationabstractAlthough the task seems trivial for human listeners, research in automating source separation still lags far behind human performance and is especially difficult for single-channel signals. One of the latest and most promising methods of single-channel source separation is non-negative matrix factorization, which works by synthesizing signals from a learned set of bases for each source. In this paper, we present a new method of creating these learned sets of bases used in the matrix factorization technique for single-channel source separation. This new method does not suffer the complication of choosing an optimal number of bases as in previous methods. In addition, this paper further explores the new method of complex matrix factorization and compares its performance to non-negative, real matrix factorization for automatic speech recognition of two-talker mixtures. Brian King, Les E. Atlas |
ICASSP | 2 |
| 2010 | Harmonic coherent demodulation for improving sound coding in cochlear implantsabstractThe paper presents a promising application of the coherent demodulation technique to improving hearing with cochlear implants. It has been a challenge to encode temporal fine structure (TFS) in cochlear implants for better speech and music perception. We propose a pitch-synchronized coherent demodulation strategy-Harmonic Single Sideband Encoder (HSSE)-to efficiently encode TFS cues. A dynamic carrier estimation approach based on harmonic detection was chosen to enhance temporal pitch coding in cochlear implants. Acoustic analysis results showed that the HSSE strategy preserves temporal fine structure cues including fundamental frequency (F0) information as well. The coding of temporal fine structure cues in cochlear implants could be potentially improved with the F0-synchronized coherent demodulation. Kaibao Nie, Les E. Atlas, Jay T. Rubinstein |
ICASSP | 3 |
| 2010 | Some Properties of an Empirical Mode Type Signal Decomposition AlgorithmabstractThe empirical mode decomposition (EMD) has seen widespread use for analysis of nonlinear and nonstationary time-series. Despite some practical success, it lacks a firm theoretical foundation. This work addresses two important theoretical properties. The original EMD algorithm is slightly modified, in a way that facilitates this analysis. For periodic, band-limited, signals the convergence and time scale separation of the algorithm are proved. Stephen D. Hawley, Les E. Atlas, Howard Jay Chizeck |
IEEE Signal Process. Lett. | 2 |
| 2009 | A sum-of-products model for effective coherent modulation filteringabstractModulation filtering is a technique for filtering slowly-varying envelopes of frequency subbands of a nonstationary signal, ideally without affecting the signal's phase and fine-structure. Coherent modulation filtering is a promising subtype of such techniques where subband envelopes are determined through demodulation of the subband signal with a coherently detected subband carrier. In this paper we demonstrate how modulation filtering, when done coherently, is far more effective than standard incoherent methods. We show that empirical results can be made to be almost ideal, and significantly better than previous coherent attempts, as long as fine-structure information is retained as side information and the filterbank reduces subband interference. Pascal Clark, Les E. Atlas |
ICASSP | 2 |
| 2008 | Modulation decompositions for the interpolation of long gaps in acoustic signalsabstractThis paper presents a modulation-based reconstruction method for audio signals across long gaps of missing samples. We use LTI filterbanks followed by a multiplicative model that decomposes subbands into constituent modulators and carriers. This processing separates slowly-varying envelopes, or modulators, from the high-frequency fine structure of the carriers. Since modulators can be downsampled, this decomposition allows faster gap reconstruction when applying standard interpolation algorithms on the downsampled modulators, particularly when interpolation requires matrix inversion. Pascal Clark, Les E. Atlas |
ICASSP | 2 |
| 2008 | Some properties of an empirical mode type signal decomposition algorithmabstractThe empirical mode decomposition (EMD) has seen widespread use for analysis of nonlinear and nonstationary time-series. Despite some practical success, it lacks a firm theoretical foundation. This work addresses this. The original EMD algorithm is slightly modified, in a way that facilitates its analysis. We prove three theorems that give conditions for the convergence and time scale separation of the EMD for stationary, band-limited signals. Stephen D. Hawley, Les E. Atlas, Howard Jay Chizeck |
ICASSP | 2 |
| 2008 | Coherent modulation filtering for speechabstractModulation filtering ideally offers a new approach to modifying the dynamics of non-stationary signals, such as speech. In this paper, a new type of coherent modulation analysis and filtering method is proposed. The new method consists of two essential parts — an instantaneous frequency estimator based on conditional mean frequency, which is used for coherent modulation analysis, and a multi-component decomposition based on spectrogram peak tracking, which is used to separate multiple modulation components in signals. An important modulation filtering property, frequency shift invariance, is achieved with the new proposed method. Les E. Atlas |
ICASSP | 2 |
| 2008 | Air-coupled ultrasound time-of-flight estimation for shipping container cargo verificationabstractThe falling cost of embedded sensor systems has opened up the possibility of in-transit air-coupled ultrasound interior imaging of shipping container cargo. This new technology could allow for real-time cargo integrity verification and improved shipping and transportation security. This paper takes the initial steps in developing a temperature invariant descriptor by comparing three time-of-flight (TOF) estimation methods from the literature using in situ data. The methods compared are matched filter, leading edge envelope line fit, and a simple envelope half-peak intercept. The results show that the simple half-peak intercept method provided the most consistent TOF estimate with respect to overlapping pulse data. Patrick McVittie, Les E. Atlas |
ICASSP | 2 |
| 2008 | Single sideband encoder for music coding in cochlear implantsabstractThe restoration of melody perception is a key remaining challenge in cochlear implants. We propose a new sound coding strategy that converts an audio signal into time-varying electrically stimulating pulse trains. A sound is first split into several frequency subbands and each subband signal is coherently downward shifted to a low-frequency base band, similar to demodulation used in single sideband (SSB) radios. These resulting coherent envelope signals have Hermitian symmetric frequency spectrums and are thus real-valued. A peak detector in each subband further converts the coherent envelopes into rate-varying and interleaved pulse trains. Acoustic simulations of cochlear implants with normal hearing listeners showed significant improvement in melody recognition over the most common stimulation approach used in cochlear implants. Kaibao Nie, Les E. Atlas, Jay T. Rubinstein |
ICASSP | 2 |
| 2008 | Target talker enhancement in hearing devicesabstractWe describe a novel coherent modulation filtering technique for single channel target talker enhancement in the presence of interfering talkers. For this technique, we have expanded our previous work on coherent modulation filtering with a carrier estimator that is more robust to speech from interfering talkers, and a modulation filter that operates on shorter time-scales. We have evaluated the technique in a subjective listening test, which indicates that the novel target talker enhancement technique achieves a moderate improvement in speech reception. We summarize our observations on single channel target talker enhancement and conclude with directions for further research. Steven M. Schimmel, Les E. Atlas |
ICASSP | 2 |
| 2007 | Feasibility of Single Channel Speaker Separation Based on Modulation Frequency AnalysisabstractWe explore the use of the modulation frequency domain for single channel speaker separation. We discuss features of the modulation spectrogram of speech signals that suggest that multiple speakers are highly separable in this space. In a preliminary experiment, we separate a target speaker from an interfering speaker by manually masking out modulation spectral features of the interferer. We extend this experiment into a new automatic speaker separation algorithm, and show that it achieves an acceptable level of separation. The new algorithm only needs a rough estimate of the target speaker's pitch range. Steven M. Schimmel, Les E. Atlas, Kaibao Nie |
ICASSP (4) | 2 |
| 2007 | Beamforming Alternatives for Multi-Channel Transient Acoustic Event ClassificationabstractSignals acquired through a microphone array are typically beamformed to combine channels and improve the signal-to-noise ratio (SNR). However, it has been previously shown that alternative methods for handling multi-channel systems can outperform beamforming for speech recognition applications. In this paper, we implemented a comprehensive set of classification tests using multiple classifiers and feature extraction techniques to determine whether the alternative methods generalize beyond speech recognition applications. We show that applying the alternative methods (in a slightly simpler form) outperforms beamforming when used for classifying transient acoustic projectile weapon signals. Furthermore, an additional technique is introduced which outperforms both beamforming and previously proposed alternatives in certain classification scenarios. For the majority of classification tests, the improvements seen through the use of these alternative methods are statistically significant. Brandon Smith, Les E. Atlas, Maya R. Gupta |
ICASSP (2) | 2 |
| 2006 | Frequency Reassignment for Coherent Modulation FilteringabstractModulation filtering is a technique for filtering slowly-varying envelopes of frequency subbands of a signal, without affecting the signal's phase and fine-structure. Coherent modulation filtering is a promising subtype of such techniques where subband envelopes are determined through demodulation of the subband signal with a coherently detected subband carrier. In this paper we propose a coherent modulation filtering technique that detects the carriers using the frequency reassignment (FR) operator from time-frequency reassignment. We show how this technique avoids the use of finite differences in the computation of instantaneous frequency (IF), and that it estimates IF more accurately than a past technique as a result. We confirm that the FR-enhanced technique retains the desirable modulation filtering properties (superposition and the preservation of zero-crossings) and show that it performs better on the same single-channel music source separation task than the past technique Steven M. Schimmel, Kelly Fitz, Les E. Atlas |
ICASSP (5) | 3 |
| 2006 | Enhanced Modulation Spectrum Using Space-Time Averaging For In-Building Acoustic Signature IdentificationabstractFor most buildings, virtually all subsystems such as fans, generators, and motors generate acoustic energy. This acoustic energy can weakly penetrate walls and pass through hallways and conduits. The propagation paths are complex and, given the typically low energy of received acoustic signals, present challenges for the detection and classification of the subsystems. While conventional approaches would not be expected to work under these difficult conditions, a key observation can be made: these types of subsystems produce line spectra which consist of harmonics of a fundamental frequency. Acoustic propagation effects then strongly affect the relative energy of these harmonics. Modulation spectra, which make use of the frequency spacing instead of the relative energy of the harmonics, are especially insensitive to these frequency-dependent acoustic attenuation affects. When combined with temporal averaging and 3-dimensional spatial (over an array of acoustic sensors) processing, enhanced modulation spectra offer a new approach to the detection and classification building subsystems which produce sound. Somsak Sukittanon, Les E. Atlas, Stephen G. Dame |
ICASSP (3) | 2 |
| 2006 | Strategies and Tactics in Multiscale Modeling of Cell-to-Organ SystemsabstractModeling is essential to integrating knowledge of human physiology. Comprehensive self-consistent descriptions expressed in quantitative mathematical form define working hypotheses in testable and reproducible form, and though such models are always "wrong" in the sense of being incomplete or partly incorrect, they provide a means of understanding a system and improving that understanding. Physiological systems, and models of them, encompass different levels of complexity. The lowest levels concern gene signaling and the regulation of transcription and translation, then biophysical and biochemical events at the protein level, and extend through the levels of cells, tissues and organs all the way to descriptions of integrated systems behavior. The highest levels of organization represent the dynamically varying interactions of billions of cells. Models of such systems are necessarily simplified to minimize computation and to emphasize the key factors defining system behavior; different model forms are thus often used to represent a system in different ways. Each simplification of lower level complicated function reduces the range of accurate operability at the higher level model, reducing robustness, the ability to respond correctly to dynamic changes in conditions. When conditions change so that the complexity reduction has resulted in the solution departing from the range of validity, detecting the deviation is critical, and requires special methods to enforce adapting the model formulation to alternative reduced-form modules or decomposing the reduced-form aggregates to the more detailed lower level modules to maintain appropriate behavior. The processes of error recognition, and of mapping between different levels of model complexity and shifting the levels of complexity of models in response to changing conditions, are essential for adaptive modeling and computer simulation of large-scale systems in reasonable time. James B. Bassingthwaighte, Howard Jay Chizeck, Les E. Atlas |
Proc. IEEE | 3 |
| 2005 | Coherent modulation spectral filtering for single-channel music source separationabstractModulation spectral filtering, if effective and distortion-free, would offer a new tool for signal modification. Previous approaches to modulation spectral filtering, which made use of incoherent detection of real and positive modulating envelopes for each frequency sub-band, have not offered effective and distortion-free signal modification. Based upon a recent observation that the modulating envelopes are potentially complex, coherent detection is instead proposed. Details are provided for accurate carrier estimation, and tests on both synthetic signals and music, show that modulation filtering is indeed distortion-free. The coherent modulation filtering method is applied to single-channel music sound source separation with promising results for music and other signal separation and modification applications. Les E. Atlas, Christian Janssen |
ICASSP (4) | 1 |
| 2005 | Properties for modulation spectral filteringabstractA two-dimensional representation, the "modulation spectrum", where the modulation frequency exists jointly with a regular Fourier frequency or other filter channel index, has previously been investigated. Accurate modulation filters offer, for example, new approaches for signal separation and noise reduction. However, a filtering operation on modulation frequency components has yet to be carefully defined. Most previous studies on modulation filtering assumed that the amplitude modulation envelope is real and non-negative, which has recently been shown to be incorrect. Distortions appear when the non-negative envelope assumption fails. Beginning with a more appropriate envelope assumption that allows the envelope to go negative, we propose three properties which modulation filtering systems should satisfy. Any modulation filtering method which satisfies these properties yields distortion-free results. An implementation of modulation filtering, based on a short-time Fourier transform followed by independent coherent demodulation for each frequency channel, is then proposed. Satisfaction of the properties is confirmed and an example result of modulation filtering on a speech signal is illustrated. Les E. Atlas |
ICASSP (4) | 2 |
| 2005 | Use of modulation spectra for representation and classification of acoustic transients from sniper fireabstractThere are many applications for classification of acoustic transients produced by supersonic projectile fire. Analysis of existing models for such transients suggests they have properties which may be well-captured by the transform of a signal into joint acoustic and modulation frequency: a modulation spectral representation. Simple features are extracted from this representation which enables successful use in such an important classification application. Lane M. D. Owsley, Les E. Atlas, Chad Heinemann |
ICASSP (4) | 2 |
| 2005 | Coherent Envelope Detection for Modulation Filtering of SpeechabstractModulation filtering, which has been previously described as several related approaches to achieve modification of speech temporal dynamics, is shown to be less effective than intended. In particular, past Hilbert envelope approaches generate distortion which spreads across frequency sub-bands and modulation rejection is far from the amount intended. The source of this distortion is analyzed and a solution, based upon coherent envelope detection in each sub-band is proposed. This coherent approach is shown to be substantially more effective than conventional incoherent approaches on speech samples. Steven M. Schimmel, Les E. Atlas |
ICASSP (1) | 2 |
| 2005 | Improved modulation spectrum through multi-scale modulation frequency decompositionabstractThe modulation spectrum is a promising method to incorporate dynamic information in pattern classification. It contains important cues about the nonstationary content of a signal and yields complementary improvements when it is combined with conventional features derived from short-term analysis. Many prior modulation spectrum approaches are based on uniform modulation frequency decomposition. The drawbacks of these approaches are high dimensionality and a lack of a connection to human perception of modulation. The paper presents multi-scale modulation frequency decomposition and shows an improvement over the standard modulation spectrum in a digital communication signal classification task. Features derived from this representation provide lower classification error rates than those from a constant-bandwidth modulation spectrum, whether used alone or in combination with short-term features. Somsak Sukittanon, Les E. Atlas, James W. Pitton, Karim Filali |
ICASSP (4) | 2 |
| 2005 | Data sampling for improved speech recognizer training
Takahiro Shinozaki, Mari Ostendorf, Les E. Atlas |
INTERSPEECH | 3 |
| 2004 | Homomorphic modulation spectraabstractPhysical evidence points to the importance of a concept called "modulation frequency". This dimension exists jointly with standard Fourier or acoustic frequency. Thus, akin to other time-varying analysis, we seek a two-dimensional representation, the "modulation spectrum", where the first dimension is the well-known acoustic frequency and the second dimension is modulation frequency. We describe some deficiencies in previous discussions of this concept, and then address those deficiencies via a homomorphic approach. We also reduce previous difficulties in homomorphic demultiplication by integrating this processing into modulation spectra and, in particular, show how assumption of analytic and relatively narrowband sub-bands allows more accurate and practical use of homomorphic demultiplication. Lastly, we show how an unambiguous demultiplication concept is only consistent with complex modulator envelopes. The assumption of complex envelopes is necessary for accurate modulation spectral analysis and filtering. Les E. Atlas, Jeffrey Thompson |
ICASSP (2) | 1 |
| 2003 | Strategies for improving audible quality and speech recognition accuracy of reverberant speechabstractWe showed previously (Gillespie and Atlas, Int. Conf. on Acoustics, Speech, and Sig. Processing, 2002) that penalizing long-term reverberation energy is more effective than maximizing the signal-to-reverberation ratio (SRR) for improving audible quality and automatic speech recognition (ASR) accuracy. Using this knowledge, we propose a blind approach to speech dereverberation that reduces the length of the equalized speaker-to-receiver impulse response. The approach reduces the long-term correlation in the linear prediction (LP) residual of reverberant speech. We show that this approach improves both the audible quality (measured with subjective listening tests) and ASR accuracy (measured with two commercial ASR systems) of reverberant speech. Bradford W. Gillespie, Les E. Atlas |
ICASSP (1) | 2 |
| 2003 | Time-variant least squares harmonic modelingabstractAn algorithm for harmonic decomposition of time-variant signals is derived from a least squares harmonic (LSH) technique. The estimates of harmonic amplitudes and phases are formulated as the solution of a set of linear equations which minimizing mean square error; the signal frequency is modeled by a linear or quadratic polynomial and obtained via a local search over polynomial coefficients. An initial estimate of signal frequency is necessary to reduce computation time. This method is capable of producing accurate and robust harmonic estimation in low SNR situations. We show applicability to high accuracy speech pitch and heart sound beat epoch estimation. Les E. Atlas |
ICASSP (2) | 2 |
| 2003 | Non-stationary signal classification using joint frequency analysisabstractTime-varying short-term spectral estimates have been successfully applied in many classification tasks. However, they are still insufficient for many non-stationary signals where time-varying information is useful. We propose to improve the deficiencies of current short-term feature analysis by adding information to describe the time-varying behavior of the signals. Our proposed method, which is motivated by the human auditory system, can be applied to several non-stationary signal types. Real world communication signals were used for experimental verification. These experimental results, assessed with a conventional probabilistic classifier, showed significant improvement when the new features were added to short-term spectral estimates. Somsak Sukittanon, Les E. Atlas, James W. Pitton, Jack McLaughlin |
ICASSP (6) | 2 |
| 2003 | A non-uniform modulation transform for audio coding with increased time resolutionabstractPerceptual audio coders exploit two properties to achieve coding gain: perceptual irrelevancy and source redundancy. Recently, a two-dimensional modulation transform was introduced which efficiently extracts perceptual irrelevancy and source redundancy not accessible in a one-dimensional transform. In this paper, we propose an alternative modulation transform design with an octave-band non-uniform modulation dimension. This non-uniform modulation dimension approximately mimics the spacing of modulation filter subbands of the human auditory system, while simultaneously increasing the time resolution of the modulation transform providing improved temporal control of coding noise. Jeffrey Thompson, Les E. Atlas |
ICASSP (5) | 2 |
| 2003 | Modulation spectral filtering of speechabstractRecent auditory physiological evidence points to a modulation frequency dimension in the auditory cortex. This dimension exists jointly with the tonotopic acoustic frequency dimension. Thus, audition can be considered as a relatively slowly-varying two-dimensional representation, the “modulation spectrum,” where the first dimension is the well-known acoustic frequency and the second dimension is modulation frequency. We have recently developed a fully invertible analysis/synthesis approach for this modulation spectral transform. A general application of this approach is removal or modification of different modulation frequencies in audio or speech signals, which, for example, causes major changes in perceived dynamic character. A specific application of this modification is single-channel multipletalker separation. Les E. Atlas |
INTERSPEECH | 1 |
| 2002 | Acoustic diversity for improved speech recognition in reverberant environmentsabstractWe show that even moderate reverberation has a detrimental effect on the audible quality of speech and automatic speech recognition (ASR) accuracy. In the presence of room reverberation, we assess the performance of several important speech enhancement techniques, and show that little improvement is offered. We experimentalIy show that multiple microphones are necessary for complete equalization of the speaker-to-receiver impulse response. Furthermore, if complete equalization is not possible, long reverberation time (RT60) is shown to affect ASR accuracy far more negatively than a low signal-to-reverberation ratio (SRR). Using this knowledge we develop an equalizing strategy that improves ASR accuracy by reducing RT60. Bradford W. Gillespie, Les E. Atlas |
ICASSP | 2 |
| 2002 | Molecular signal processing and detection in T-cellsabstractBiological cells in mammalian immune systems use groups of molecules to process signals for the purposes of detecting and eliminating pathogens from the body. Immunologists strive to understand the key molecular components of these living detectors. The size and complexity of cells makes detailed, quantitative data very difficult to collect. One approach to modeling with limited quantitative data is to assume that the system is optimal (e.g., due to natural selection) for the function it performs. In our work, we apply this approach to an extended, stochastic version of McKeithan's model for T-cell signal transduction [9]. This model is interpreted as a binary detector on which we impose mutual information as the optimality criterion. With a model structure, a performance metric, and an optimization algorithm, we then estimate the parameters of a molecular signal processing model. John F. Keane, Les E. Atlas |
ICASSP | 2 |
| 2002 | Modulation frequency features for audio fingerprintingabstractThis paper explores modulation frequency features with subband normalization for audio identification. Our main goal is to find features for audio fingerprinting that are invariant to time and frequency distortions, both unintentional and intentional. Two-dimensional features, called “joint acoustic and modulation frequency,” are proposed. The paper describes these features and corresponding cross entropy classification. Experimental results show that standard spectral features are inadequate when frequency distortion occurs, as in low bit rate coding or equalization. In contrast, our proposed normalized modulation frequency features can provide accurate fingerprints, even when time and frequency distortions are imposed on music passages. Somsak Sukittanon, Les E. Atlas |
ICASSP | 2 |
| 2001 | Impulses and stochastic arithmetic for signal processingabstractExplores the use of Poisson point processes and stochastic arithmetic to perform signal processing, functions. Our work is inspired by the asynchrony and fault tolerance of biological neural systems. The essence of our approach is to code the input signal as the rate parameter of a Poisson point process, perform stochastic computing operations on the signal in the arrival or "pulse" domain, and decode the output signal by estimating the rate of the resulting process. An analysis of the Poisson pulse frequency modulation encoding error is performed. Asynchronous, stochastic computing operations are applied to the impulse stream and analyzed. A special finite impulse response (FIR) filtering scheme is proposed that preserves the Poisson properties and allows filters to be cascaded without compromising the ideal signal statistics. John F. Keane, Les E. Atlas |
ICASSP | 2 |
| 2001 | Joint use of dynamical classifiers and ambiguity plane featuresabstractThis paper argues for using ambiguity plane features within dynamic statistical models for classification problems. The relative contribution of the two model components are investigated in the context of acoustically monitoring cutter wear during milling of titanium, an application where it is known that standard static classification techniques work poorly. Experiments show that explicit modeling of long-term context via a hidden Markov model state improves performance, but mainly by using this to augment sparsely labelled training data. An additional performance gain is achieved by using the shorter-term context of ambiguity plane features. Mari Ostendorf, Les E. Atlas, Robert Fish, Özgür Çetin, Somsak Sukittanon, Gary D. Bernard |
ICASSP | 2 |
| 2001 | Scalable and progressive audio codecabstractA source coding technique for variable, bandwidth-constrained channels such as the Internet must do two things: offer high quality at low data rates, and adapt gracefully to changes in available bandwidth. Here we propose an audio coding algorithm that is superior on both counts. It is inherently scalable, meaning that channel conditions can be matched without the need for additional computation. Moreover, it is compact: in subjective tests our algorithm, coded at 32 kb/s/channel, outperformed MPEG-1 Layer 3 (MP3) coded at 56 kb/s/channel (both at 44.1 kHz). We achieve this simultaneous increase in compression and scalability through use of a two-dimensional transform that concentrates relevant information into a small number of coefficients. Mark S. Vinton, Les E. Atlas |
ICASSP | 2 |
| 2000 | Hidden Markov models for monitoring machining tool-wearabstractAs summarized by Atlas, Bernard, and Narayanan (1996), the sensing of acoustic vibrations can remotely estimate the state of wear at the tool edge. This form of monitoring offers the potential to characterize, in real time, the efficiency of metal removal processes such as drilling and milling. For example, information about sudden increases in tool wear, if manifest as a change in acoustic vibration, could be valuable to a machine operator. The nature of this monitoring problem has some similarities to automatic speech recognition. For example, there is significant tool-to-tool variation in details of vibration and lifetime. Also, the easy adaptability of monitoring systems across manufacturing processes is important. In this work we model the evolution of vibration signals with the same technique which has shown to be successful in speech recognition: hidden Markov models (HMMs). We focus on the monitoring of milling processes at three different time scales and show the how HMMs can give accurate wear prediction. Les E. Atlas, Mari Ostendorf, Gary D. Bernard |
ICASSP | 1 |
| 2000 | Data-driven time-frequency classification techniques applied to tool-wear monitoringabstractIn many pattern recognition applications features are traditionally extracted from standard time-frequency representations (e.g. the spectrogram) and input to a classifier. This assumes that the implicit smoothing of, say, a spectrogram is appropriate for the classification task. It is better to begin with no implicit smoothing assumptions and optimize the time-frequency representation for each specific classification task. Here we describe two different approaches to data-driven time-frequency classification techniques, one supervised and one unsupervised. We show that a certain class of quadratic time-frequency representations will always provide best classification performance. Using our techniques we explore the wear process of milling cutters. Our initial experiments give strong evidence to the nonlinear nature of the wear process and the importance of capturing nonstationary information about each flute-strike to accurately understand the wear process. Bradford W. Gillespie, Les E. Atlas |
ICASSP | 2 |
| 1999 | Optimization of time and frequency resolution for radar transmitter identificationabstractAn entirely new set of criteria for the design of kernels for time-frequency representations (TFRs) has been previously proposed. The goal of these criteria is to produce kernels (and thus, TFRs) which will enable accurate classification without explicitly defining, a priori, the underlying structure that differentiates individual classes. These kernels, which are optimized to discriminate among multiple classes of signals, are referred to as signal class-dependent kernels, or simply class-dependent kernels. Until now, our technique has utilized the Rihaczek TFR as the base representation, deriving the optimal smoothing in time and frequency from this representation. Here the performance of the class-dependent approach is investigated in relation to the choice of the base representation. Classifier performance using several base TFRs is analyzed within the context of radar transmitter identification. It is shown that both the Rihaczek and the Wigner-Ville distributions yield equivalent results, far superior to the short-time Fourier transform. In addition, a correlation reduction step is presented. This improves performance and extensibility of the class-dependent approach. Bradford W. Gillespie, Les E. Atlas |
ICASSP | 2 |
| 1997 | Class-dependent, discrete time-frequency distributions via operator theoryabstractWe propose a property for kernel design which results in distributions for each of two classes of signals which maximally separates their energies in the time-frequency plane. Such maximally separated distributions may result in improved classification because the signal representation is optimized to accentuate the differences in signal classes. This is not the case with other time-frequency kernels which are optimized based upon some criteria unrelated to the classification task. Using our operator theory formulation for time-frequency representations, our "maximal separation" criteria takes on a very easily solved form. Analysis of the solution in both the time-frequency and ambiguity planes is given along with an example on discrete signals. Jack McLaughlin, James Droppo, Les E. Atlas |
ICASSP | 3 |
| 1997 | Automatic clustering of vector time-series for manufacturing machine monitoringabstractOur research in online monitoring of industrial milling tools has focused on the occurrence of certain wide-band transient events. Time-frequency representations of these events appear to reveal a variety of classes of transients, and a time-structure to these classes which would be well modeled using hidden Markov models. However, the identities of these classes are not known, and obtaining a labeled training set based on a priori information is not possible for reasons both theoretical and practical. Unsupervised clustering algorithms which exist are only appropriate for single vector patterns. We introduce an approach to unsupervised clustering of vector series based around the hidden Markov model. This system is justified as a generalization of a common single-vector approach, and applied to a set of vector patterns from a milling data set. Results presented illustrate the value of this approach in the milling application. Lane M. D. Owsley, Les E. Atlas, Gary D. Bernard |
ICASSP | 2 |
| 1996 | Self-organizing feature maps with perfect organizationabstractThe self-organizing feature maps (SOFMs) introduced by Kohonen (1990) have found use in a wide variety of signal processing applications. The goal of SOFMs is to allow encoding of high-dimensional data vectors in such a manner that a vector's relative position in the codebook is related to the information in the vector in as simple a fashion as possible. One measure of the SOFM's performance in achieving this goal has been proposed by Zrehen and Blayo (1992). According to this measure, many feature maps produced by previous algorithms were disorganized. We review the Zrehen disorganization measure and some of its characteristics. We then show how our version of the SOFM algorithms can accept some simple modifications to produce feature maps which achieve perfect organization under the Zrehen measure of feature map performance. We discuss the emergent geometric properties of resulting feature maps, and illustrate the results using industrial drill vibration data. We note that applying the Zrehen-constrained algorithm to a data set implies certain assumptions about the set's distribution. We discuss the implications of these assumptions in the context of a feature extraction system for industrial milling data. Lane M. D. Owsley, Les E. Atlas, Gary D. Bernard |
ICASSP | 2 |
| 1996 | Applications of time-frequency analysis to signals from manufacturing and machine monitoring sensorsabstractManufacturing industries are now demanding substantial increases in flexibility, productivity and reliability from their process machines as well as increased quality and value of their products. One important strategy to support this goal is sensor-based, on-line, real-time evaluation of key characteristics of both machines and products, throughout the manufacturing process. Recent advances in time-frequency (TF) analysis are particularly well suited to extracting key vibrational characteristics from monitoring sensors. Thus this paper presents applications of TF analysis to several important manufacturing and machine monitoring tasks, to show the value of these forms of digital signal processing applied to manufacturing. Les E. Atlas, Gary D. Bernard, Siva Bala Narayanan |
Proc. IEEE | 1 |
| 1995 | Feature extraction networks for dull tool monitoringabstractAutomatic feature extraction is a need in many current applications, including the monitoring of industrial tools. Currently available approaches suffer from a number of shortcomings. The Kohonen (1989) self-organizing neural network (SONN) has the potential to act as a feature extractor, but we find it benefits from several modifications. The purpose of these modifications is to cause feature variations to be aligned with the SONN indices so that the indices themselves can be used as measures of the features. The modified SONN is applied to the dull tool monitoring problem, and it is shown that the new algorithm extracts and characterizes useful features of the data. Lane M. D. Owsley, Les E. Atlas, Gary D. Bernard |
ICASSP | 2 |
| 1995 | The challenge of spoken language systems: research directions for the ninetiesabstractA spoken language system combines speech recognition, natural language processing and human interface technology. It functions by recognizing the person's words, interpreting the sequence of words to obtain a meaning in terms of the application, and providing an appropriate response back to the user. Potential applications of spoken language systems range from simple tasks, such as retrieving information from an existing database (traffic reports, airline schedules), to interactive problem solving tasks involving complex planning and reasoning (travel planning, traffic routing), to support for multilingual interactions. We examine eight key areas in which basic research is needed to produce spoken language systems: (1) robust speech recognition; (2) automatic training and adaptation; (3) spontaneous speech; (4) dialogue models; (5) natural language response generation; (6) speech synthesis and speech generation; (7) multilingual systems; and (8) interactive multimodal systems. In each area, we identify key research challenges, the infrastructure needed to support research, and the expected benefits. We conclude by reviewing the need for multidisciplinary research, for development of shared corpora and related resources, for computational support and far rapid communication among researchers. The successful development of this technology will increase accessibility of computers to a wide range of users, will facilitate multinational communication and trade, and will create new research specialties and jobs in this rapidly expanding area.> Ronald A. Cole, Lynette Hirschman, Les E. Atlas, Mary E. Beckman, Alan Biermann, Marcia A. Bush, Mark A. Clements, Jordan Cohen, Oscar Garcia, Brian A. Hanson, Hynek Hermansky, Steve Levinson, Kathy McKeown, Nelson Morgan, David G. Novick, Mari Ostendorf, Sharon L. Oviatt, Patti Price, Harvey F. Silverman, Judy Spitz, Alex Waibel, Clifford J. Weinstein, Stephen A. Zahorian, Victor Zue |
IEEE Trans. Speech Audio Process. | 3 |
| 1994 | Feature representations for monitoring of tool wearabstractWe address the general problem of reliable, real-time detection of faults in metal-removal processes in manufacturing. As has long been recognized by skilled machine operators, mechanical and acoustic vibrations can be reliable sources of cues for such monitoring. However, conventional dull-tool monitoring systems, which are generally based on stationary signal processing methods, are inadequate for real-time control of drilling procedure. Making use of a database from nine different drill bits, we (a) identify different features which seem to contain tool wear information, (b) document what we found to be superior signal processing tools to identify, extract and process these non-stationary features, and (c) stress the need for a fully annotated public-domain manufacturing signal database.> Siva Bala Narayanan, Gary D. Bernard, Les E. Atlas |
ICASSP (6) | 4 |
| 1994 | Improving Generalization with Active Learning
David A. Cohn, Les E. Atlas, Richard E. Ladner |
Mach. Learn. | 2 |
| 1994 | Applications of positive time-frequency distributions to speech processingabstractMuch of our current knowledge and intuition of speech is derived from analyses involving assumptions of short-time stationarity (e.g., the speech spectrogram). Such methods are, by their very nature, incapable of revealing the true nonstationary nature of speech. A careful consideration of the theory of time-frequency distributions (TFDs), however, allows the construction of methods that reveal far more of the nonstationarities of speech, thereby highlighting just what it is that conventional approaches miss. We apply two iterative methods for generating positive time-frequency distributions (TFDs) to speech analysis. Both methods make use of multiple sources of information (e.g., multiple spectrograms) to yield a high-resolution estimate of the joint time-frequency energy density of speech. Plosive events and formant harmonic structure are simultaneously preserved in these TFDs. Rapidly time-varying formants are also resolved by these TFDs, and harmonic structure is revealed, independent of sweep rate; this result is quite different from that seen with conventional speech spectrograms. The speech features observed in these distributions demonstrate that conventional sliding window techniques lose or distort much of the rich nonstationary structure of speech. Examples for synthetic formants and real speech are provided. The differences between joint distributions and conditional distributions are also illustrated.> James W. Pitton, Les E. Atlas, Patrick J. Loughlin |
IEEE Trans. Speech Audio Process. | 2 |
| 1994 | Index assignment for progressive transmission of full-search vector quantizationabstractThe authors study codeword index assignment to allow for progressive image transmission of fixed rate full-search vector quantization (VQ). They develop three new methods of assigning indices to a vector quantization codebook and formulate these assignments as labels of nodes of a full-search progressive transmission tree. The tree is used to design intermediate codewords for the decoder so that full-search VQ has a successive approximation character. The binary representation for the path through the tree represents the progressive transmission code. The methods of designing the tree that they apply are the generalized Lloyd algorithm, minimum cost perfect matching from optimization theory, and a method of principal component partitioning. Their empirical results show that the final method gives intermediate signal-to-noise ratios (SNRs) that are close to those obtained with tree-structured vector quantization, yet they have higher final SNRs. Eve A. Riskin, Richard E. Ladner, Ren-Yuh Wang, Les E. Atlas |
IEEE Trans. Image Process. | 4 |
| 1994 | Recurrent neural networks and robust time series predictionabstractWe propose a robust learning algorithm and apply it to recurrent neural networks. This algorithm is based on filtering outliers from the data and then estimating parameters from the filtered data. The filtering removes outliers from both the target function and the inputs of the neural network. The filtering is soft in that some outliers are neither completely rejected nor accepted. To show the need for robust recurrent networks, we compare the predictive ability of least squares estimated recurrent networks on synthetic data and on the Puget Power Electric Demand time series. These investigations result in a class of recurrent neural networks, NARMA(p,q), which show advantages over feedforward neural networks for time series with a moving average component. Conventional least squares methods of fitting NARMA(p,q) neural network models are shown to suffer a lack of robustness towards outliers. This sensitivity to outliers is demonstrated on both the synthetic and real data sets. Filtering the Puget Power Electric Demand time series is shown to automatically remove the outliers due to holidays. Neural networks trained on filtered data are then shown to give better predictions than neural networks trained on unfiltered time series. Jerome T. Connor, R. Douglas Martin, Les E. Atlas |
IEEE Trans. Neural Networks | 3 |
| 1993 | Coding Theory and RegularizationabstractThis paper uses two principles, the robust encoding of residuals and the efficient coding of parameters, to obtain a new learning rule for neural networks. In particular, it examines how different coding techniques give rise to different learning rules. The storage space requirements of parameters and residuals are considered. A 'group regularizer' is derived from encoding of the parameters as a whole group rather than individually.> Jerome T. Connor, Les E. Atlas |
Data Compression Conference | 2 |
| 1993 | Positive time-frequency distributions via maximum entropy deconvolution of the evolutionary spectrum
James W. Pitton, Patrick J. Loughlin, Les E. Atlas |
ICASSP (4) | 3 |
| 1992 | Quadratic detectors for general nonlinear analysis of speechabstractGeneral quadratic detectors are defined and linked to the discrete-time version of Teager's model (1980), to quadratic (also known as bilinear) time-frequency representations, and to other theory. As demonstrated on speech, these detectors have resolution advantages in simultaneous frequency selectivity and temporal resolution over linear detectors. Preprocessing for pitch detection using the detector is investigated, and it is shown that fast preprocessing for pitch tracking can be obtained by a suitable choice of coefficients. Simple speaker-independent phoneme classification based on a quadratic preprocessor and an artificial neural network is also examined, and comparative results indicate that a quadratic preprocessor facilitates the accurate estimation of waveform features. The preliminary results also suggest that these waveform features may provide useful information for fine phonetic distinctions.> Les E. Atlas |
ICASSP | 1 |
| 1992 | An information-theoretic approach to positive time-frequency distributionsabstractThe principle of minimum cross entropy (MCE) is used to generate positive distributions in the Cohen-Posch (1985) class of proper time-frequency distributions (TFDs). The MCE-TFDs are not only intuitively satisfying, but they also yield the correct marginals, have strong finite support, and are everywhere nonnegative. The usual cross-term artifacts that hinder interpretation of other time-frequency representations (most notably, the Wigner-Ville) are not a problem with the MCE-TFDs. Examples of speech and chirps are given and compared to the spectrogram. An interesting observation in the case of speech is that spectrograms more closely resemble time-conditional-frequency distributions (narrowband) or frequency-conditional-time distributions (wideband) than they do joint time-frequency distributions.> Patrick J. Loughlin, James W. Pitton, Les E. Atlas |
ICASSP | 3 |
| 1992 | Parameter estimation and restoration of noisy images using Gibbs distributions in hidden Markov models
Yunxin Zhao, Xinhua Zhuang, Les E. Atlas, Lars Anderson |
CVGIP Graph. Model. Image Process. | 3 |
| 1991 | Truly nonstationary techniques for the analysis and display of voiced speechabstractThe assumption of quasi-stationarity in the analysis of speech is questioned from the standpoint of the best resolution of global periodicity and vocal tract formant frequency locations. The generalized class of time-frequency representations is discussed from the standpoint of speech analysis and is integrated with a description of the spectrogram. The spectrogram is known to have minimal interference artifacts, and the truly nonstationary cone-kernel time-frequency representation (CK-TFR) is shown to be similarly free of interference. The CK-TFR is observed to give higher resolution of the time point of glottal closure, and the clarity of some formant frequencies, especially for a nasal consonant, is shown to be better than the spectrogram. It is shown that interference-free representations of speech with a higher simultaneous resolution in time and frequency than the spectrogram are possible, and that these new representations may be applicable to the better analysis and understanding of speech.> Les E. Atlas, Patrick J. Loughlin, James W. Pitton |
ICASSP | 1 |
| 1991 | New properties to alleviate interference in time-frequency representationsabstractTwo properties are presented that constrain the cross-terms of Cohen-class time-frequency representations (TFRs) to appear only at signal frequencies, and only when the signal is nonzero. These properties thus guarantee strong finite support, i.e. the TFR is zero everywhere the signal or its spectrum is zero. When combined with cross-term attenuation, one can obtain TFRs with spectrogram-like interference suppression, but without the inherent time-frequency resolution tradeoff of the spectrogram.> Patrick J. Loughlin, James W. Pitton, Les E. Atlas |
ICASSP | 3 |
| 1991 | Recurrent Networks and NARMA Modeling
Jerome T. Connor, Les E. Atlas, R. Douglas Martin |
NIPS | 2 |
| 1990 | New nonstationary techniques for the analysis and display of speech transientsabstractA nonstationary analysis technique is applied to speech, with encouraging results. This technique makes use of a generalized time-frequency representation (GTFR) which uses a kernel that allows finite-time support while suppressing interference terms. This kernel has a cone shape in the (t, tau ) plane, where t is time of a signal and tau is an autocorrelationlike lag. The processing thus allows time and frequency resolution equivalent to the Wigner distribution but does not have significant interference terms. It is shown that the effect of this distribution on the long-term (1.5-s) display of speech is a visible enhancement of formant tracks. Potentially more important shorter-term (30-ms or less) advantages are illustrated and quantified as more accurate locations of bursts in the time-frequency distribution. It is shown that the GTFR technique offers a representation which is less sensitive to threshold settings and is thus more robust.> Les E. Atlas, William Koolman, Patrick J. Loughlin, Randy Cole |
ICASSP | 1 |
| 1990 | A comparison of the LBG algorithm and Kohonen neural network paradigm for image vector quantizationabstractThe creation of an acceptable codebook, as defined by three methods of measuring performance (peak signal-to-noise ratio, image quality, and entropy), is discussed and how the Linde-Buzo-Gray (LBG) and Kohonen neural network (KNN) methods differ detailed. The results show that the codebooks generated by these two methods both enable low bits-per-pixel coding with low distortion. When using fewer training vectors, and when given a suboptimal initial codebook, the KNN method outperformed the LBG. For a theoretical lower bound, mean square error comparisons to an optimal N-level k-dimensional quantizer lower bound were made using a Gaussian source. As k increased, the KNN performance came quite close to the optimal quantizer.> Jean D. Mc Auliffe, Les E. Atlas, Carlos Rivera |
ICASSP | 2 |
| 1989 | A comparison of processor topologies for a fast trainable neural network for speech recognitionabstractA fast processing system is necessary to provide adequate learning speed in multilayer neural networks (NNs). Some schemes for mapping from a multilayer NN to a parallel digital processor topology are discussed. For a mesh topology there exists an optimal point where the computation count is minimum. In order to allow for applications such as a speaker-independent speech recognizer, the authors extend this mesh architecture to operate on sequential or, specifically, spatio-temporal inputs. A pipelining scheme is thus revealed, making it possible to improve the processing throughput. An extension of the processing element structure is obtained by introducing dynamic neurons and a consequent pipelining architecture.> Yoshitake Suzuki, Les E. Atlas |
ICASSP | 2 |
| 1989 | Parameter estimation and restoration of noisy images using Gibbs distributions in hidden Markov modelsabstractA noisy image model is formulated by integrating the image and noise models into the hidden and observation layers of a hidden Markov model (HMM). The true image is modeled by a Markov random field (MRF), and the noise is a flip error that changes the gray level of image pixels according to a stochastic matrix. An algorithm for parameter estimation is developed on the basis of the reestimation formulation in HMM. At each iteration of the reestimation, the Gibbs distribution (GD) parameters are estimated using gradient ascent, and the noise parameter is estimated as the percentage of pixels in the unobserved image having the same gray levels as the observed image, where the percentage is the posterior expectation over all possible configurations of the unobserved image. Gibbs samplers are used to generate the samples of MRFs, and sample averages are taken to approximate the expectation terms. Images are restored using the minimum misclassification technique. Experiments on binary images contaminated by 20-30% noise showed good restoration results.> Yunzin Zhao, Lars S. Andersen, Les E. Atlas |
ICASSP | 3 |
| 1989 | Performance Comparisons Between Backpropagation Networks and Classification Trees on Three Real-World Applications
Les E. Atlas, Ronald A. Cole, Jerome T. Connor, Mohamed A. El-Sharkawi, Robert J. Marks II, Yeshwant K. Muthusamy, Etienne Barnard |
NIPS | 1 |
| 1989 | Training Connectionist Networks with Queries and Selective Sampling
Les E. Atlas, David A. Cohn, Richard E. Ladner |
NIPS | 1 |
| 1989 | A performance comparison of trained multilayer perceptrons and trained classification treesabstractMultilayer perceptrons and trained classification trees are two very different techniques which have recently become popular. Giving enough data and time, both methods are capable of performing arbitrary nonlinear classification. The two techniques have not previously been compared on real-world problems. The authors first consider the important differences between multilayer perceptrons and classification trees and conclude that there is not enough theoretical basis for the clear-cut superiority of one technique over the other. They then present results of a number of empirical tests on quite different problems in power system load forecasting and speaker-independent vowel identification. They compare the performance for classification and prediction in terms of accuracy outside the training set. In all cases, even with various sizes of training sets, the multilayer perceptron performed as well as or better than the trained classification trees. The authors are confident that the univariate version of the trained classification trees do not perform as well as the multilayer perceptron. More studies are needed, however, on the comparative performance of the linear combination version of the classification trees.> Les E. Atlas, Jerome T. Connor, Dong-Chul Park 0001, Mohamed A. El-Sharkawi, Robert J. Marks II, Alan F. Lippman, Ronald A. Cole, Yeshwant K. Muthusamy |
SMC | 1 |
| 1988 | Application of the Gibbs distribution to hidden Markov modeling in isolated word recognitionabstractA new method of formulating hidden Markov models (HMM) for isolated word recognition is presented. The authors model probabilities of hidden state sequences as Gibbs distributions (GDs) instead of the conventional products of transition probabilities. This formulation is based on the Hammersley-Clifford theorem which establishes the equivalence between Markov random fields (MRF) and GDs. The Markov chains in HMM are equivalent to one-dimensional, first order neighborhood MRFs. The observation sequences are modeled by the usual autoregressive Gaussian densities. The flexibility in the choice of energy functions in GDs makes it possible to use only a few parameters while maintaining a powerful model. The authors have developed a learning algorithm to estimate the parameters using maximum likelihood estimation and an algorithm to efficiently compute 1-D, first order neighborhood GDs using a lattice structure.> Yunxin Zhao, Les E. Atlas, Xinhua Zhuang |
ICASSP | 2 |
| 1987 | An Artificial Neural Network for Spatio-Temporal Bipolar Patterns: Application to Phoneme Classification
Les E. Atlas, Toshiteru Homma, Robert J. Marks II |
NIPS | 1 |
| 1987 | The Performance of Convex Set Projection Based Neural Networks
Robert J. Marks II, Les E. Atlas, Seho Oh, James A. Ritcey |
NIPS | 2 |
| 1986 | Image recognition with inexact processingabstractShannon has shown that one can communicate exactly over a noisy channel. It thus stands to reason that one can compute accurately with an inexact processor. For example, optical processors offer advantages of massive parallelism and great speed but suffer from significant inexactitude. Another example is wafer-scale VLSI processors, which can be inexact due to a distribution of defects. We propose an architecture whereby the accuracy of matched filter processors can be improved significantly at the cost of a modest increase in computational intensity. Errors due to noisy data and/or inexact computing can be detected and in some cases corrected. We show that over one million objects can potentially be detected with single error correction capability using only 25 (composite) matched filters. Robert J. Marks II, Les E. Atlas |
ICASSP | 2 |
| 1985 | Cross-channel correlation for the enhancement of noisy speechabstractA solution to the problem of the degradation of voiced speech by variable noise is proposed. The proposed method uses a filterbank analyzer to determine which areas of the spectrum contain periodicities. It is assumed that these periodicities are characteristic of voiced speech and that spectral areas which are lacking in periodicity arise from additive random noise. Correlation values, which indicate the amount of periodicity across the spectrum, are used for weights in the resynthesis of enhanced speech. Preliminary experimental results were obtained. These results indicate that the algorithm is useful for identifying noise-free and noisy areas of the spectrum and that some restoration of the original magnitude spectrum is possible. Les E. Atlas, Leo M. Hengky |
ICASSP | 1 |