EDBT 2026 Demo / reviewers in the wild / expert
Michael S. Brandstein
dblp:52/3848
· DBLP profile ↗
29ranked-venue papers
10as first author
5since 2021 · last 2024
0009-0008-7883-3658ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
8 papers |
Audio and music processing · 92% Multimedia analysis and retrieval · 8% Visualization and visual analytics · 0% | |
| Artificial intelligence
3 papers |
Deep learning architectures and training · 67% Speech recognition and synthesis · 33% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing
speech enhancement |
1.3 | 2 | 2024 | A Multiscale Autoencoder (MSAE) Framework for End-to-End Neural Network Speech Enhancement · IEEE ACM Trans. Audio Speech Lang. Process. 2024 Speech Enhancement via Attention Masking Network (SEAMNET): An End-to-End System for Joint Suppression of Noise and Reverberation · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Machine learning › Deep learning architectures and training
autoencoder |
0.4 | 2 | 2024 | A Multiscale Autoencoder (MSAE) Framework for End-to-End Neural Network Speech Enhancement · IEEE ACM Trans. Audio Speech Lang. Process. 2024 Speech Enhancement via Attention Masking Network (SEAMNET): An End-to-End System for Joint Suppression of Noise and Reverberation · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Multimedia analysis and retrieval › object tracking
contour tracking |
0.1 | 3 | 2000 | Contour Tracking in Clutter: A Subset Approach · Int. J. Comput. Vis. 2000 Provably Fast Algorithms for Contour Tracking · CVPR 2000 A Subset Approach to Contour Tracking in Clutter · ICCV 1999 |
Natural language and speech › Speech recognition and synthesis › speech coding
low-bit-rate speech coding |
0.1 | 1 | 2006 | Exploiting nonacoustic sensors for speech encoding · IEEE Trans. Speech Audio Process. 2006 |
Natural language and speech › Speech recognition and synthesis
speech coding |
0.1 | 1 | 2006 | Exploiting nonacoustic sensors for speech encoding · IEEE Trans. Speech Audio Process. 2006 |
Natural language and speech › Speech recognition and synthesis
speech enhancement |
0.1 | 1 | 2006 | Exploiting nonacoustic sensors for speech encoding · IEEE Trans. Speech Audio Process. 2006 |
Multimedia analysis and retrieval
object tracking |
0.1 | 2 | 2000 | Provably Fast Algorithms for Contour Tracking · CVPR 2000 A Subset Approach to Contour Tracking in Clutter · ICCV 1999 |
Audio and music processing
microphone array processing |
0.0 | 2 | 2000 | Cell-based beamforming (CE-BABE) for speech acquisition with microphone arrays · IEEE Trans. Speech Audio Process. 2000 A closed-form location estimator for use with room environment microphone arrays · IEEE Trans. Speech Audio Process. 1997 |
Audio and music processing
sound source localization |
0.0 | 2 | 2000 | Cell-based beamforming (CE-BABE) for speech acquisition with microphone arrays · IEEE Trans. Speech Audio Process. 2000 A closed-form location estimator for use with room environment microphone arrays · IEEE Trans. Speech Audio Process. 1997 |
Audio and music processing
beamforming |
0.0 | 1 | 2000 | Cell-based beamforming (CE-BABE) for speech acquisition with microphone arrays · IEEE Trans. Speech Audio Process. 2000 |
Audio and music processing › beamforming
robust beamforming |
0.0 | 1 | 2000 | Cell-based beamforming (CE-BABE) for speech acquisition with microphone arrays · IEEE Trans. Speech Audio Process. 2000 |
Audio and music processing
spatial audio |
0.0 | 1 | 1997 | A closed-form location estimator for use with room environment microphone arrays · IEEE Trans. Speech Audio Process. 1997 |
Audio and music processing › sound source localization
time delay estimation |
0.0 | 1 | 1997 | A closed-form location estimator for use with room environment microphone arrays · IEEE Trans. Speech Audio Process. 1997 |
Audio and music processing › acoustic signal processing › audio signal reconstruction › audio restoration
reverberation suppression |
0.0 | 1 | 2000 | Cell-based beamforming (CE-BABE) for speech acquisition with microphone arrays · IEEE Trans. Speech Audio Process. 2000 |
Audio and music processing
speech acquisition |
0.0 | 1 | 2000 | Cell-based beamforming (CE-BABE) for speech acquisition with microphone arrays · IEEE Trans. Speech Audio Process. 2000 |
Visualization and visual analytics
visual clutter |
0.0 | 1 | 2000 | Contour Tracking in Clutter: A Subset Approach · Int. J. Comput. Vis. 2000 |
Methods — techniques the papers use, named apart from their topics
multiscale autoencoder · 1.5end-to-end training · 1.5constant-q transform · 1.5perceptual loss · 1.0end-to-end neural network · 1.0attention masking · 1.0skin vibration sensor · 0.1microwave radar · 0.1bone conduction sensor · 0.1MELPe coder · 0.1global optimization · 0.0condensation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Neurophysiological-Auditory "Listen Receipt" for Communication EnhancementabstractInformation overload, and specifically auditory overload, is common in critical situations and detrimental to communication. Currently, there is no auditory equivalent of an email read receipt to know if a person has heard a message, other than waiting for a reply. This work hypothesizes that it may be possible to decode whether a person has indeed heard a message, or in other words, create an an auditory “listen receipt,” through use of non-invasive physiological or neural monitoring. We extracted a variety of features derived from Electrodermal activity (EDA), Electroencephalography (EEG), and the correlations between the acoustic envelope of the radio message and EEG to use in the decoder. We were able to classify the cases in which the subject responded correctly to the question in the message, versus the cases where they missed or heard the message incorrectly, with an accuracy of 79% and a receiver operating characteristic (ROC) area under the curve (AUC) of 0.83. This work suggests that the concept of a “listen receipt” may be possible, and future wearable machine-brain interface technologies may be able to automatically determine if an important radio message has been missed for both human-to-human and human-to-machine communication. Christine Beauchene, Michael S. Brandstein, Thomas F. Quatieri, Eric Thompson, Christopher J. Smalt |
ICASSP | 2 |
| 2024 | A Multiscale Autoencoder (MSAE) Framework for End-to-End Neural Network Speech EnhancementabstractNeural network approaches to single-channel speech enhancement have received much recent attention. In particular, mask-based architectures have achieved significant performance improvements over conventional methods. This paper proposes a multiscale autoencoder (MSAE) for mask-based end-to-end neural network speech enhancement. The MSAE performs spectral decomposition of an input waveform within separate band-limited branches, each operating with a different rate and scale, to extract a sequence of multiscale embeddings. The proposed framework features intuitive parameterization of the autoencoder, including a flexible spectral band design based on the Constant-Q transform. Additionally, the MSAE is constructed entirely of differentiable operators, allowing it to be implemented within an end-to-end neural network, and be discriminatively trained. The MSAE draws motivation both from recent multiscale network topologies and from traditional multiresolution transforms in speech processing. Experimental results show the MSAE to provide clear performance benefits relative to conventional single-branch autoencoders. Additionally, the proposed framework is shown to outperform a variety of state-of-the-art enhancement systems, both in terms of objective speech quality metrics and automatic speech recognition accuracy. Bengt J. Borgstrom, Michael S. Brandstein |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Subject-Specific Adaptation for a Causally-Trained Auditory-Attention Decoding SystemabstractFuture hearing-aid technology may allow a listener to isolate a single talker of interest from a mixture by shifting their attention as measured by Electroencephalography (EEG). Such decoding algorithms are often trained with data from a single individual or a pool of several participants (i.e., group model). Performance in either approach is limited: group models suffer due to the variability across subjects and time, while individual models are constrained by the limited data samples available. To overcome this challenge, we introduce a subject-specific adaptive form of auditory attention decoding (AAD) over short time windows to account for the variability across EEG recording sessions. Our subject-specific augmented model, adapts a group model to an individual, significantly improving decoding accuracy by approximately 10% as compared to an individual model. This result has implications for real-time applications of neuro-steered hearing aids, where causal-training data and real-time algorithms are necessary. Christine Beauchene, Michael S. Brandstein, Stephanie Haro, Thomas F. Quatieri, Christopher J. Smalt |
ICASSP | 2 |
| 2021 | Speaker separation in realistic noise environments with applications to a cognitively-controlled hearing aidabstractFuture wearable technology may provide for enhanced communication in noisy environments and for the ability to pick out a single talker of interest in a crowded room simply by the listener shifting their attentional focus. Such a system relies on two components, speaker separation and decoding the listener's attention to acoustic streams in the environment. To address the former, we present a system for joint speaker separation and noise suppression, referred to as the Binaural Enhancement via Attention Masking Network (BEAMNET). The BEAMNET system is an end-to-end neural network architecture based on self-attention. Binaural input waveforms are mapped to a joint embedding space via a learned encoder, and separate multiplicative masking mechanisms are included for noise suppression and speaker separation. Pairs of output binaural waveforms are then synthesized using learned decoders, each capturing a separated speaker while maintaining spatial cues. A key contribution of BEAMNET is that the architecture contains a separation path, an enhancement path, and an autoencoder path. This paper proposes a novel loss function which simultaneously trains these paths, so that disabling the masking mechanisms during inference causes BEAMNET to reconstruct the input speech signals. This allows dynamic control of the level of suppression applied by BEAMNET via a minimum gain level, which is not possible in other state-of-the-art approaches to end-to-end speaker separation. This paper also proposes a perceptually-motivated waveform distance measure. Using objective speech quality metrics, the proposed system is demonstrated to perform well at separating two equal-energy talkers, even in high levels of background noise. Subjective testing shows an improvement in speech intelligibility across a range of noise levels, for signals with artificially added head-related transfer functions and background noise. Finally, when used as part of an auditory attention decoder (AAD) system using existing electroencephalogram (EEG) data, BEAMNET is found to maintain the decoding accuracy achieved with ideal speaker separation, even in severe acoustic conditions. These results suggest that this enhancement system is highly effective at decoding auditory attention in realistic noise environments, and could possibly lead to improved speech perception in a cognitively controlled hearing aid. Bengt J. Borgstrom, Michael S. Brandstein, Gregory A. Ciccarelli, Thomas F. Quatieri, Christopher J. Smalt |
Neural Networks | 2 |
| 2021 | Speech Enhancement via Attention Masking Network (SEAMNET): An End-to-End System for Joint Suppression of Noise and ReverberationabstractThis paper proposes the Speech Enhancement via Attention Masking Network (SEAMNET), a neural network-based end-to-end single-channel speech enhancement system designed for joint suppression of noise and reverberation. It formalizes an end-to-end network architecture, referred to as b-Net, which accomplishes noise suppression through attention masking in a learned embedding space. A key contribution of SEAMNET is that the b-Net architecture contains both an enhancement and an autoencoder path. This paper proposes a novel loss function which simultaneously trains both the enhancement and the autoencoder paths, so that disabling the masking mechanism during inference causes SEAMNET to reconstruct the input speech signal. This allows dynamic control of the level of suppression applied by SEAMNET via a minimum gain level, which is not possible in other state-of-the-art approaches to end-to-end speech enhancement. This paper also proposes a perceptually-motivated waveform distance measure. In addition to the b-Net architecture, this paper proposes a novel method for designing target waveforms for network training, so that joint suppression of additive noise and reverberation can be performed by an end-to-end enhancement system, which has not been previously possible. Experimental results show the SEAMNET system to outperform a variety of state-of-the-art baselines systems, both in terms of objective speech quality measures and subjective listening tests. Finally, this paper draws parallels between SEAMNET and conventional statistical model-based enhancement approaches, offering interpretability of many network components. Bengt J. Borgstrom, Michael S. Brandstein |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Speaker Recognition Using Real vs Synthetic Parallel Data for DNN Channel CompensationabstractAbstract : The effective use of synthetic multi-channel data for training denoising DNNs has been demonstrated for several speech technologies such as ASR and speaker recognition. This paper compares the use of real and synthetic data for training denoising DNNs for multi-microphone speaker recognition. Large reductions in error rates (37% and 50% for the AVG and POOL EERs and 20% and 30% for the AVG and POOL min DCFs) are attained on Mixer 6 microphone data using Mixer 1 and 2 multi-microphone data to train a denoising DNN. Nearly the same reduction in error rate is realized using room impulse response and noise estimates (RIRs) derived from the Mixer 1and 2 data and applied to just the telephone channel. Applying RIRs from three publicly available databases used in the Kaldi Aspire evaluation system yields lower but significant reductions in error rate (16% and 34% relative improvement in AVG and POOL EER and 13% and 25% relative improvement in AVG and POOL min DCFs). In all cases, the telephone channel performance on SRE10 is improved by the denoising DNNs with the real Mixer 1 and 2 trained DNN reducing EER by 12% and min DCF by 8.9%. Fred Richardson, Michael S. Brandstein, Jennifer Melot, Douglas A. Reynolds |
INTERSPEECH | 2 |
| 2007 | An Evaluation of Audio-Visual Person Recognition on the XM2VTS Corpus using the Lausanne ProtocolsabstractA multimodal person recognition architecture has been developed for the purpose of improving overall recognition performance and for addressing channel-specific performance shortfalls. This multimodal architecture includes the fusion of a face recognition system with the MIT/LL GMM/UBM speaker recognition architecture. This architecture exploits the complementary and redundant nature of the face and speech modalities. The resulting multimodal architecture has been evaluated on the XM2VTS corpus using the Lausanne open set verification protocols, and demonstrates excellent recognition performance. The multimodal architecture also exhibits strong recognition performance gains over the performance of the individual modalities. Kevin Brady 0001, Michael S. Brandstein, Thomas F. Quatieri, Robert B. Dunn |
ICASSP (4) | 2 |
| 2006 | Exploiting nonacoustic sensors for speech encodingabstractThe intelligibility of speech transmitted through low-rate coders is severely degraded when high levels of acoustic noise are present in the acoustic environment. Recent advances in nonacoustic sensors, including microwave radar, skin vibration, and bone conduction sensors, provide the exciting possibility of both glottal excitation and, more generally, vocal tract measurements that are relatively immune to acoustic disturbances and can supplement the acoustic speech waveform. We are currently investigating methods of combining the output of these sensors for use in low-rate encoding according to their capability in representing specific speech characteristics in different frequency bands. Nonacoustic sensors have the ability to reveal certain speech attributes lost in the noisy acoustic signal; for example, low-energy consonant voice bars, nasality, and glottalized excitation. By fusing nonacoustic low-frequency and pitch content with acoustic-microphone content, we have achieved significant intelligibility performance gains using the DRT across a variety of environments over the government standard 2400-bps MELPe coder. By fusing quantized high-band 4-to-8-kHz speech, requiring only an additional 116 bps, we obtain further DRT performance gains by exploiting the ear's insensitivity to fine spectral detail in this frequency region. Thomas F. Quatieri, Kevin Brady 0001, D. Messing, Joseph P. Campbell, William M. Campbell, Michael S. Brandstein, Clifford J. Weinstein, John D. Tardelli, Paul D. Gatewood |
IEEE Trans. Speech Audio Process. | 6 |
| 2004 | Multisensor MELPe using parameter substitutionabstractThe estimation of speech parameters and the intelligibility of speech transmitted through low-rate coders, such as MELP (mixed excitation linear prediction), are severely degraded when there are high levels of acoustic noise in the speaking environment. The application of nonacoustic and nontraditional sensors, which are less sensitive to acoustic noise than the standard microphone, is being investigated as a means to address this problem. Sensors being investigated include the general electromagnetic motion sensor (GEMS) and the physiological microphone (P-mic). As an initial effort in this direction, a multisensor MELPe coder (MELP coder with the addition of a noise preprocessor) using parameter substitution has been developed, where pitch and voicing parameters are obtained from GEMS and P-Mic sensors, respectively, and the remaining parameters are obtained as usual from a standard acoustic microphone. This parameter substitution technique is shown to produce significant and promising DRT (diagnostic rhyme test) intelligibility improvements over the standard 2400 bps MELPe coder in several high-noise military environments. Further work is in progress aimed at utilizing the nontraditional sensors for additional intelligibility improvements and for more effective lower-rate coding in noise. Kevin Brady 0001, Thomas F. Quatieri, Joseph P. Campbell, William M. Campbell, Michael S. Brandstein, Clifford J. Weinstein |
ICASSP (1) | 5 |
| 2004 | Automated lip-reading for improved speech intelligibilityabstractVarious psycho-acoustical experiments have concluded that visual features strongly affect the perception of speech. This contribution is most pronounced in noisy environments where the intelligibility of audio-only speech is quickly degraded. The paper explores the effectiveness of using extracted visual features, such as lip height and width, for improving speech intelligibility in noisy environments. The intelligibility content of these extracted visual features is investigated through an intelligibility test on an animated rendition of the video generated from the extracted visual features, as well as on the original video. These experiments demonstrate that the extracted video features do contain important aspects of intelligibility that may be utilized in augmenting speech enhancement and coding applications. Alternatively, these extracted visual features can be transmitted in a bandwidth effective way to augment speech coders. Matthew McClain, Kevin Brady 0001, Michael S. Brandstein, Thomas F. Quatieri |
ICASSP (1) | 3 |
| 2001 | Microphone array speech dereverberation using coarse channel modelingabstractThis paper presents a model-based method for the enhancement of multichannel speech acquired under reverberant conditions. A very coarse estimate of the channel responses associated with each source-microphone pair is derived directly from the received data on a short-term basis. These estimates are employed to modify the LPC residuals of the channel data in an effort to deemphasize the effects of reverberant energy in the resulting synthesized signal. The approach is robust to conditions of partial and approximate channel information. Specifically, the incorporated channel model requires only approximate times and amplitudes of the initial multipath reflections. In practice these impulses are responsible for the bulk of reverberant energy in the received speech signal and can be estimated to a sufficient degree on a time-varying basis. Scott M. Griebel, Michael S. Brandstein |
ICASSP | 2 |
| 2000 | Provably Fast Algorithms for Contour TrackingabstractA new tracker is presented. Two sets are identified: one which contains all possible curves as found in the image, and a second which contains all curves which characterize the object of interest. The former is constructed out of edge-points in the image, while the latter is learned prior to running. The tracked curve is taken to be the element of the first set which is nearest the second set. The formalism for the learned set of curves allows for mathematically well understood groups of transformations (e.g. affine, projective) to be treated on the same footing as less well understood deformations, which may be learned from training curves. An algorithm is proposed to solve the tracking problem, and its properties are theoretically demonstrated: it solves the global optimization problem, and does so with certain complexity bounds. Experimental results applying the proposed algorithm to the tracking of a moving finger are presented, and compared with the results of a condensation approach. Daniel Freedman, Michael S. Brandstein |
CVPR | 2 |
| 2000 | Robust Head Pose Estimation by Machine LearningabstractSupport vector machines are applied for estimating the head orientation angle of talkers in a video environment. The procedure is capable of accurately evaluating head orientations over a complete 360 degree interval and has been designed to function as part of an existing real-time, multi-talker tracking system. By relying on a facial criterion that is easily extracted from video images acquired across a range of lighting and zooming conditions, the estimator is designed to be effective in practical situations such as those encountered in video conferencing or surveillance scenarios. Michael S. Brandstein |
ICIP | 2 |
| 2000 | Head Pose Estimation for Video-Conferencing with Multiple Cameras and Microphones
Michael S. Brandstein |
ICMI | 2 |
| 2000 | Contour Tracking in Clutter: A Subset Approach
Daniel Freedman, Michael S. Brandstein |
Int. J. Comput. Vis. | 2 |
| 2000 | Cell-based beamforming (CE-BABE) for speech acquisition with microphone arraysabstractThis paper introduces a microphone array processing method that possesses the robustness of fixed beamforming along with the ability to be dynamically reconfigured to limit interference and reverberation. The basic approach is to partition the environment into two regions: an interior region (containing sources that are physically present within the room enclosure), and an exterior region (containing virtual sources of reverberation). The interior region is further divided into cells, and standard source localization techniques are used to identify those cells containing the desired source as well as sources of interference (e.g., competing talkers). Beamforming weights are then found to pass the desired signal, while simultaneously minimizing a weighted combination of interior interference and exterior reverberation. Simulation results are presented to demonstrate the effectiveness of the proposed technique when compared with conventional beamforming methods. Michael S. Brandstein, Darren B. Ward |
IEEE Trans. Speech Audio Process. | 1 |
| 1999 | An event-based method for microphone array speech enhancementabstractThis paper presents the multi-channel multi-pulse (MCMP) algorithm for the enhancement of speech degraded by reverberations and additive noise. The enhanced speech is synthesized from a sequence of impulses exciting a linear predictive filter. The excitation signal is computed from a nonlinear process which uses impulse clustering of the multi-channel speech data to discriminate portions of the linear prediction residual produced by the desired speech signal from those due to multipath effects and uncorrelated noise. The MCMP algorithm is shown to be capable of identifying and attenuating reverberant portions of the speech signal as well as reducing the effects of additive noise. Michael S. Brandstein |
ICASSP | 1 |
| 1999 | A Subset Approach to Contour Tracking in ClutterabstractA new method for tracking contours of moving objects in clutter is presented. For a given object, a model of its contours is learned from training data in the form of a subset of contour space. Greater complexity is added to the contour model by analyzing rigid and non-rigid transformations of contours separately. In the course of tracking, multiple contours may be observed due to the presence of extraneous edges in the form of clutter; the learned model guides the algorithm in picking out the correct one. The algorithm, which is posed as a solution to a minimization problem, is made efficient by the use of several iterative schemes. Results applying the proposed algorithm to the tracking of a flexing finger and to a conversing individual's lips are presented. Daniel Freedman, Michael S. Brandstein |
ICCV | 2 |
| 1999 | Multi-source face tracking with audio and visual dataabstractA real-time face tracker based on both sound and visual cues is presented. Initial talker locations are estimated acoustically from microphone array data while precise localization and tracking are derived from visual data. The image processing employs a hierarchical structure which utilizes source motion, contour geometry, color data, and facial features. The resulting system is capable of tracking multiple persons in complex backgrounds and robustly discriminating faces from similar objects. While the direct focus of this work is automated videoconferencing, the face tracking capability has utility to many multimedia and virtual reality applications. Michael S. Brandstein |
MMSP | 2 |
| 1998 | On the use of explicit speech modeling in microphone array applicationsabstractThis paper addresses the limitations of current approaches to distant-talker speech acquisition and advocates the development of techniques which explicitly incorporate the nature of the speech signal (e.g. statistical non-stationarity, method of production, pitch, voicing, formant structure, and source radiator model) into a multi-channel context. The goal is to combine the advantages of spatial filtering achieved through beamforming with knowledge of the desired time-series attributes. The potential utility of such an approach is demonstrated through the application of a multi-channel version of the dual excitation speech model. Michael S. Brandstein |
ICASSP | 1 |
| 1998 | A hybrid real-time face tracking systemabstractA hybrid real-time face tracker based on both sound and visual cues is presented. Initial talker locations are estimated acoustically from microphone array data while precise localization and tracking are derived from image information. A computationally efficient algorithm for face detection via motion analysis is employed to track individual faces at rates up to 30 frames per second. The system is robust to nonlinear source motions, complex backgrounds, varying lighting conditions, and a variety of source-camera depths. While the direct focus of this work is automated video conferencing, the face tracking capability has utility to many multimedia and virtual reality applications. Michael S. Brandstein |
ICASSP | 2 |
| 1997 | A robust method for speech signal time-delay estimation in reverberant roomsabstractConventional time-delay estimators exhibit dramatic performance degradations in the presence of multipath signals. This limits their application in reverberant enclosures, particularly when the signal of interest is speech and it may not possible to estimate and compensate for channel effects prior to time-delay estimation. This paper details an alternative approach which reformulates the problem as a linear regression of phase data and then estimates the time-delay through minimization of a robust statistical error measure. The technique is shown to be less susceptible to room reverberation effects. Simulations are performed across a range of source placements and room conditions to illustrate the utility of the proposed time-delay estimation method relative to conventional methods. Michael S. Brandstein, Harvey F. Silverman |
ICASSP | 1 |
| 1997 | Tracking multiple talkers using microphone-array measurementsabstractA method for tracking the positional estimates of multiple talkers in the operating region of an acoustic microphone array is presented. Initial talker location estimates are provided by a time-delay-based localization algorithm. These raw estimates are spatially smoothed by a Kalman filter derived from a set of potential source motion models. Data association techniques based on the estimate clusterings and source trajectories are incorporated to match location observations with individual talkers. Experimental results are presented for array recorded data using multiple talkers in a variety of scenarios. Douglas E. Sturim, Michael S. Brandstein, Harvey F. Silverman |
ICASSP | 2 |
| 1997 | A practical methodology for speech source localization with microphone arrays
Michael S. Brandstein, Harvey F. Silverman |
Comput. Speech Lang. | 1 |
| 1997 | A closed-form location estimator for use with room environment microphone arraysabstractThe linear intersection (LI) estimator, which is a closed-form method for the localization of source positions given sensor array time-delay estimate information, is presented. The LI estimator is shown to be robust and accurate, to closely model the search-based ML estimator, and to outperform a benchmark algorithm. The computational complexity of the LI estimator is suitable for use in real-time microphone-array applications where search-based location algorithms may be infeasible. Michael S. Brandstein, John E. Adcock, Harvey F. Silverman |
IEEE Trans. Speech Audio Process. | 1 |
| 1996 | A localization-error-based method for microphone-array designabstractThis paper presents a means for predicting the error region associated with a speech-source location estimate obtained from a set of microphones in a room environment. The error predictor presented is derived assuming a specific source-sensor geometry consisting of pairs of closely-spaced sensors for which a delay estimate associated with the potential source has been evaluated. The accuracy of the predictor is evaluated through a set of Monte Carlo simulations and an application of the predictor to microphone-array design in the context of a video-teleconferencing scenario is presented. Michael S. Brandstein, John E. Adcock, Harvey F. Silverman |
ICASSP | 1 |
| 1995 | A closed-form method for finding source locations from microphone-array time-decay estimatesabstractThe linear intersection (LI) estimator, a closed-form method for the localization of source positions given only the sensor array time-delay estimate information, is presented. The array is constrained to be composed of 4-element sub-arrays configured in 2 centered orthogonal pairs. A bearing line in 3-space is estimated from each sub-array and potential source locations are found via closest intersection of bearing line pairs. The final location estimate is determined by a probabilistic weighting of these potential locations. The LI estimator is shown to be robust and accurate, to closely model the ML estimator, and to outperform a representative algorithm. The computational complexity of the LI estimator is suitable for use in real-time microphone-array applications. Michael S. Brandstein, John E. Adcock, Harvey F. Silverman |
ICASSP | 1 |
| 1995 | A practical time-delay estimator for localizing speech sources with a microphone array
Michael S. Brandstein, John E. Adcock, Harvey F. Silverman |
Comput. Speech Lang. | 1 |
| 1990 | A real-time implementation of the improved MBE speech coderabstractA real-time, single digital signal processing (DSP) chip implementation of a 2.4-, 4.8-, and 8.0-kb/s improved multiband excitation (IMBE) vocoder is presented. The IMBE vocoder is based on the MBE speech model, and it is shown to generate high-quality speech under both clean and noisy conditions. In addition, the IMBE vocoder is well suited for real-time implementation since it does not require excessive computation or storage. Full-duplex operation is demonstrated using a single AT&T WE DSP 32. Aspects of the hardware architecture, algorithm implementation, and system performance are addressed.> Michael S. Brandstein, Peter A. Monta, John C. Hardwick, Jae S. Lim |
ICASSP | 1 |