Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Malcolm Slaney

dblp:75/6837 · DBLP profile ↗
← Back
61ranked-venue papers
18as first author
4since 2021 · last 2023
0000-0001-9733-4864ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 45 · 14 first-author · 2 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
13 papers
Audio and music processing · 56% Multimedia analysis and retrieval · 34% Multimedia systems and quality of experience · 4%
Human-computer interaction and pervasive computing
2 papers
Wearable and physiological sensing · 39% Haptics and multimodal interaction · 39% Human-AI interaction · 17%
Databases, data mining, and information retrieval
4 papers
Web and social media mining · 36% Recommender systems · 30% Data integration and cleaning · 18%
Theoretical computer science
1 paper
Algorithms and data structures · 67% Mathematical optimization · 33%
Artificial intelligence
5 papers
Vision and language · 65% Probabilistic and Bayesian machine learning · 25% Image recognition and object detection · 10%

Topics — the 30 heaviest of 45, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
music information retrieval
0.232008
Acoustic Chord Transcription and Key Extraction From Audio Using Key-Dependent HMMs Trained on Synthesized Audio · IEEE Trans. Speech Audio Process. 2008
Analysis of Minimum Distances in High-Dimensional Musical Spaces · IEEE Trans. Speech Audio Process. 2008
Content-Based Music Information Retrieval: Current Directions and Future Challenges · Proc. IEEE 2008
Computer vision › Vision and language
multimodal dialogue
0.212015
A Study of Multimodal Addressee Detection in Human-Human-Computer Interaction · IEEE Trans. Multim. 2015
Human-AI interaction › conversational interaction
addressee detection
0.212015
A Study of Multimodal Addressee Detection in Human-Human-Computer Interaction · IEEE Trans. Multim. 2015
Audio and music processing › audio analysis › audio content analysis
audio fingerprinting
0.222008
Analysis of Minimum Distances in High-Dimensional Musical Spaces · IEEE Trans. Speech Audio Process. 2008
Content-Based Music Information Retrieval: Current Directions and Future Challenges · Proc. IEEE 2008
Algorithms and data structures › data structure design › search structures › hashing
locality-sensitive hashing
0.112012
Optimal Parameters for Locality-Sensitive Hashing · Proc. IEEE 2012
Algorithms and data structures › similarity search
nearest neighbor search
0.112012
Optimal Parameters for Locality-Sensitive Hashing · Proc. IEEE 2012
Mathematical optimization
parameter optimization
0.112012
Optimal Parameters for Locality-Sensitive Hashing · Proc. IEEE 2012
Web and social media mining › social network analysis
centrality measures
0.112011
Identifying authoritative sources of multimedia content: mining specificity and expertise from large-scale multimedia databases · ACM Multimedia 2011
Data integration and cleaning
missing data
0.112011
Recommender Systems, Missing Data and Statistical Model Estimation · IJCAI 2011
Web and social media mining
web mining
0.112011
Identifying authoritative sources of multimedia content: mining specificity and expertise from large-scale multimedia databases · ACM Multimedia 2011
Multimedia analysis and retrieval
image retrieval
0.112011
Identifying authoritative sources of multimedia content: mining specificity and expertise from large-scale multimedia databases · ACM Multimedia 2011
Multimedia analysis and retrieval
image classification
0.112010
Image classification using the web graph · ACM Multimedia 2010
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.112008
Resolving tag ambiguity · ACM Multimedia 2008
Information retrieval
retrieval models
0.112008
Resolving tag ambiguity · ACM Multimedia 2008
Recommender systems
tag recommendation
0.112008
Resolving tag ambiguity · ACM Multimedia 2008
Audio and music processing › music transcription
chord recognition
0.112008
Acoustic Chord Transcription and Key Extraction From Audio Using Key-Dependent HMMs Trained on Synthesized Audio · IEEE Trans. Speech Audio Process. 2008
Audio and music processing › music information retrieval
content-based music retrieval
0.112008
Content-Based Music Information Retrieval: Current Directions and Future Challenges · Proc. IEEE 2008
Audio and music processing › music information retrieval › music similarity
cover song identification
0.112008
Analysis of Minimum Distances in High-Dimensional Musical Spaces · IEEE Trans. Speech Audio Process. 2008
Audio and music processing
music analysis
0.112008
Content-Based Music Information Retrieval: Current Directions and Future Challenges · Proc. IEEE 2008
Audio and music processing › music information retrieval
music similarity
0.112008
Analysis of Minimum Distances in High-Dimensional Musical Spaces · IEEE Trans. Speech Audio Process. 2008
Human-robot interaction
multi-party interaction
0.112015
A Study of Multimodal Addressee Detection in Human-Human-Computer Interaction · IEEE Trans. Multim. 2015
Audio and music processing
audio classification
0.112006
Discrimination of speech from nonspeech based on multiscale spectro-temporal Modulations · IEEE Trans. Speech Audio Process. 2006
Audio and music processing › speech processing
speech/nonspeech discrimination
0.112006
Discrimination of speech from nonspeech based on multiscale spectro-temporal Modulations · IEEE Trans. Speech Audio Process. 2006
Multimedia analysis and retrieval
near-duplicate detection
0.012011
Identifying authoritative sources of multimedia content: mining specificity and expertise from large-scale multimedia databases · ACM Multimedia 2011
Computer vision › Image recognition and object detection
image classification
0.012010
Image classification using the web graph · ACM Multimedia 2010
Image and video processing
signal decomposition
0.012010
Solving Demodulation as an Optimization Problem · IEEE Trans. Speech Audio Process. 2010
Image and video processing
video segmentation
0.012001
Multimedia edges: finding hierarchy in all dimensions · ACM Multimedia 2001
Multimedia systems and quality of experience › multimedia synchronization
audio-visual synchronization
0.012000
FaceSync: A Linear Operator for Measuring Synchronization of Video Facial Images and Audio Tracks · NIPS 2000
Computational social science and digital humanities
social media analysis
0.012008
Resolving tag ambiguity · ACM Multimedia 2008
Information retrieval › multimedia analysis and retrieval › music retrieval
query by humming
0.012008
Content-Based Music Information Retrieval: Current Directions and Future Challenges · Proc. IEEE 2008

Methods — techniques the papers use, named apart from their topics

multimodal fusion · 0.4beamforming · 0.4statistical model estimation · 0.2graph centrality · 0.2copy detection · 0.2probabilistic framework · 0.2metadata analysis · 0.2locality-sensitive hashing · 0.2web graph features · 0.2semi-supervised learning · 0.2multimodal feature fusion · 0.2kd-tree · 0.1convex optimization · 0.1symbolic music representation · 0.1near-neighbor retrieval · 0.1hidden markov model · 0.1audio shingling · 0.1audio feature extraction · 0.1
YearPublicationVenuePosition
2023 Disentangling Speech from Surroundings with Neural Embeddings
abstract
We present a method to separate speech signals from noisy environments in the embedding space of a neural audio codec. We introduce a new training procedure that allows our model to produce structured encodings of audio waveforms given by embedding vectors, where one part of the embedding vector represents the speech signal, and the rest represent the environment. We achieve this by partitioning the embeddings of different input waveforms and training the model to faithfully reconstruct audio from mixed partitions, thereby ensuring each partition encodes a separate audio attribute. As use cases, we demonstrate the separation of speech from background noise or from reverberation characteristics. Our method also allows for targeted adjustments of the audio output characteristics.
Ahmed Omran, Neil Zeghidour, Zalan Borsos, Félix de Chaumont Quitry, Malcolm Slaney, Marco Tagliasacchi
ICASSP5
2023 Neural architecture search for energy-efficient always-on audio machine learning
abstract
Abstract Mobile and edge computing devices for always-on classification tasks require energy-efficient neural network architectures. In this paper we present several changes to neural architecture searches that improve the chance of success in practical situations. Our search simultaneously optimizes for network accuracy, energy efficiency and memory usage. We benchmark the performance of our search on real hardware, but since running thousands of tests with real hardware is difficult, we use a random forest model to roughly predict the energy usage of a candidate network. We present a search strategy that uses both Bayesian and regularized evolutionary search with particle swarms, and employs early stopping to reduce the computational burden. Our search, evaluated on a sound event classification dataset based upon AudioSet, results in an order of magnitude less energy per inference and a much smaller memory footprint than our baseline MobileNetV1/V2 implementations while slightly improving task accuracy. We also demonstrate how combining a 2D spectrogram with a convolution with many filters causes a computational bottleneck for audio classification and that alternative approaches reduce the computational burden but sacrifice task accuracy.
Daniel T. Speckhard, Karolis Misiunas, Sagi Perel, Tenghui Zhu, Simon Carlile, Malcolm Slaney
Neural Comput. Appl.6
2022 Multi-Channel Speech Denoising for Machine Ears
abstract
This work describes a speech denoising system for machine ears that aims to improve speech intelligibility and the overall listening experience in noisy environments. We recorded approximately 100 hours of audio data with reverberation and moderate environmental noise using a pair of microphone arrays placed around each of the two ears and then mixed sound recordings to simulate adverse acoustic scenes. Then, we trained a multi-channel speech denoising network (MCSDN) on the mixture of recordings. To improve the training, we employ an unsupervised method, complex angular central Gaussian mixture model (cACGMM), to acquire cleaner speech from noisy recordings to serve as the learning target. We propose a MCSDN-Beamforming-MCSDN framework in the inference stage. The results of the subjective evaluation show that the cACGMM improves the training data, resulting in better noise reduction and user preference, and the entire system improves the intelligibility and listening experience in noisy situations.
Cong Han 0001, Emine Merve Kaya, Kyle Hoefer, Malcolm Slaney, Simon Carlile
ICASSP4
2021 VHP: Vibrotactile Haptics Platform for On-body Applications
abstract
Wearable vibrotactile devices have many potential applications, including sensory substitution for accessibility and notifications. Currently, vibrotactile experimentation is done using large lab setups. However, most practical applications require standalone on-body devices and integration into small form factors. Such integration is time-consuming and requires expertise.
Artem Dementyev, Pascal Getreuer, Dimitri Kanevsky, Malcolm Slaney, Richard F. Lyon
UIST4
2018 Using audio-visual information to understand speaker activity: Tracking active speakers on and off screen
abstract
We present a system that associates faces with voices in a video by fusing information from the audio and visual signals. The thesis underlying our work is that an extreme simple approach to generating (weak) speech clusters can be combined with strong visual signals to effectively associate faces and voices by aggregating statistics across a video. This approach does not need any training data specific to this task and leverages the natural coherence of information in the audio and visual streams. It is particularly applicable to tracking speakers in videos on the web where a priori information about the environment (e.g., number of speakers, spatial signals for beamforming) is not available.
Kenneth Hoover, Sourish Chaudhuri, Caroline Pantofaru, Ian Sturdy, Malcolm Slaney
ICASSP5
2017 CNN architectures for large-scale audio classification
abstract
Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of a dataset of 70M training videos (5.24 million hours) with 30,871 video-level labels. We examine fully connected Deep Neural Networks (DNNs), AlexNet [1], VGG [2], Inception [3], and ResNet [4]. We investigate varying the size of both training set and label vocabulary, finding that analogs of the CNNs used in image classification do well on our audio classification task, and larger training and label sets help up to a point. A model using embeddings from these classifiers does much better than raw features on the Audio Set [5] Acoustic Event Detection (AED) classification task.
Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis, Jort F. Gemmeke, Aren Jansen, R. Channing Moore, Manoj Plakal, Devin Platt, Rif A. Saurous, Bryan Seybold, Malcolm Slaney, Ron J. Weiss, Kevin W. Wilson
ICASSP11
2015 Probabilistic features for connecting eye gaze to spoken language understanding
abstract
Many users obtain content from a screen and want to make requests of a system based on items that they have seen. Eye-gaze information is a valuable signal in speech recognition and spoken-language understanding (SLU) because it provides context for a user's next utterance-what the user says next is probably conditioned on what they have seen. This paper investigates three types of features for connecting eye-gaze information to an SLU system: lexical, and two types of eye-gaze features. These features help us to understand which object (i.e. a link) that a user is referring to on a screen. We show a 17% absolute performance improvement in the referenced-object F-score by adding eye-gaze features to conventional methods based on a lexical comparison of the spoken utterance and the text on the screen.
Anna Prokofieva, Malcolm Slaney, Dilek Hakkani-Tür
ICASSP2
2015 Multimodal addressee detection in multiparty dialogue systems
abstract
Addressee detection answers the question, “Are you talking to me?” When multiple users interact with a dialogue system, it is important to know when a user is speaking to the computer and when he or she is speaking to another person. We approach this problem from a multimodal perspective, using lexical, acoustic, visual, dialog state, and beam-forming information. Using data from a multiparty dialogue system, we demonstrate the benefit of using multiple modalities over using a single modality. We also assess the relative importance of the various modalities in predicting the addressee. In our experiments, we find that acoustic features are by far the most important, that ASR and system-state information are useful, and that visual and beamforming features provide little additional benefit. Our study suggests that acoustic, lexical, and system state information are an effective, economical combination of modalities to use in addressee detection.
T. J. Tsai 0001, Andreas Stolcke, Malcolm Slaney
ICASSP3
2015 A Study of Multimodal Addressee Detection in Human-Human-Computer Interaction
abstract
The goal of addressee detection is to answer the question , “Are you talking to me?” When a dialogue system interacts with multiple users, it is crucial to detect when a user is speaking to the system as opposed to another person. We study this problem in a multimodal scenario, using lexical, acoustic, visual, dialogue state, and beamforming information. Using data from a multiparty dialogue system, we quantify the benefits of using multiple modalities over using a single modality. We also assess the relative importance of the various modalities, as well as of key individual features, in estimating the addressee. We find that energy-based acoustic features are by far the most important, that information from speech recognition and system state is useful as well, and that visual and beamforming features provide little additional benefit. While we find that head pose is affected by whom the speaker is addressing, it yields little nonredundant information due to the system acting as a situational attractor. Our findings would be relevant to multiparty, open-world dialogue systems in which the agent plays an active, conversational role, such as an interactive assistant deployed in a public, open space. For these scenarios , our study suggests that acoustic, lexical, and system-state information is an effective and practical combination of modalities to use for addressee detection. We also consider how our analyses might be affected by the ongoing development of more realistic, natural dialogue systems.
T. J. Tsai 0001, Andreas Stolcke, Malcolm Slaney
IEEE Trans. Multim.3
2014 Gaze-enhanced speech recognition
abstract
This work demonstrates through simulations and experimental work the potential of eye-gaze data to improve speech-recognition results. Multimodal interfaces, where users see information on a display and use their voice to control an interaction, are of growing importance as mobile phones and tablets grow in popularity. We demonstrate an improvement in speech-recognition performance, as measured by word error rate, by rescoring the output from a large-vocabulary speech-recognition system. We use eye-gaze data as a spotlight and collect bigram word statistics near to where the user looks in time and space. We see a 25% relative reduction in the word-error rate over a generic language model, and approximately a 10% reduction in errors over a strong, page-specific baseline language model.
Malcolm Slaney, Rahul Rajan, Andreas Stolcke, Partha Parthasarathy
ICASSP1
2014 Eye Gaze for Spoken Language Understanding in Multi-modal Conversational Interactions
abstract
When humans converse with each other, they naturally amalgamate information from multiple modalities (i.e., speech, gestures, speech prosody, facial expressions, and eye gaze). This paper focuses on eye gaze and its combination with speech. We develop a model that resolves references to visual (screen) elements in a conversational web browsing system. The system detects eye gaze, recognizes speech, and then interprets the user's browsing intent (e.g., click on a specific element) through a combination of spoken language understanding and eye gaze tracking. We experiment with multi-turn interactions collected in a wizard-of-Oz scenario where users are asked to perform several web-browsing tasks. We compare several gaze features and evaluate their effectiveness when combined with speech-based lexical features. The resulting multi-modal system not only increases user intent (turn) accuracy by 17%, but also resolves the referring expression ambiguity commonly observed in dialog systems with a 10% increase in F-measure.
Dilek Hakkani-Tür, Malcolm Slaney, Asli Celikyilmaz, Larry Heck
ICMI2
2014 The Relation of Eye Gaze and Face Pose: Potential Impact on Speech Recognition
abstract
We are interested in using context to improve speech recognition and speech understanding. Knowing what the user is attending to visually helps us predict their utterances and thus makes speech recognition easier. Eye gaze is one way to access this signal, but is often unavailable (or expensive to gather) at longer distances. In this paper we look at joint eye-gaze and facial-pose information while users perform a speech reading task. We hypothesize, and verify experimentally, that the eyes lead, and then the face follows. Face pose might not be as fast, or as accurate a signal of visual attention as eye gaze, but based on experiments correlating eye gaze with speech recognition, we conclude that face pose provides useful information to bias a recognizer toward higher accuracy.
Malcolm Slaney, Andreas Stolcke, Dilek Hakkani-Tür
ICMI1
2014 Towards better performance with heterogeneous training data in acoustic modeling using deep neural networks
abstract
Modeling heterogeneous data sources remains a fundamental chal-lenge of acoustic modeling in speech recognition. We call this the multi-condition problem because the speech data come from many different conditions. In this paper, we introduce the fundamen-tal confusability problem in multi-condition learning, then discuss the problem formalization, the taxonomy, and the architectures for multi-condition learning. While the ideas presented are applicable to all classifiers, we focus our attention in this work on acoustic models based on deep neural networks (DNN). We propose four different strategies for multi-condition learning of a DNN that we refer to as a mixed-condition model, a condition-dependent model, a condition-normalizing model, and a condition-aware model. Based on the experimental results on the voice search and short message dictation task and the Aurora 4 task, we show that the confusabil-ity introduced when modeling heterogeneous data depends on the source of acoustic distortion itself, the front-end feature extractor, and the classifier. We also demonstrate the best approach for dealing with heterogeneous data may not be to let the model sort it out blindly, even with a classifier as sophisticated as a DNN. Index Terms — Multi-task learning, deep learning, CD-DNN-HMM, noise robustness, channel compensation
Yan Huang 0028, Malcolm Slaney, Michael L. Seltzer, Yifan Gong 0001
INTERSPEECH2
2014 The influence of pitch and noise on the discriminability of filterbank features
abstract
Most features used for speech recognition are derived from the out-put of a filterbank inspired by the auditory system. The two most commonly used filter shapes are the triangular filters used in MFCC (mel-frequency cepstral coefficients) and the gammatone filters that model psychoacoustic critical bands. However, for both of these fil-terbanks there are free parameters that must be chosen by the system designer. In this paper, we explore the effect that different parameter settings have on the discriminability of speech sound classes. Specif-ically, we focus our attention on two primary parameters: the filter shape (triangular or gammatone) and the filter bandwidth. We use variations in the noise level and the pitch to explore the behavior of different filterbanks. We use the Fisher linear discriminant to give us insight about why some filterbanks perform better than others. We observe three things: 1) there are significant differences even among different implementations of the same filterbank, 2) wider filters help remove the non-informative pitch information, and 3) the Fisher cri-teria helps us understand why. We validate the Fisher measure with speech recognition experiments on the Aurora-4 speech corpus.
Malcolm Slaney, Michael L. Seltzer
INTERSPEECH1
2014 Eye gaze for understanding conversational speech
abstract
Eye gaze is a useful indication of attention and, as such, can be a valuable feature to improve spoken-language understanding in human-computer interaction. Based on the hypothesis that users look at a link before selecting it, we investigate the use of novel eye-gaze features to improve link click event prediction. Our data comprises users performing a variety of online tasks such as form filling and web browsing, and we show significant performance improvement by incorporating the use of gaze features. In addition, our analysis shows that there is much user-specific variation in gaze, so we are also looking to improve the modeling of gaze by user- and task-specific adaptation.
Anna Prokofieva, Dilek Hakkani-Tür, Malcolm Slaney
SLT3
2014 Artificial neural network features for speaker diarization
abstract
Speaker diarization finds contiguous speaker segments in an audio recording and clusters them by speaker identity, without any a-priori knowledge. Diarization is typically based on short-term spectral features such as Mel-frequency cepstral coefficients (MFCCs). Though these features carry average information about the vocal tract characteristics of a speaker, they are also susceptible to factors unrelated to the speaker identity. In this study, we propose an artificial neural network (ANN) architecture to learn a feature transform that is optimized for speaker diarization. We train a multi-hidden-layer ANN to judge whether two given speech segments came from the same or different speakers, using a shared transform of the input features that feeds into a bottleneck layer. We then use the bottleneck layer activations as features, either alone or in combination with baseline MFCC features in a multistream mode, for speaker diarization on test data. The resulting system is evaluated on various corpora of multi-party meetings. A combination of MFCC and ANN features gives up to 14% relative reduction in diarization error, demonstrating that these features are providing an additional independent source of knowledge.
Sree Harsha Yella, Andreas Stolcke, Malcolm Slaney
SLT3
2013 Characteristic contours of syllabic-level units in laughter
abstract
Trying to automatically detect laughter and other nonlin-guistic events in speech raises a fundamental question: Is it appropriate to simply adopt acoustic features that have tradi-tionally been used for analyzing linguistic events? Thus we take a step back and propose syllabic-level features that may show a contrast between laughter and speech in their intensity-, pitch-, and timbral-contours and rhythmic patterns. We mo-tivate and define our features and evaluate their effectiveness in correctly classifying laughter from speech. Inclusion of our features in the baseline feature set for the Social Signals Sub-Challenge of the Computational Paralinguistics Challenge yielded an improvement of 2.4 % in Unweighted Average Area Under the Curve (UAAUC). But beyond objective metrics, an-alyzing laughter at a phonetically meaningful level has allowed us to examine the characteristic contours of laughter and to rec-ognize the importance of the shape of its intensity envelope.
Jieun Oh, Eunjoon Cho, Malcolm Slaney
INTERSPEECH3
2013 Pitch-gesture modeling using subband autocorrelation change detection
abstract
Calculating speaker pitch (or f0) is typically the first computational step in modeling tone and intonation for spoken language understanding. Usually pitch is treated as a fixed, single-valued quantity. The inherent ambiguity judging the octave of pitch, as well as spurious values, leads to errors in modeling pitch gestures that propagate in a computational pipeline. We present an alternative that instead measures changes in the harmonic structure using a subband autocorrelation change detector (SACD). This approach builds upon new machine-learning ideas for how to integrate autocorrelation information across subbands. Importantly however, for modeling gestures, we preserve multiple hypotheses and integrate information from all harmonics over time. The benefits of SACD over standard pitch approaches include robustness to noise and amount of voicing. This is important for real-world data in terms of both acoustic conditions and speaking style. We discuss applications in tone and intonation modeling, and demonstrate the efficacy of the approach in a Mandarin Chinese tone-classification experiment. Results suggest that SACD could replace conventional pitch-based methods for modeling gestures in selected spoken-language processing tasks.
Malcolm Slaney, Elizabeth Shriberg, Jui-Ting Huang
INTERSPEECH1
2013 Introduction to the special section on the 20th anniversary of the ACM international conference on multimedia
abstract
introduction Introduction to the special section on the 20th anniversary of the ACM international conference on multimedia Authors: Klara Nahrstedt View Profile , Rainer Lienhart View Profile , Malcolm Slaney View Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 9Issue 1sOctober 2013 Article No.: 32pp 1–3https://doi.org/10.1145/2523001.2523003Published:17 October 2013Publication History 0citation112DownloadsMetricsTotal Citations0Total Downloads112Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Klara Nahrstedt, Rainer Lienhart, Malcolm Slaney
ACM Trans. Multim. Comput. Commun. Appl.3
2012 A model of attention-driven scene analysis
abstract
Parsing complex acoustic scenes involves an intricate interplay between bottom-up, stimulus-driven salient elements in the scene with top-down, goal-directed, mechanisms that shift our attention to particular parts of the scene. Here, we present a framework for exploring the interaction between these two processes in a simulated cocktail party setting. The model shows improved digit recognition in a multi-talker environment with a goal of tracking the source uttering the highest value. This work highlights the relevance of both data-driven and goal-driven processes in tackling real multi-talker, multi-source sound analysis.
Malcolm Slaney, Trevor Agus, Shih-Chii Liu, Emine Merve Kaya, Mounya Elhilali
ICASSP1
2012 Coulda, woulda, shoulda: 20 years of multimedia opportunities
abstract
The ACM Special Interest Group on Multimedia (SIGMM) is celebrating the 20th anniversary of establishing its premier conference, the ACM International Conference on Multimedia (ACM Multimedia). The panel "Coulda, Woulda, Shoulda" is part of the celebration at the ACM Multimedia 2012. The panelists and the audience will discuss the 20 years of multimedia opportunities that our community has seen, took upon and pushed forward to advance the state of the art.
Klara Nahrstedt, Malcolm Slaney
ACM Multimedia2
2012 Web-Scale Multimedia Processing and Applications [Scanning the Issue]
abstract
The articles in this special issue focus on web-scale multimedia processing as well as applications for its use.
Edward Y. Chang, Shih-Fu Chang, Alex Hauptmann 0001, Thomas S. Huang, Malcolm Slaney
Proc. IEEE5
2012 Optimal Parameters for Locality-Sensitive Hashing
abstract
Locality-sensitive hashing (LSH) is the basis of many algorithms that use a probabilistic approach to find nearest neighbors. We describe an algorithm for optimizing the parameters and use of LSH. Prior work ignores these issues or suggests a search for the best parameters. We start with two histograms: one that characterizes the distributions of distances to a point's nearest neighbors and the second that characterizes the distance between a query and any point in the data set. Given a desired performance level (the chance of finding the true nearest neighbor) and a simple computational cost model, we return the LSH parameters that allow an LSH index to meet the performance goal and have the minimum computational cost. We can also use this analysis to connect LSH to deterministic nearest-neighbor algorithms such as$k$-$d$trees and thus start to unify the two approaches.
Malcolm Slaney, Yury Lifshits, Junfeng He
Proc. IEEE1
2011 Recommender Systems, Missing Data and Statistical Model Estimation
abstract
The goal of rating-based recommender systems is to make personalized predictions and recommendations for individual users by leveraging the preferences of a community of users with respect to a collection of items like songs or movies. Recommender systems are often based on intricate statistical models that are estimated from data sets containing a very high proportion of missing ratings. This work describes evidence of a basic incompatibility between the properties of recommender system data sets and the assumptions required for valid estimation and evaluation of statistical models in the presence of missing data. We discuss the implications of this problem and describe extended modelling and evaluation frameworks that attempt to circumvent it. We present prediction and ranking results showing that models developed and tested under these extended frameworks can significantly outperform standard models. 1
Benjamin M. Marlin, Richard S. Zemel, Sam T. Roweis, Malcolm Slaney
IJCAI4
2011 Identifying authoritative sources of multimedia content: mining specificity and expertise from large-scale multimedia databases
abstract
We present a framework for identifying authoritative sources (such as web sites or individual users) that are likely to produce high-quality or interesting images. We construct a directed graph across sources based on the propensity of one source to "cite" the content from another. A graph-centrality measure scores the authority for each source, which could then be applied for retrieval purposes. We apply this method to web image retrieval, where web sites are the sources, and citations are found via copy detection; and on a photo sharing site, where individuals are the sources and citations are users' favorites. We are able to identify primary or influential sources of media while avoiding the computational cost of other approaches.
Lyndon Kennedy, Malcolm Slaney
ACM Multimedia2
2010 The information content of demodulated speech
abstract
In this paper we describe the effect of demodulation on speech signals. We compare two different algorithms for demodulating audio: the classic approach based on the Hilbert transform and a new approach based on solving a convex optimization problem. We show that convex demodulation better separates the speech information between the modulator and the carrier. We demonstrate this advantage by measuring the speech-information content using a speech-recognition experiment. Finally, we explore the effect of subband filtering on the demodulation process and the shift of information from the modulator to the carrier as the subbands become wider.
Gregory Sell, Malcolm Slaney
ICASSP2
2010 Image classification using the web graph
abstract
Image classification is a well-studied and hard problem in computer vision. We extend a proven solution for classifying web spam to handle images. We exploit the link structure of the web graph: a web page related to a given category is normally linked to other pages describing related objects. Our approach combines information from the webgraph structure with semi-supervised learning from all the unlabeled images to create a superior image-classification model for multimedia data. We show that fusing image, text and web-graph features gives a 12% improvement (in the area under the ROC curve) over content features alone in an adult image-classification experiment.
Dhruv Mahajan 0001, Malcolm Slaney
ACM Multimedia2
2010 Processing web-scale multimedia data
abstract
The Internet brings us access to multimedia databases with billions of data instances. The massive amount of data available to researchers and application developers brings both opportunities and challenges. In particular, massive amount of data makes data-driven approach feasible, but at the same time, it demands scalable algorithms.
Malcolm Slaney, Edward Y. Chang
ACM Multimedia1
2010 Solving Demodulation as an Optimization Problem
abstract
We introduce two new methods for the demodulation of acoustic signals by posing the problem in a convex optimization framework. This allows the parameters of the modulator and carrier to be explicitly defined as constraints in an optimization problem. We first show the theory used to define the demodulation relationship within the rules of convex programming. Then, for the two approaches introduced, we derive specific cost functions and constraints to solve for modulators specifically motivated by perceptual rules. The methods described here perform well with simple, harmonic, and stochastic carriers, and also in the presence of noise.
Gregory Sell, Malcolm Slaney
IEEE Trans. Speech Audio Process.2
2009 Reconciliation of human and machine speech recognition performance
abstract
This paper focuses on resolving a number of issues that appear when the performance of human speech recognition is compared to that of automatic speech recognition. In particular human experimental data suggest that the resulting error is a product of the individual streams. On the other hand, Bayesian combination requires a multiplication of the estimates of prior probabilities and likelihoods. We show that, in principle, there is no discrepancy. The product of errors is a performance measure and human and machine performance may be consistent with this empirically established regularity. The product of probabilities is step in an algorithm to achieve the performance that may or may not be consistent with the product of errors. The main problem is that most of prior discussions failed to distinguish the performance measures from the estimates of the parameters used in the algorithm.
Misha Pavel, Malcolm Slaney, Hynek Hermansky
ICASSP2
2009 Periodicity Detection and Localization using Spike Timing from the AER EAR
abstract
We present a system consisting of a spiking cochlea chip and real-time event-based processing software that is able to discriminate between two sets of sounds based on their periodicity content. The periodicity measurements are computed from the spike timing information of asynchronous output spikes from the binaural spiking-cochlea chip. The chip consists of a matched pair of silicon cochlea with an address event interface for the output. Each section of the cochlea is modeled by a second-order low-pass filter followed by a simplified Inner Hair Cell circuit and a Spiking Neuron circuit. We show discrimination results using the periodicity measure for 2 classes of sound and preliminary localization results based on a discriminated sound.
Theodore Yu, John G. Harris, Malcolm Slaney, Shih-Chii Liu
ISCAS4
2008 Resolving tag ambiguity
abstract
Tagging is an important way for users to succinctly describe the content they upload to the Internet. However, most tag-suggestion systems recommend words that are highly correlated with the existing tag set, and thus add little information to a user's contribution. This paper describes a means to determine the ambiguity of a set of (user-contributed) tags and suggests new tags that disambiguate the original tags. We introduce a probabilistic framework that allows us to find two tags that appear in different contexts but are both likely to co-occur with the original tag set. If such tags can be found, the current description is considered "ambiguous" and the two tags are recommended to the user for further clarification. In contrast to previous work, we only query the user when information is most needed and good suggestions are available. We verify the efficacy of our approach using geographical, temporal and semantic metadata, and a user study. We built our system using statistics from a large (100M) database of images and their tags.
Kilian Q. Weinberger, Malcolm Slaney, Roelof van Zwol
ACM Multimedia2
2008 Content-Based Music Information Retrieval: Current Directions and Future Challenges
abstract
The steep rise in music downloading over CD sales has created a major shift in the music industry away from physical media formats and towards online products and services. Music is one of the most popular types of online information and there are now hundreds of music streaming and download services operating on the World-Wide Web. Some of the music collections available are approaching the scale of ten million tracks and this has posed a major challenge for searching, retrieving, and organizing music content. Research efforts in music information retrieval have involved experts from music perception, cognition, musicology, engineering, and computer science engaged in truly interdisciplinary activity that has resulted in many proposed algorithmic and methodological solutions to music search using content-based methods. This paper outlines the problems of content-based music information retrieval and explores the state-of-the-art methods using audio cues (e.g., query by humming, audio fingerprinting, content-based music retrieval) and other cues (e.g., music notation and symbolic representation), and identifies some of the major challenges for the coming years.
Michael A. Casey, Remco C. Veltkamp, Masataka Goto, Marc Leman, Christophe Rhodes, Malcolm Slaney
Proc. IEEE6
2008 Analysis of Minimum Distances in High-Dimensional Musical Spaces
abstract
We propose an automatic method for measuring content-based music similarity, enhancing the current generation of music search engines and recommended systems. Many previous approaches to track similarity require brute-force, pair-wise processing between all audio features in a database and therefore are not practical for large collections. However, in an Internet-connected world, where users have access to millions of musical tracks, efficiency is crucial. Our approach uses features extracted from unlabeled audio data and near-neigbor retrieval using a distance threshold, determined by analysis, to solve a range of retrieval tasks. The tasks require temporal features-analogous to the technique of shingling used for text retrieval. To measure similarity, we count pairs of audio shingles, between a query and target track, that are below a distance threshold. The distribution of between-shingle distances is different for each database; therefore, we present an analysis of the distribution of minimum distances between shingles and a method for estimating a distance threshold for optimal retrieval performance. The method is compatible with locality-sensitive hashing (LSH)-allowing implementation with retrieval times several orders of magnitude faster than those using exhaustive distance computations. We evaluate the performance of our proposed method on three contrasting music similarity tasks: retrieval of mis-attributed recordings (fingerprint), retrieval of the same work performed by different artists (cover songs), and retrieval of edited and sampled versions of a query track by remix artists (remixes). Our method achieves near-perfect performance in the first two tasks and 75% precision at 70% recall in the third task. Each task was performed on a test database comprising 4.5 million audio shingles.
Michael A. Casey, Christophe Rhodes, Malcolm Slaney
IEEE Trans. Speech Audio Process.3
2008 Acoustic Chord Transcription and Key Extraction From Audio Using Key-Dependent HMMs Trained on Synthesized Audio
abstract
We describe an acoustic chord transcription system that uses symbolic data to train hidden Markov models and gives best-of-class frame-level recognition results. We avoid the extremely laborious task of human annotation of chord names and boundaries-which must be done to provide machine learning models with ground truth-by performing automatic harmony analysis on symbolic music files. In parallel, we synthesize audio from the same symbolic files and extract acoustic feature vectors which are in perfect alignment with the labels. We, therefore, generate a large set of labeled training data with a minimal amount of human labor. This allows for richer models. Thus, we build 24 key-dependent HMMs, one for each key, using the key information derived from symbolic data. Each key model defines a unique state-transition characteristic and helps avoid confusions seen in the observation vector. Given acoustic input, we identify a musical key by choosing a key model with the maximum likelihood, and we obtain the chord sequence from the optimal state path of the corresponding key model, both of which are returned by a Viterbi decoder. This not only increases the chord recognition accuracy, but also gives key information. Experimental results show the models trained on synthesized data perform very well on real recordings, even though the labels automatically generated from symbolic data are not 100% accurate. We also demonstrate the robustness of the tonal centroid feature, which outperforms the conventional chroma feature.
Kyogu Lee, Malcolm Slaney
IEEE Trans. Speech Audio Process.2
2007 Varying Time Constants and Gain Adaptation in Feature Extraction for Speech Processing
abstract
Previously we showed that band-pass filtered MFCC-like features are useful for noise robust speech discrimination and recognition. In this paper we aim to improve the previously presented features by incorporating varying time constants and gain adaptation in each frequency channel. We show that varying the time constants leads to a representation that is less prone to the effects of noise. Further, we show that gain adaptation can not only provide better performance in clean condition but can also be used to improve the noise robustness of the features. These improvements come at a very small increase in computational cost. Speech discrimination and recognition results are presented.
David V. Anderson, Sourabh Ravindran, Malcolm Slaney
ICASSP (4)3
2007 Fast Recognition of Remixed Music Audio
abstract
We present an efficient algorithm for automatically detecting remixes of pop songs in large commercial collections. Remixes are closely related as commercial products but they are not closely related in their audio spectral content because of the nature of the remixing process. Therefore spectral modelling approaches to audio similarity fail to recognize them. We propose a new approach - that chops songs into small chunks called audio shingles - to recognize remixed songs. We model the distribution of pair-wise distances between shingles by two independent processes - one corresponding to remix content and the other corresponding to non-remix content in a database. A nearest neighbour algorithm groups songs if they share shingles drawn from the remix process. Our results show 1) log-chromagram shingles separate remixed from non-remixed content with 75%-75% precision-recall performance, cepstral coefficient features do not separate the two distributions adequately 2) increasing the observations from the remix distribution increases the separability. Efficient implementation follows from the separability of the distributions using locality sensitive hashing (LSH) which speeds up automatic grouping of remixes by between one to two orders of magnitude in a 2018-song test set.
Michael A. Casey, Malcolm Slaney
ICASSP (4)2
2007 PLSA on Large Scale Image Databases
abstract
The Web and image repositories such as Fickrtrade are the largest image databases in the world. There are billions of images on the web, and hundreds of million high-quality images in image repositories. Currently, these images are indexed based on manually-entered tags and individual and group usage patterns. In this work we a exploring a third information dimension: image features. We are exploring probabilistic latent semantic analysis in order to infer which visual patterns describe each object. We wish to build models that connect words and image features, and use content features and tags to better find similar images.
Rainer Lienhart, Malcolm Slaney
ICASSP (4)2
2007 Collaborative Filtering and the Missing at Random Assumption
Benjamin M. Marlin, Richard S. Zemel, Sam T. Roweis, Malcolm Slaney
UAI4
2006 The Importance of Sequences in Musical Similarity
abstract
This paper demonstrates the importance of temporal sequences for passage-level music information retrieval. A number of audio analysis problems are solved successfully by using models that throw away the temporal sequence data. This paper suggests that we do not have this luxury when we consider a more difficult problem: that is finding musically similar passages within a narrow range of musical styles or within a single musical piece. Our results demonstrate a significant improvement in performance for audio similarity measures using temporal sequences of features, and we show that quantizing the features to string-based representations also performs well, thus admitting efficient implementations based on string matching
Michael A. Casey, Malcolm Slaney
ICASSP (5)2
2006 Discrimination of speech from nonspeech based on multiscale spectro-temporal Modulations
abstract
We describe a content-based audio classification algorithm based on novel multiscale spectro-temporal modulation features inspired by a model of auditory cortical processing. The task explored is to discriminate speech from nonspeech consisting of animal vocalizations, music, and environmental sounds. Although this is a relatively easy task for humans, it is still difficult to automate well, especially in noisy and reverberant environments. The auditory model captures basic processes occurring from the early cochlear stages to the central cortical areas. The model generates a multidimensional spectro-temporal representation of the sound, which is then analyzed by a multilinear dimensionality reduction technique and classified by a support vector machine (SVM). Generalization of the system to signals in high level of additive noise and reverberation is evaluated and compared to two existing approaches (Scheirer and Slaney, 2002 and Kingsbury et al., 2002). The results demonstrate the advantages of the auditory model over the other two systems, especially at low signal-to-noise ratios (SNRs) and high reverberation.
Nima Mesgarani, Malcolm Slaney, Shihab A. Shamma
IEEE Trans. Speech Audio Process.2
2005 Analytic Worksheets: A Framework to Support Human Analysis of Large Streaming Data Volumes
Grace Crowder, Sterling Foster, Daniel M. Russell, Malcolm Slaney, Lisa Yanguas
INTERACT4
2005 A timbre space for speech
abstract
We describe a perceptual space for timbre, define an objective metric that takes into account perceptual orthogonality and measure the quality of timbre interpolation. We discuss two timbre representations and measure perceptual judgments. We determine that a timbre space based on Mel-frequency cepstral coefficients (MFCC) is a good model for perceptual timbre space.
Hiroko Terasawa, Malcolm Slaney, Jonathan Berger
INTERSPEECH2
2004 Speech discrimination based on multiscale spectro-temporal modulations
abstract
A novel approach for content based audio classification is presented based on multiscale spectro-temporal modulation features extracted using a model of auditory cortex. The task is to discriminate speech from non-speech which consists of animal vocalizations, music and environmental sounds. Generalization of the system to signals in high level of additive noise and reverberation is evaluated and compared to two existing approaches. The results demonstrate the advantages of the auditory model over the other two systems, especially at low SNR and high reverberation.
Nima Mesgarani, Shihab A. Shamma, Malcolm Slaney
ICASSP (1)3
2004 Low-power audio classification for ubiquitous sensor networks
abstract
In the past researchers have proposed a variety of features that are based on the human auditory system. However none of these features have been able to replace mel-frequency cepstral coefficients (MFCC) as the preferred feature for audio classification problems, either because of computational costs involved or because of their poor performance in the presence of noise. In this paper we present new features derived from a model of the early auditory system. We compare the performance of the new features with MFCC in a four-class audio classification problem and show that they perform better. We also test the noise robustness of the new features in a two-way audio classification problem and show that it outperforms the MFCC. Further, these new features can be implemented in low-power analog VLSI circuitry making them ideal for low-power sensor networks.
Sourabh Ravindran, David V. Anderson, Malcolm Slaney
ICASSP (4)3
2003 BabyEars: A recognition system for affective vocalizations
Malcolm Slaney, Gerald McRoberts
Speech Commun.1
2002 Semantic-audio retrieval
abstract
This paper describes a system for connecting sounds and words in linked multi-dimensional vector spaces. The acoustic space is represented using anchor models and partitioned using agglomerative clustering. The semantic space is modeled by a hierarchical multinomial clustering model. Nodes in one space are linked by probabilistic models to the other space. With these linked models, users retrieve sounds with natural language, and the system describes new sounds with words.
Malcolm Slaney
ICASSP1
2002 Mixtures of probability experts for audio retrieval and indexing
abstract
This paper describes a system for connecting nonspeech sounds and words using linked multidimensional vector spaces. An approach based on a mixture of experts learns the mapping between one space and the other. This paper describes the conversion of audio and semantic data into their respective vector spaces. Two different mixture-of-probability-expert models are trained to learn the association between acoustic queries and the corresponding semantic explanation, and vice versa. Test results are presented based on commercial sound effects CD.
Malcolm Slaney
ICME (1)1
2001 FastMPEG: time-scale modification of bit-compressed audio information
abstract
This paper describes techniques to change the playback speed of MPEG-compressed audio, without first decompressing the audio file. There are two primary contributions in this paper. (1) We describe three techniques to perform time-scale modification in the maximally decimated domain. (2) We show how to infer the output of the auditory masking model on the new audio stream, using the information in the original file. This new FastMPEG algorithm is more than an order of magnitude more efficient than decompressing the audio, performing time-scale modification in the conventional time-domain, and recompressing.
Michele Covell, Malcolm Slaney, Art Rothstein
ICASSP2
2001 Hierarchical segmentation using latent semantic indexing in scale space
abstract
This paper describes a new algorithm which discovers the hierarchical organization of a document or media presentation. We use latent semantic indexing to describe the semantic content of the signal, and scale-space segmentation to describe its features at many different scales. We present results from a text document and a video transcript.
Malcolm Slaney, Dulce B. Ponceleon
ICASSP1
2001 Multimedia edges: finding hierarchy in all dimensions
abstract
This paper describes a new unified representation for the information in a video. We reduce the dimensionality of the signal with either a singular-value decomposition (on the semantic and image data) or mel-frequency cepstral coefficients (on the audio data) and then concatenate the vectors to form a multi-dimensional representation of the video. Using scale-space techniques we find large jumps in the video's path, which we call edges. We use these techniques to analyze the temporal properties of the audio and image data in a video. This analysis creates a hierarchical segmentation of the video, or a table-of-contents, from the audio, semantic and image data.
Malcolm Slaney, Dulce B. Ponceleon, James H. Kaufman
ACM Multimedia1
2000 FaceSync: A Linear Operator for Measuring Synchronization of Video Facial Images and Audio Tracks
abstract
FaceSync is an optimal linear algorithm that finds the degree of syn(cid:173) chronization between the audio and image recordings of a human speaker. Using canonical correlation, it finds the best direction to com(cid:173) bine all the audio and image data, projecting them onto a single axis. FaceSync uses Pearson's correlation to measure the degree of synchro(cid:173) nization between the audio and image data. We derive the optimal linear transform to combine the audio and visual information and describe an implementation that avoids the numerical problems caused by comput(cid:173) ing the correlation matrices. 1 Motivation In many applications, we want to know about the synchronization between an audio signal and the corresponding image data. In a teleconferencing system, we might want to know which of the several people imaged by a camera is heard by the microphones; then, we can direct the camera to the speaker. In post-production for a film, clean audio dialog is often dubbed over the video; we want to adjust the audio signal so that the lip-sync is perfect. When analyzing a film, we want to know when the person talking is in the shot, instead of off camera. When evaluating the quality of dubbed films, we can measure of how well the translated words and audio fit the actor's face. This paper describes an algorithm, FaceSync, that measures the degree of synchronization between the video image of a face and the associated audio signal. We can do this task by synthesizing the talking face, using techniques such as Video Rewrite [1], and then com(cid:173) paring the synthesized video with the test video. That process, however, is expensive. Our solution finds a linear operator that, when applied to the audio and video signals, generates an audio-video-synchronization-error signal. The linear operator gathers information from throughout the image and thus allows us to do the computation inexpensively. Hershey and Movellan [2] describe an approach based on measuring the mutual informa(cid:173) tion between the audio signal and individual pixels in the video. The correlation between the audio signal, x, and one pixel in the image y, is given by Pearson's correlation, r. The mutual information between these two variables is given by f(x,y) = -1/2 log(l-?). They create movies that show the regions of the video that have high correlation with the audio; Currently at IBM Almaden Research, 650 Harry Road, San Jose, CA 95120. 2. Currently at Yes Video. com, 2192 Fortune Drive, San Jose, CA 95131.
Malcolm Slaney, Michele Covell
NIPS1
1998 MACH1: nonuniform time-scale modification of speech
abstract
We propose a new approach to nonuniform time compression, called Mach1, designed to mimic the natural timing of fast speech. At identical overall compression rates, listener comprehension for Mach1-compressed speech increased between 5 and 31 percentage points over that for linearly compressed speech, and the response times dropped by 15%. For rates between 2.5 and 4.2 times real time, there was no significant comprehension loss with increasing Mach1 compression rates. In A-B preference tests, Mach1-compressed speech was chosen 95% of the time. This paper describes the Mach1 technique and our listener-test results.
Michele Covell, Margaret Withgott, Malcolm Slaney
ICASSP3
1998 Baby Ears: a recognition system for affective vocalizations
abstract
We collected more than 500 utterances from adults talking to their infants. We automatically classified 65% of the strongest utterances correctly as approval, attentional bids, or prohibition. We used several pitch and formant measures, and a multidimensional Gaussian mixture-model discriminator to perform this task. As previous studies have shown, changes in pitch are an important cue for affective messages; we found that timbre or cepstral coefficients are also important. The utterances of female speakers, in this test, were easier to classify than were those of male speakers. We hope this research will allow us to build machines that sense the "emotional state" of a user.
Malcolm Slaney, Gerald McRoberts
ICASSP1
1997 Construction and evaluation of a robust multifeature speech/music discriminator
abstract
We report on the construction of a real-time computer system capable of distinguishing speech signals from music signals over a wide range of digital audio input. We have examined 13 features intended to measure conceptually distinct properties of speech and/or music signals, and combined them in several multidimensional classification frameworks. We provide extensive data on system performance and the cross-validated training/test setup used to evaluate the system. For the datasets currently in use, the best classifier classifies with 5.8% error on a frame-by-frame basis, and 1.4% error when integrating long (2.4 second) segments of sound.
Eric D. Scheirer, Malcolm Slaney
ICASSP2
1997 Video Rewrite: driving visual speech with audio
abstract
Article Video Rewrite: driving visual speech with audio Share on Authors: Christoph Bregler Interval Research Corporation, 1801 Page Mill Road, Building C, Palo Alto, CA Interval Research Corporation, 1801 Page Mill Road, Building C, Palo Alto, CAView Profile , Michele Covell Interval Research Corporation, 1801 Page Mill Road, Building C, Palo Alto, CA Interval Research Corporation, 1801 Page Mill Road, Building C, Palo Alto, CAView Profile , Malcolm Slaney Interval Research Corporation, 1801 Page Mill Road, Building C, Palo Alto, CA Interval Research Corporation, 1801 Page Mill Road, Building C, Palo Alto, CAView Profile Authors Info & Claims SIGGRAPH '97: Proceedings of the 24th annual conference on Computer graphics and interactive techniquesAugust 1997 Pages 353–360https://doi.org/10.1145/258734.258880Published:03 August 1997 385citation1,950DownloadsMetricsTotal Citations385Total Downloads1,950Last 12 Months132Last 6 weeks23 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Christoph Bregler, Michele Covell, Malcolm Slaney
SIGGRAPH3
1996 Automatic audio morphing
abstract
This paper describes techniques to automatically morph from one sound to another. Audio morphing is accomplished by representing the sound in a multi-dimensional space that is warped or modified to produce a desired result. The multi-dimensional space encodes the spectral shape and pitch on orthogonal axes. After matching components of the sound, a morph smoothly interpolates the amplitudes to describe a new sound in the same perceptual space. Finally, the representation is inverted to produce a sound. This paper describes representations for morphing, techniques for matching, and algorithms for interpolating and morphing each sound component. Spectrographic images of a complete morph are shown at the end.
Malcolm Slaney, Michele Covell, Bud Lassiter
ICASSP1
1994 Auditory model inversion for sound separation
abstract
Techniques to recreate sounds from perceptual displays known as cochleagrams and correlograms are developed using a convex projection framework. Prior work on cochlear-model inversion is extended to account for rectification and gain adaptation. A prior technique for phase recovery in spectrogram inversion is combined with the synchronized overlap-and-add technique of speech rate modification, and is applied to inverting the short-time autocorrelation function representation in the auditory correlogram. Improved methods of initial phase estimation are explored. A range of computational cost options, with and without iteration, produce a range of quality levels from fair to near perfect.>
Malcolm Slaney, Daniel Naar, Richard F. Lyon
ICASSP (2)1
1994 Pattern Playback in the 90s
abstract
Deciding the appropriate representation to use for modeling human auditory processing is a critical issue in auditory science. While engi(cid:173) neers have successfully performed many single-speaker tasks with LPC and spectrogram methods, more difficult problems will need a richer representation. This paper describes a powerful auditory representation known as the correlogram and shows how this non-linear representation can be converted back into sound, with no loss of perceptually impor(cid:173) tant information. The correlogram is interesting because it is a neuro(cid:173) physiologically plausible representation of sound. This paper shows improved methods for spectrogram inversion (conventional pattern playback), inversion of a cochlear model, and inversion of the correlo(cid:173) gram representation.
Malcolm Slaney
NIPS1
1990 Speaker-independent vowel recognition: spectrograms versus cochleagrams
abstract
The ability of multilayer perceptrons (MLPs) trained with backpropagation to classify vowels excised from natural continuous speech is examined. Two spectral representations are compared: spectrograms and cochleagrams. The features used to train the MLPs include discrete Fourier transform (DFT) or cochleagram coefficients from a single frame in the middle of the vowel, or coefficients from each third of the vowel. The effects of estimates of pitch, duration, and the relative amplitude of the vowel were investigated. The experiments show that with coefficients alone, the cochleagram is superior to the spectrogram in classification performance for all experimental conditions. With the three additional features, however, the results are comparable. Perceptual experiments with trained human listeners on the same data revealed that MLPs perform much better than humans on vowels excised from context.>
Yeshwant K. Muthusamy, Ronald A. Cole, Malcolm Slaney
ICASSP3
1990 A perceptual pitch detector
abstract
A pitch detector based on Licklider's (1979) duplex theory of pitch perception was implemented and tested on a variety of stimuli from human perceptual tests. It is believed that this approach accurately models how people perceive pitch. It is shown that it correctly identifies the pitch of complex harmonic and inharmonic stimuli and that it is robust in the face of noise and phase changes. This perceptual pitch detector combines a cochlear model with a bank of autocorrelators. By performing an independent autocorrelation for each channel, the pitch detector is relatively insensitive to phase changes across channels. The information in the correlogram is filtered, nonlinearly enhanced, and summed across channels. Peaks are identified and a pitch is then proposed that is consistent with the peaks.>
Malcolm Slaney, Richard F. Lyon
ICASSP1