EDBT 2026 Demo / reviewers in the wild / expert
Shiva Sundaram
dblp:46/2430
· DBLP profile ↗
36ranked-venue papers
14as first author
7since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 14 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Multi-Scale Compositional Constraints for Representation Learning on VideosabstractCombining simple concepts to form structured thoughts and decomposing complex concepts into their constituents is one key characteristic of human cognition. In this work we extract video representations by combining multi-scale processing with compositional constraints, i.e., we constrain the latent space created by the network so that coarse grained video features are composed from a set of fine-grained video features using simple functions. We integrate the proposed constraints in a state-of-the-art contrastive learning frame-work. In our ablations, we evaluate different formulations of the compositional constraints and composition functions. We evaluate the proposed approach for the downstream tasks of action detection in UCF-101, and video summarization in the SumMe dataset. We achieve significant improvements over the baseline, i.e., 3.9% and 6.3% relative improvements for UCF-101 and SumMe respectively, showcasing the importance of compositional video representations. Georgios Paraskevopoulos, Chandrashekhar Lavania, Lovish Chum, Shiva Sundaram |
ICASSP | 4 |
| 2022 | Enhancing Contrastive Learning with Temporal Cognizance for Audio-Visual Representation GenerationabstractAudio-visual data allows us to leverage different modalities for downstream tasks. The idea being individual streams can complement each other in the given task, thereby resulting in a model with improved performance. In this work, we present our experimental results on action recognition and video summarization tasks. The proposed modeling approach builds upon the recent advances in contrastive loss based audio-visual representation learning. Temporally cognizant audio-visual discrimination is achieved in a Transformer model by learning with a masked feature reconstruction loss over a fixed time window in addition to learning via contrastive loss. Overall, our results indicate that the addition of temporal information significantly improved the performance of the contrastive loss based framework. We achieve an action classification accuracy of 66.2% versus the next best baseline at 64.7% on the HMDB dataset. For video summarization, we attain an F1 score of 43.5 verses 42.2 on the SumMe dataset. Chandrashekhar Lavania, Shiva Sundaram, Sundararajan Srinivasan, Katrin Kirchhoff |
ICASSP | 2 |
| 2022 | Scene Representation Learning from Videos Using Self-Supervised and Weakly-Supervised TechniquesabstractHolistic understanding of videos requires the recognition of the overall scene beyond detecting foreground activity and objects. It provides valuable information for various video understanding tasks such as video summarization, scene change detection and content filtering. While significant effort has been put into developing models for scene classification in images (e.g. Places365), video-level scene recognition is relatively nascent. The scope of this paper is to address this problem of going from image representations to video for scene classification. In particular, we compare self-supervised deep learning methods on video scene recognition task using the HVU dataset. Starting from strong image level scene representations, with triplets based contrastive loss, we train a video-level scene classifier. We propose triplet sampling strategies that aid the self-supervision. We compare the self-supervised techniques against the image level scene representations, as well as a weakly supervised classifier trained on image labels. We observe that the models learned using self-supervised method outperform both baselines (with statistical significance), showing that we are able to retain the representative power of the video-level scene representations compared to a competitive image-level scene recognition model trained on Places365, while showing benefits over weakly supervised techniques. Raghuveer Peri, Srinivas Parthasarathy, Shiva Sundaram |
ICIP | 3 |
| 2021 | Audiovisual Highlight Detection in VideosabstractIn this paper, we test the hypothesis that interesting events in unstructured videos are inherently audiovisual. We combine deep image representations for object recognition and scene understanding with representations from an audiovisual affect recognition model. To this set, we include content agnostic audio-visual synchrony representations and mel-frequency cepstral coefficients to capture other intrinsic properties of audio. These features are used in a modular supervised model. We present results from two experiments: efficacy study of single features on the task, and an ablation study where we leave one feature out at a time. For the video summarization task, our results indicate that the visual features carry most information, and including audiovisual features improves over visual-only information. To better study the task of highlight detection, we run a pilot experiment with highlights annotations for a small subset of video clips and fine-tune our best model on it. Results indicate that we can transfer knowledge from the video summarization task to a model trained specifically for the task of highlight detection. Karel Mundnich, Alexandra Fenster, Aparna Khare, Shiva Sundaram |
ICASSP | 4 |
| 2021 | Disentanglement for Audio-Visual Emotion Recognition Using Multitask SetupabstractDeep learning models trained on audio-visual data have been successfully used to achieve state-of-the-art performance for emotion recognition. In particular, models trained with multitask learning have shown additional performance improvements. However, such multitask models entangle information between the tasks, encoding the mutual dependencies present in label distributions in the real world data used for training. This work explores the disentanglement of multimodal signal representations for the primary task of emotion recognition and a secondary person identification task. In particular, we developed a multitask framework to extract low-dimensional embeddings that aim to capture emotion specific information, while containing minimal information related to person identity. We evaluate three different techniques for disentanglement and report results of up to 13% disentanglement while maintaining emotion recognition performance. Raghuveer Peri, Srinivas Parthasarathy, Charles Bradshaw, Shiva Sundaram |
ICASSP | 4 |
| 2021 | Self-Supervised Learning with Cross-Modal Transformers for Emotion RecognitionabstractEmotion recognition is a challenging task due to limited availability of in-the-wild labeled datasets. Self-supervised learning has shown improvements on tasks with limited labeled datasets in domains like speech and natural language. Models such as BERT learn to incorporate context in word embeddings, which translates to improved performance in downstream tasks like question answering. In this work, we extend self-supervised training to multi-modal applications. We learn multi-modal representations using a transformer trained on the masked language modeling task with audio, visual and text features. This model is fine-tuned on the downstream task of emotion recognition. Our results on the CMU-MOSEI dataset show that this pre-training technique can improve the emotion recognition performance by up to 3% compared to the baseline. Aparna Khare, Srinivas Parthasarathy, Shiva Sundaram |
SLT | 3 |
| 2021 | Detecting Expressions with Multimodal TransformersabstractDeveloping machine learning algorithms to understand person-to-person engagement can result in natural user experiences for communal devices such as Amazon Alexa. Among other cues such as voice activity and gaze, a person’s audio-visual expression that includes tone of the voice and facial expression serves as an implicit signal of engagement between parties in a dialog. This study investigates deep-learning algorithms for audio-visual detection of user’s expression. We first implement an audio-visual baseline model with recurrent layers that shows competitive results compared to current state of the art. Next, we propose the transformer architecture with encoder layers that better integrate audio-visual features for expressions tracking. Performance on the Aff-Wild2 database shows that the proposed methods perform better than baseline architecture with recurrent layers with absolute gains approximately 2% for arousal and valence descriptors. Further, multimodal architectures show significant improvements over models trained on single modalities with gains of up to 3.6%. Ablation studies show the significance of the visual modality for the expression detection on the Aff-Wild2 database. Srinivas Parthasarathy, Shiva Sundaram |
SLT | 2 |
| 2020 | Multimodal and Multiresolution Speech Recognition with TransformersabstractThis paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture.We particularly focus on the scene context provided by the visual information, to ground the ASR.We extract representations for audio features in the encoder layers of the transformer and fuse video features using an additional crossmodal multihead attention layer.Additionally, we incorporate a multitask training criterion for multiresolution ASR, where we train the model to generate both character and subword level transcriptions.Experimental results on the How2 dataset, indicate that multiresolution training can speed up convergence by around 50% and relatively improves word error rate (WER) performance by upto 18% over subword prediction models.Further, incorporating visual information improves performance with relative gains upto 3.76% over audio only models.Our results are comparable to state-of-the-art Listen, Attend and Spell-based architectures. Georgios Paraskevopoulos, Srinivas Parthasarathy, Aparna Khare, Shiva Sundaram |
ACL | 4 |
| 2020 | Robust Multi-Channel Speech Recognition Using Frequency Aligned NetworkabstractConventional speech enhancement technique such as beamforming has known benefits for far-field speech recognition. Our own work in frequency-domain multi-channel acoustic modeling has shown additional improvements by training a spatial filtering layer jointly within an acoustic model. In this paper, we further develop this idea and use frequency aligned network for robust multi-channel automatic speech recognition (ASR). Unlike an affine layer in the frequency domain, the proposed frequency aligned component prevents one frequency bin influencing other frequency bins. We show that this modification not only reduces the number of parameters in the model but also significantly and improves the ASR performance. We investigate effects of frequency aligned network through ASR experiments on the real-world far-field data where users are interacting with an ASR system in uncontrolled acoustic environments. We show that our multi-channel acoustic model with a frequency aligned network shows up to 18% relative reduction in word error rate. Taejin Park, Ken'ichi Kumatani, Minhua Wu, Shiva Sundaram |
ICASSP | 4 |
| 2020 | Fully Learnable Front-End for Multi-Channel Acoustic Modeling Using Semi-Supervised LearningabstractIn this work, we investigated the teacher-student training paradigm to train a fully learnable multi-channel acoustic model for far-field automatic speech recognition (ASR). Using a large offline teacher model trained on beamformed audio, we trained a simpler multi-channel student acoustic model used in the speech recognition system. For the student, both multi-channel feature extraction layers and the higher classification layers were jointly trained using the logits from the teacher model. In our experiments, compared to a baseline model trained on about 600 hours of transcribed data, a relative word-error rate (WER) reduction of about 27.3% was achieved when using an additional 1800 hours of untran-scribed data. We also investigated the benefit of pre-training the multi-channel front end to output the beamformed log-mel filter bank energies (LFBE) using L2 loss. We find that pre-training improves the word error rate by 10.7% when compared to a multi-channel model directly initialized with a beamformer and mel-filter bank coefficients for the front end. Finally, combining pre-training and teacher-student training produces a WER reduction of 31% compared to our baseline. Sanna Wager, Aparna Khare, Minhua Wu, Ken'ichi Kumatani, Shiva Sundaram |
ICASSP | 5 |
| 2020 | Multi-Modal Embeddings Using Multi-Task Learning for Emotion RecognitionabstractGeneral embeddings like word2vec, GloVe and ELMo have shown a lot of success in natural language tasks. The embeddings are typically extracted from models that are built on general tasks such as skip-gram models and natural language generation. In this paper, we extend the work from natural language understanding to multi-modal architectures that use audio, visual and textual information for machine learning tasks. The embeddings in our network are extracted using the encoder of a transformer model trained using multi-task training. We use person identification and automatic speech recognition as the tasks in our embedding generation framework. We tune and evaluate the embeddings on the downstream task of emotion recognition and demonstrate that on the CMU-MOSEI dataset, the embeddings can be used to improve over previous state of the art results. Aparna Khare, Srinivas Parthasarathy, Shiva Sundaram |
INTERSPEECH | 3 |
| 2019 | Multi-geometry Spatial Acoustic Modeling for Distant Speech RecognitionabstractThe use of spatial information with multiple microphones can improve far-field automatic speech recognition (ASR) accuracy. However, conventional microphone array techniques degrade speech enhancement performance when there is an array geometry mismatch between design and test conditions. Moreover, such speech enhancement techniques do not always yield ASR accuracy improvement due to the difference between speech enhancement and ASR optimization objectives. In this work, we propose to unify an acoustic model framework by optimizing spatial filtering and long short-term memory (LSTM) layers from multi-channel (MC) input. Our acoustic model subsumes beamformers with multiple types of array geometry. In contrast to deep clustering methods that treat a neural network as a black box tool, the network encoding the spatial filters can process streaming audio data in real time without the accumulation of target signal statistics. We demonstrate the effectiveness of such MC neural networks through ASR experiments on the real-world far-field data. We show that our two-channel acoustic model can on average reduce word error rates (WERs) by 13.4 and 12.7% compared to a single channel ASR system with the log-mel filter bank energy (LFBE) feature under the matched and mismatched microphone placement conditions, respectively. Our result also shows that our two-channel network achieves a relative WER reduction of over 7.0% compared to conventional beamforming with seven microphones overall. Ken'ichi Kumatani, Minhua Wu, Shiva Sundaram, Nikko Strom, Björn Hoffmeister |
ICASSP | 3 |
| 2019 | Improving Noise Robustness of Automatic Speech Recognition via Parallel Data and Teacher-student LearningabstractFor real-world speech recognition applications, noise robustness is still a challenge. In this work, we adopt the teacher-student (T/S) learning technique using a parallel clean and noisy corpus for improving automatic speech recognition (ASR) performance under multimedia noise. On top of that, we apply a logits selection method which only preserves the k highest values to prevent wrong emphasis of knowledge from the teacher and to reduce bandwidth needed for transferring data. We incorporate up to 8000 hours of untranscribed data for training and present our results on sequence trained models apart from cross entropy trained ones. The best sequence trained student model yields relative word error rate (WER) reductions of approximately 10.1%, 28.7% and 19.6% on our clean, simulated noisy and real test sets respectively comparing to a sequence trained teacher. Ladislav Mosner, Minhua Wu, Anirudh Raju, Sree Hari Krishnan Parthasarathi, Ken'ichi Kumatani, Shiva Sundaram, Roland Maas, Björn Hoffmeister |
ICASSP | 6 |
| 2019 | Frequency Domain Multi-channel Acoustic Modeling for Distant Speech RecognitionabstractConventional far-field automatic speech recognition (ASR) systems typically employ microphone array techniques for speech enhancement in order to improve robustness against noise or reverberation. However, such speech enhancement techniques do not always yield ASR accuracy improvement because the optimization criterion for speech enhancement is not directly relevant to the ASR objective. In this work, we develop new acoustic modeling techniques that optimize spatial filtering and long short-term memory (LSTM) layers from multi-channel (MC) input based on an ASR criterion directly. In contrast to conventional methods, we incorporate array processing knowledge into the acoustic model. Moreover, we initialize the network with beamformers’ coefficients. We investigate effects of such MC neural networks through ASR experiments on the real-world far-field data where users are interacting with an ASR system in uncontrolled acoustic environments. We show that our MC acoustic model can reduce a word error rate (WER) by 16.5% compared to a single channel ASR system with the traditional log-mel filter bank energy (LFBE) feature on average. Our result also shows that our network with the spatial filtering layer on two-channel input achieves a relative WER reduction of 9.5% compared to conventional beamforming with seven microphones. Minhua Wu, Ken'ichi Kumatani, Shiva Sundaram, Nikko Strom, Björn Hoffmeister |
ICASSP | 3 |
| 2018 | Detecting Media Sound Presence in Acoustic Scenes
Constantinos Papayiannis, Justice Amoh, Viktor Rozgic, Shiva Sundaram, Chao Wang 0018 |
INTERSPEECH | 4 |
| 2013 | Affective classification of generic audio clips using regression modelsabstractWe investigate acoustic modeling, feature extraction and feature selection for the problem of affective content recognition of generic, non-speech, non-music sounds. We annotate and analyze a database of generic sounds containing a subset of the BBC sound effects library. We use regression models, longterm features and wrapper-based feature selection to model affect in the continuous 3-D (arousal, valence, dominance) emotional space. The frame-level features for modeling are extracted from each audio clip and combined with functionals to estimate long term temporal patterns over the duration of the clip. Experimental results show that the regression models provide similar categorical performance as the more popular Gaussian Mixture Models. They are also capable of predicting accurate affective ratings on continuous scales, achieving 62-67% 3-class accuracy and 0.69-0.75 correlation with human ratings, higher than comparable numbers in literature. Nikos Malandrakis, Shiva Sundaram, Alexandros Potamianos |
INTERSPEECH | 2 |
| 2013 | An Overview on Perceptually Motivated Audio Indexing and ClassificationabstractAn audio indexing system aims at describing audio content by identifying, labeling, or categorizing different acoustic events. Since the resulting audio classification and indexing is meant for direct human consumption, it is highly desirable that it produces perceptually relevant results. This can be obtained by integrating specific knowledge of the human auditory system in the design process to various extent. In this paper, we highlight some of the important concepts used in audio classification and indexing that are perceptually motivated or that exploit some principles of perception. In particular, we discuss several different strategies to integrate human perception, including: 1) the use of generic audition models; 2) the use of perceptually relevant features for the analysis stage that are perceptually justified either as a component of a hearing model or as being correlated with a perceptual dimension of sound similarity; and 3) the involvement of the user in the audio indexing or classification task. In this paper, we also illustrate some of the recent trends in semantic audio retrieval that approximate higher level perceptual processing and cognitive aspects of human audio recognition capabilities, including affect-based audio retrieval. Gaël Richard, Shiva Sundaram, Shri Narayanan |
Proc. IEEE | 2 |
| 2012 | Latent perceptual mapping with data-driven variable-length acoustic units for template-based speech recognitionabstractIn recent work, we introduced Latent Perceptual Mapping (LPM) [1], a new framework for acoustic modeling suitable for template-like speech recognition. The basic idea is to leverage a reduced dimensionality description of the observations to derive acoustic prototypes that are closely aligned with perceived acoustic events. Our initial work adopted a bag-of-frames strategy to represent relevant acoustic information within speech segments. In this paper, we extend this approach by better integrating temporal information into the LPM feature extraction. Specifically, we use variable-length units to represent acoustic events at the supra-frame level, in order to benefit from finer temporal alignments when deriving the acoustic prototypes. The outcome can be viewed as a generalization of both conventional template-based approaches and recently proposed sparse representation solutions. This extension is experimentally validated on a context-independent phoneme classification task using the TIMIT corpus. Shiva Sundaram, Jerome R. Bellegarda |
ICASSP | 1 |
| 2011 | Experiments in context-independent recognition of non-lexical 'yes' or 'no' responsesabstractWe present our experiments in context-free recognition of non-lexical responses. Non-lexical verbal responses such as mmm-hmm or uh-huh are used by listeners to signal confirmation, uncertainty in understanding, agreement or disagreement in speech-based interaction between humans. Correct recognition of these utterances by speech interfaces can lead to a more natural interaction paradigm with computers. We present our study on both human and automatic recognition of positive (yes) and negative (no) non-lexical utterances. Approximately 3000 isolated utterances from 26 German native speakers were collected in a human-human spoken interaction setting. Our experiments indicate that human recognition accuracy is close to 99% and up to 90% accuracy can be obtained by standard automatic recognition techniques. Shiva Sundaram, Robert Schleicher, Nathalie Diehl |
ICASSP | 1 |
| 2011 | Hotflashes: Thumbnailing videos of social gatherings by detecting camera flash illuminated framesabstractAutomatic annotation of video clips is desired to efficiently thumbnail user generated content available in the Internet. Automatic techniques typically focus on a selected set of inherent video features (such as scene-cut or shot boundaries) that are deemed to be salient. Along this line, the presence of camera flash light illumination is a feature of interest. This event is usually triggered manually (by a photographer) within a scene and the selected frame(s) they occur in are deemed to be interesting in the recording: for instance, the appearance of a celebrity in a party. In this paper we present a method to detect video frames that contain flashes originating from still cameras in user generated video clips. We focus on designing features for flash illumination detection using various measures of luminance change within video sequences. Using the proposed method, we obtain detection performance of approximately 89% which is 8% absolute improvement over the baseline method that uses only change in average illumination. We also illustrate a case where flashes are automatically detected in video clips of social gatherings that can be used for thumbnailing and browsing. Shiva Sundaram, Vladan Velisavljevic |
ICME | 1 |
| 2010 | Using naïve text queries for robust audio information retrievalabstractThe goal of this work is to build an audio information retrieval system which provides users with flexibility in formulating their queries: from audio examples to naïve text. Specifically, the focus of this paper is on using naïve text to create input queries describing the desired information of the users. Using naïve text queries, however, raises interoperability issues between annotation and retrieval processes due to the wide variety of available audio descriptions. In this paper, we propose an intermediate audio description layer (iADL) to solve the interoperability issues between the annotation and retrieval processes. The iADL comprises two axes corresponding to semantic and onomatopoeic descriptions based on human-to-human communication experiments on how humans express sounds verbally. Various text modeling schemes, such as latent semantic analysis (LSA) and latent topic model, are utilized to transform the naïve text onto the proposd iADL. Samuel Kim, Panayiotis G. Georgiou, Shri Narayanan, Shiva Sundaram |
ICASSP | 4 |
| 2010 | Towards evaluation of example-based audio retrieval system using affective dimensionsabstractReal world audio clips contain numerous acoustic sources. The rich acoustic information they carry cannot be fully described with single or even multiple terms about the acoustic sources alone. For instance, the label birds assigned to a birds singing clip that includes sounds of trees and a small river does not properly capture the experience it creates in the person listening to it. In this paper we introduce a novel scheme where the subjective experience of listening to sound clips containing mixture of sources is captured using affective measures. Furthermore, in contrast to the conventional approach of simple label-based methods, the affective ratings are then used to evaluate the performance of an example-based audio retrieval system. We argue that audio retrieval systems can benefit from using affective measures which are well established in experimental psychology, especially when dealing with real world audio clips. We present experimental results of our pilot study to support this motivation where the latent indexing framework has been employed for example-based retrieval on a collection of clips from the BBC sound effects library. The result of our study indicates that using the scheme of affective measures for representation and evaluation is indeed a promising direction to explore. Shiva Sundaram, Robert Schleicher |
ICME | 1 |
| 2010 | Latent perceptual mapping: a new acoustic modeling framework for speech recognition
Shiva Sundaram, Jerome R. Bellegarda |
INTERSPEECH | 1 |
| 2010 | An N-gram model for unstructured audio signals toward information retrievalabstractAn N-gram modeling approach for unstructured audio signals is introduced with applications to audio information retrieval. The proposed N-gram approach aims to capture local dynamic information in acoustic words within the acoustic topic model framework which assumes an audio signal consists of latent acoustic topics and each topic can be interpreted as a distribution over acoustic words. Experimental results on classifying audio clips from BBC Sound Effects Library according to both semantic and onomatopoeic labels indicate that the proposed N-gram approach performs better than using only a bag-of-words approach by providing complementary local dynamic information. Samuel Kim, Shiva Sundaram, Panayiotis G. Georgiou, Shri Narayanan |
MMSP | 2 |
| 2010 | A demonstration of automatic recognition of 'yes' or 'no' non-lexical verbal responses for speech-based interactionabstractIn this paper, we demonstrated our work in context-independent classification of non-lexical `yes' and `no' responses. A typical use case in speech-based interaction with a computer was presented. Related data collection procedure adopted from human-human communication was also described. Shiva Sundaram, Robert Schleicher, Nathalie Diehl |
SLT | 1 |
| 2009 | A divide-and-conquer approach to Latent Perceptual Indexing of audio for large Web 2.0 applicationsabstractIn the recently proposed latent perceptual indexing of audio, a collection of clips is indexed using unit-document frequency measures between a set of reference clusters as units and the clips as the documents. The reference units are derived by clustering the bag-of-feature vectors extracted from the whole audio library using an unsupervised clustering technique. Indexing is achieved through reduced-rank approximation (using singular-value decomposition) of the unit-document co-occurrence measure matrix that is obtained for the given set of reference clusters and the collection of audio clips. In our initial investigation, the k-means algorithm was used to derive the reference units. In this paper, we attempt to reduce the computation load requirements for the k-means algorithm and singular-value decomposition by randomly splitting the training data into smaller sized parts instead of working on it as a whole. We present results of classification experiments on the BBC sound effects library and our results indicate this approach can significantly reduce the computation time without significant loss in classification performance. Shiva Sundaram, Shri Narayanan |
ICME | 1 |
| 2009 | Emotion classification in children's speech using fusion of acoustic and linguistic featuresabstractThis paper describes a system to detect angry vs. non-angry utterances of children who are engaged in dialog with an Aibo robot dog. The system was submitted to the Interspeech2009 Emotion Challenge evaluation. The speech data consist of short utterances of the children’s speech, and the proposed system is designed to detect anger in each given chunk. Frame-based cepstral features, prosodic and acoustic features as well as glottal excitation features are extracted automatically, reduced in dimensionality and classified by means of an artificial neural network and a support vector machine. An automatic speech recognizer transcribes the words in an utterance and yields a separate classification based on the degree of emotional salience of the words. Late fusion is applied to make a final decision on anger vs. non-anger of the utterance. Preliminary results show 75.9% unweighted average recall on the training data and 67.6 % on the test set. Index Terms: speech processing, meta-data extraction, emotion recognition, evaluation Tim Polzehl, Shiva Sundaram, Hamed Ketabdar, Michael Wagner 0004, Florian Metze |
INTERSPEECH | 2 |
| 2009 | Saliency-driven unstructured acoustic scene classification using latent perceptual indexingabstractAutomatic acoustic scene classification of real life, complex and unstructured acoustic scenes is a challenging task as the number of acoustic sources present in the audio stream are unknown and overlapping in time. In this work, we present a novel approach to classification such unstructured acoustic scenes. Motivated by the bottom-up attention model of the human auditory system, salient events of an audio clip are extracted in an unsupervised manner and presented to the classification system. Similar to latent semantic indexing of text documents, the classification system uses unit-document frequency measure to index the clip in a continuous, latent space. This allows for developing a completely class-independent approach to audio classification. Our results on the BBC sound effects library indicates that using the saliency-driven attention selection approach presented in this paper, 17.5% relative improvement can be obtained in frame-based classification and 25% relative improvement can be obtained using the latent audio indexing approach. Ozlem Kalinli, Shiva Sundaram, Shri Narayanan |
MMSP | 2 |
| 2008 | Audio retrieval by latent perceptual indexingabstractWe present a query-by-example audio retrieval framework by indexing audio clips in a generic database as points in a latent perceptual space. First, feature-vectors extracted from the clips in the database are grouped into reference clusters using an unsupervised clustering technique. An audio clip-to-cluster matrix is constructed by keeping count of the number of features that are quantized into each of the reference clusters. By singular-value decomposition of this matrix, each audio clip of the database is mapped into a a point in the latent perceptual space. This is used for indexing the retrieval system. Since each of the initial reference clusters represents a specific perceptual quality in a perceptual space (similar to words that represent specific concepts in the semantic space), querying-by-example results in clips that have similar perceptual qualities. Subjective human evaluation indicates about 75% retrieval performance. Evaluation on semantic categories reveals that the system performance is comparable to other proposed methods. Shiva Sundaram, Shri Narayanan |
ICASSP | 1 |
| 2008 | Classification of sound clips by two schemes: Using onomatopoeia and semantic labelsabstractUsing the recently proposed framework for latent perceptual indexing of audio clips, we present classification of whole clips categorized by two schemes: high-level semantic labels and the mid-level perceptually motivated onomatopoeia labels. First, feature-vectors extracted from the clips in the database are grouped into reference clusters using an unsupervised clustering technique. A unit-document co-occurrence matrix is then obtained by quantizing the feature-vectors extracted from the audio clips into the reference clusters. The audio clips are then mapped to a latent perceptual space by the reduced rank approximation of this matrix. The classification experiments are performed in this representation space using corresponding semantic and onomatopoeic labels of the clips. Using the proposed method, classification accuracy of about sixty percent was obtained when tested on the BBC sound effects library using over twenty categories. Having the two labeling schemes together in a single framework makes the classification system more flexible as each scheme addresses the limitation of the other. These aspects are the main motivation of the work presented here. Shiva Sundaram, Shri Narayanan |
ICME | 1 |
| 2007 | Discriminating Two Types of Noise Sources using Cortical Representation and Dimension Reduction TechniqueabstractContent-based audio classification techniques have focused on classifying events that are both semantically and perceptually distinct (such as speech, music, environmental sounds etc.). However, it is both useful and challenging to develop systems that can also discern sources that are semantically and perceptually close. In this paper we present results of our experiments on discriminating two types of noise sources. Particularly, we focus on machine-generated versus natural noise sources. A bio-inspired tensor representation of audio that models the processing at the primary auditory cortex is used for feature extraction. To handle large tensor feature sets, we use a generalized discriminant analysis method to reduce the dimension. We also present a novel technique of partitioning data into smaller subsets and combining the results of individual analysis before training pattern classifiers. The results of the classification experiments indicate that cortical representation performs 25% better than the common perceptual feature set used in audio classification systems (MFCCs). Shiva Sundaram, Shri Narayanan |
ICASSP (1) | 1 |
| 2007 | Analysis of Audio Clustering using Word DescriptionsabstractWe present an analysis of clustering audio clips using word descriptions that are imitative of sounds. These onomatopoeia words describe the acoustic properties of sources, and they can be useful in annotating a medium that cannot embed audio (e.g. text). First, an audio-to-word relationship is established by manually tagging a variety of audio clips (from a sound effects library) with onomatopoeia words. Using a newly proposed distance metric for word-level similarities, the feature vectors from the audio are clustered according to their tags, resulting in clusters with similarities in their onomatopoeic descriptions. By discriminant analysis of the clusters at the feature level, we present results on separability of these clusters. Our results indicate that by just using onomatopoeic descriptions, meaningful clusters with similar acoustic properties can be formed. However, in terms of audio feature level representation, clusters formed by some word groups such as buzz, fizz etc are better represented by signal features than percussive sounds such as clang, clank, tap. Shiva Sundaram, Shri Narayanan |
ICASSP (2) | 1 |
| 2007 | Experiments in Automatic Genre Classification of Full-length Music Tracks using Audio Activity RateabstractThe activity rate of an audio clip in terms of three defined attributes results in a generic, quantitative measure of various acoustic sources present in it. The objective of this work is to verify if the acoustic structure measured in terms of these three attributes can be used for genre classification of music tracks. For this, we experiment on classification of full-length music tracks by using a dynamic time warping approach for time-series similarity (derived from the activity rate measure) and also a Hidden Markov Model based classifier. The performance of directly using timbral (Mel-frequency Cepstral Coefficients) features is also presented. Using only the activity rate measure we obtain classification performance that is about 35% better than baseline chance and this compares well with other proposed systems that use musical information such as beat histogram or pitch based melody information. Shiva Sundaram, Shri Narayanan |
MMSP | 1 |
| 2006 | Speech Recognition Engineering Issues in Speech to Speech Translation System Design for Low Resource Languages and DomainsabstractEngineering automatic speech recognition (ASR) for speech to speech (S2S) translation systems, especially targeting languages and domains that do not have readily available spoken language resources, is immensely challenging due to a number of reasons. In addition to contending with the conventional data-hungry speech acoustic and language modeling needs, these designs have to accommodate varying requirements imposed by the domain needs and characteristics, target device and usage modality (such as phrase-based, or spontaneous free form interactions, with or without visual feedback) and huge spoken language variability arising due to socio-linguistic and cultural differences of the users. This paper, using case studies of creating speech translation systems between English and languages such as Pashto and Farsi, describes some of the practical issues and the solutions that were developed for multilingual ASR development. These include novel acoustic and language modeling strategies such as language adaptive recognition, active-learning based language modeling, class-based language models that can better exploit resource poor language data, efficient search strategies, including N-best and confidence generation to aid multiple hypotheses translation, use of dialog information and clever interface choices to facilitate ASR, and audio interface design for meeting both usability and robustness requirements Shri Narayanan, Panayiotis G. Georgiou, Abhinav Sethy, Dagen Wang, Murtaza Bulut, Shiva Sundaram, Emil Ettelaie, Sankaranarayanan Ananthakrishnan, Horacio Franco, Kristin Precoda, Dimitra Vergyri, Jing Zheng 0001, Wen Wang 0001, Venkata Ramana Rao Gadde, Martin Graciarena, Victor Abrash, Michael W. Frandsen, Colleen Richey |
ICASSP (5) | 6 |
| 2006 | An attribute-based approach to audio description applied to segmenting vocal sections in popular music songsabstractWe present a descriptive approach for analyzing audio scenes that can comprise a mixture of audio sources. We apply this method to segment popular music songs into vocal and non-vocal sections. Unlike existing methods that directly rely on within-class feature similarities of acoustic sources, the proposed data-driven system is based on a training set where the acoustic sources are grouped by their perceptual or semantic attributes. Our audio analysis approach is based on a quantitative time-varying metric to measure the interaction between acoustic sources present in a scene developed using pattern recognition methods. Using the proposed system that is trained on a general sound effects library, we achieve less than ten percent vocal-section segmentation error and less than five percent false alarm rates when evaluated on a database of popular music recordings that spans four different genres (rock, hiphop, pop, and easy listening) Shiva Sundaram, Shri Narayanan |
MMSP | 1 |
| 2003 | An empirical text transformation method for spontaneous speech synthesizers
Shiva Sundaram, Shri Narayanan |
INTERSPEECH | 1 |