Ramani Duraiswami

dblp:d/RamaniDuraiswami · DBLP profile ↗
← Back
90ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-5596-8460ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 61 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 41 · 1 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
abstract
Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benchmark for assessing audio intelligence in AI systems. MMAU-Pro contains 5,305 instances, where each instance has one or more audios paired with human expert-generated question-answer pairs, spanning speech, sound, music, and their combinations. Unlike existing benchmarks, MMAU-Pro evaluates auditory intelligence across 49 unique skills and multiple complex dimensions, including long-form audio comprehension, spatial audio reasoning, multi-audio understanding, among others. All questions are meticulously designed to require deliberate multi-hop reasoning, including both multiple-choice and open-ended response formats. Importantly, audio data is sourced directly ``from the wild" rather than from existing datasets with known distributions. We evaluate 22 leading open-source and proprietary multimodal AI models, revealing significant limitations: even state-of-the-art models such as Gemini 2.5 Flash and Audio Flamingo 3 achieve only 59.2% and 51.7% accuracy, respectively, approaching random performance in multiple categories. Our extensive analysis highlights specific shortcomings and provides novel insights, offering actionable perspectives for the community to enhance future AI systems' progression toward audio general intelligence. The benchmark and code is available at https://sonalkum.github.io/mmau-pro.
Sonal Kumar, Simon Sedlácek, Vaibhavi Lokegaonkar, Fernando López, Wenyi Yu, Nishit Anand, Hyeonggon Ryu, Lichang Chen, Maxim Plicka, Miroslav Hlavácek, William Fineas Ellingwood, Sathvik Udupa, Siyuan Hou, Allison Ferner, Sara Barahona, Cecilia Bolaños, Satish Rahi, Laura Herrera-Alarcón, Satvik Dixit, Rupali S. Patil, Soham Deshmukh, Lasha Koroshinadze, L. Paola García-Perera, Eleni Zanou, Themos Stafylakis, Joon Son Chung, David F. Harwath, Dinesh Manocha, Alicia Lozano-Diez, Santosh Kesiraju, Sreyan Ghosh, Ramani Duraiswami
AAAI34
2026 FIGMA: Towards FIne-Grained Music retrievAl
abstract
Retrieving music using natural language descriptions has improved with contrastive audio-text models such as CLAP, but current systems remain limited to coarse semantic queries.When descriptions specify fine-grained musical attributes such as tempo, key, chord progression, or rhythmic structure, existing models often fail to retrieve the correct audio.We show that this limitation stems from the contrastive learning objective itself: despite being trained on long captions, CLAP-based models effectively utilize only the first few tokens, discarding much of the information encoded in detailed prompts.Then, we propose FIGMA (FIne-Grained Music RetrievAl), a multi-view contrastive architecture that addresses this limitation by jointly optimizing global audio-text alignment and frame-level, token-wise alignment.This design enables FIGMA to capture both high-level semantic context and finegrained musical attributes within a unified representation space.Moreover, we formalize the task of Fine-Grained Music Retrieval and construct Fine-Grained Music Caption dataset (FGMCaps), a large-scale dataset of 380K music-caption pairs for training along with a 10K test set, both annotated with tempo, key, chord progression, beat count, as well as genre and mood.Extensive experiments demonstrate that FIGMA consistently outperforms existing CLAP-based music retrieval models across multiple music retrieval benchmarks, including out-of-domain evaluations, with relative improvements of up to 73.3%.
Nishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha, Ramani Duraiswami
ACL (1)5
2025 EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
abstract
Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand, Sonal Kumar, Sreyan Ghosh, Ramani Duraiswami, Chirag Agarwal, Dinesh Manocha. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Ashish Seth, Utkarsh Tyagi, Ramaneswaran S., Nishit Anand, Sonal Kumar, Sreyan Ghosh, Ramani Duraiswami, Chirag Agarwal, Dinesh Manocha
EMNLP7
2025 Efficient Spatial Audio Rendering Via Differentiable FIR To IIR Estimation
abstract
The MPEG-H standard for spatial audio proposes the rendering of multiple auditory objects (up to 16) and ambisonics to create a spatial audio scene. Convolution of these (and their early environmental reflections) with user-specific Head Related Impulse Responses (HRIRs), and a treatment of the late tail of the room reverberation, is the gold standard for creating a spatial audio scene. However, this is expensive both in terms of computational time/battery power for finite impulse response (FIR) convolution, and device memory required to store the HRIRs. If quality could be maintained, an implementation with equivalent infinite impulse response (IIR) filters would mitigate these costs. We propose a novel differentiable optimization approach for determination of a IIR filter cascade from a given FIR filter. This is done via an application specific formulation that yields a convex and differentiable cost function for such conversion. We describe our results for spatial audio rendering of HRIR convolution. We compare our work against a recent neural network based HRIR estimation in terms of accuracy and speed. Finally, we implemented our approach in a real-time setting, suitable for implementation on DSP hardware, and conducted a small user study. Results from human participants were positive.
Armin Gerami, Bowen Zhi, Dmitry N. Zotkin, Ramani Duraiswami
ICASSP4
2025 ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
abstract
Open-vocabulary audio-language models, like CLAP [1], offer a promising approach for zero-shot audio classification (ZSAC) by enabling classification with any arbitrary set of categories specified with natural language prompts. In this paper, we propose a simple but effective method to improve ZSAC with CLAP. Specifically, we shift from the conventional method of using prompts with abstract category labels (e.g., Sound of an organ) to prompts that describe sounds using their inherent descriptive features in a diverse context (e.g., The organ’s deep and resonant tones filled the cathedral.). To achieve this, we first propose ReCLAP, a CLAP model trained with rewritten audio captions for improved understanding of sounds in the wild. These rewritten captions describe each sound event in the original caption using their unique discriminative characteristics. ReCLAP outperforms all baselines on both multi-modal audio-text retrieval and ZSAC. Next, to improve zero-shot audio classification with ReCLAP, we propose prompt augmentation. In contrast to the traditional method of employing hand-written template prompts, we generate custom prompts for each unique label in the dataset. These custom prompts first describe the sound event in the label and then employ them in diverse scenes. Our proposed method improves ReCLAP’s performance on ZSAC by 1%-18% and outperforms 1 all baselines by 1% - 55%1.
Sreyan Ghosh, Sonal Kumar, Chandra Kiran Reddy Evuru, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha
ICASSP5
2025 MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
abstract
The ability to comprehend audio—which includes speech, non-speech sounds, and music—is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding models on tasks requiring expert-level knowledge and complex reasoning. MMAU comprises 10k carefully curated audio clips paired with human-annotated natural language questions and answers spanning speech, environmental sounds, and music. It includes information extraction and reasoning questions, requiring models to demonstrate 27 distinct skills across unique and challenging tasks. Unlike existing benchmarks, MMAU emphasizes advanced perception and reasoning with domain-specific knowledge, challenging models to tackle tasks akin to those faced by experts. We assess 18 open-source and proprietary (Large) Audio-Language Models, demonstrating the significant challenges posed by MMAU. Notably, even the most advanced Gemini 2.0 Flash achieves only 59.93% accuracy, and the state-of-the-art open-source Qwen2-Audio achieves only 52.50%, highlighting considerable room for improvement. We believe MMAU will drive the audio and multimodal research community to develop more advanced audio understanding models capable of solving complex audio tasks.
S. Sakshi, Utkarsh Tyagi, Sonal Kumar, Ashish Seth, Ramaneswaran S., Oriol Nieto, Ramani Duraiswami, Sreyan Ghosh, Dinesh Manocha
ICLR7
2025 ProSE: Diffusion Priors for Speech Enhancement
abstract
Sonal Kumar, Sreyan Ghosh, Utkarsh Tyagi, Anton Jeran Ratnarajah, Chandra Kiran Reddy Evuru, Ramani Duraiswami, Dinesh Manocha. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sonal Kumar, Sreyan Ghosh, Utkarsh Tyagi, Anton Ratnarajah, Chandra Kiran Reddy Evuru, Ramani Duraiswami, Dinesh Manocha
NAACL (Long Papers)6
2025 Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
abstract
We present Audio Flamingo 3 (AF3), a fully open state-of-the-art (SOTA) large audio-language model that advances reasoning and understanding across speech, sound, and music. AF3 introduces: (i) AF-Whisper, a unified audio encoder trained using a novel strategy for joint representation learning across all 3 modalities of speech, sound, and music; (ii) flexible, on-demand thinking, allowing the model to do chain-of-thought-type reasoning before answering; (iii) multi-turn, multi-audio chat; (iv) long audio understanding and reasoning (including speech) up to 10 minutes; and (v) voice-to-voice interaction. To enable these capabilities, we propose several large-scale training datasets curated using novel strategies, including AudioSkills-XL, LongAudio-XL, AF-Think, and AF-Chat, and train AF3 with a novel five-stage curriculum-based training strategy. Trained on only open-source audio data, AF3 achieves new SOTA results on over 20+ (long) audio understanding and reasoning benchmarks, surpassing both open-weight and closed-source models trained on much larger datasets.
Sreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar, Zhifeng Kong, Sang-gil Lee, Chao-Han Huck Yang, Ramani Duraiswami, Dinesh Manocha, Rafael Valle, Bryan Catanzaro
NeurIPS8
2024 GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
abstract
Sreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Reddy Evuru, Utkarsh Tyagi, S Sakshi, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Sreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Reddy Evuru, Utkarsh Tyagi, S. Sakshi, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha
EMNLP8
2024 Recap: Retrieval-Augmented Audio Captioning
abstract
We present RECAP (REtrieval-Augmented Audio CAPtioning), a novel and effective audio captioning system that generates captions conditioned on an input audio and other captions similar to the audio retrieved from a datastore. Additionally, our proposed method can transfer to any domain without the need for any additional fine-tuning. To generate a caption for an audio sample, we leverage an audio-text model CLAP [1] to retrieve captions similar to it from a replaceable datastore, which are then used to construct a prompt. Next, we feed this prompt to a GPT-2 decoder and introduce cross-attention layers between the CLAP encoder and GPT-2 to condition the audio for caption generation. Experiments on two benchmark datasets, Clotho and AudioCaps, show that RECAP achieves competitive performance in in-domain settings and significant improvements in out-of-domain settings. Additionally, due to its capability to exploit a large text-captions-only datastore in a training-free fashion, RECAP shows unique capabilities of captioning novel audio events never seen during training and compositional audios with multiple events. To promote research in this space, we also release 150,000+ new weakly labeled captions for AudioSet, AudioCaps, and Clotho1.
Sreyan Ghosh, Sonal Kumar, Chandra Kiran Reddy Evuru, Ramani Duraiswami, Dinesh Manocha
ICASSP4
2024 CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
abstract
A fundamental characteristic of audio is its compositional nature. Audio-language models (ALMs) trained using a contrastive approach (e.g., CLAP) that learns a shared representation between audio and language modalities have improved performance in many downstream applications, including zero-shot audio classification, audio retrieval, etc. However, the ability of these models to effectively perform compositional reasoning remains largely unexplored and necessitates additional research. In this paper, we propose CompA, a collection of two expert-annotated benchmarks with a majority of real-world audio samples, to evaluate compositional reasoning in ALMs. Our proposed CompA-order evaluates how well an ALM understands the order or occurrence of acoustic events in audio, and CompA-attribute evaluates attribute-binding of acoustic events. An instance from either benchmark consists of two audio-caption pairs, where both audios have the same acoustic events but with different compositions. An ALM is evaluated on how well it matches the right audio to the right caption. Using this benchmark, we first show that current ALMs perform only marginally better than random chance, thereby struggling with compositional reasoning. Next, we propose CompA-CLAP, where we fine-tune CLAP using a novel learning method to improve its compositional reasoning abilities. To train CompA-CLAP, we first propose improvements to contrastive training with composition-aware hard negatives, allowing for more focused training. Next, we propose a novel modular contrastive loss that helps the model learn fine-grained compositional understanding and overcomes the acute scarcity of openly available compositional audios. CompA-CLAP significantly improves over all our baseline models on the CompA benchmark, indicating its superior compositional reasoning capabilities.
Sreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi, Chandra Kiran Reddy Evuru, Ramaneswaran S., Sakshi Singh, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha
ICLR9
2024 A Closer Look at the Limitations of Instruction Tuning
abstract
Instruction Tuning (IT), the process of training large language models (LLMs) using instruction-response pairs, has emerged as the predominant method for transforming base pre-trained LLMs into open-domain conversational agents. While IT has achieved notable success and widespread adoption, its limitations and shortcomings remain underexplored. In this paper, through rigorous experiments and an in-depth analysis of the changes LLMs undergo through IT, we reveal various limitations of IT. In particular, we show that (1) IT fails to enhance knowledge or skills in LLMs. LoRA fine-tuning is limited to learning response initiation and style tokens, and full-parameter fine-tuning leads to knowledge degradation. (2) Copying response patterns from IT datasets derived from knowledgeable sources leads to a decline in response quality. (3) Full-parameter fine-tuning increases hallucination by inaccurately borrowing tokens from conceptually similar instances in the IT dataset for generating responses. (4) Popular methods to improve IT do not lead to performance improvements over a simple LoRA fine-tuned model. Our findings reveal that responses generated solely from pre-trained knowledge consistently outperform responses by models that learn any form of new knowledge from IT on open-source datasets. We hope the insights and challenges revealed in this paper inspire future work in related directions.
Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S., Deepali Aneja, Zeyu Jin, Ramani Duraiswami, Dinesh Manocha
ICML7
2024 LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
Sreyan Ghosh, Sonal Kumar, Ashish Seth, Purva Chiniya, Utkarsh Tyagi, Ramani Duraiswami, Dinesh Manocha
INTERSPEECH6
2023 Rapid Audiometric Evaluation for Personalized Headphone Listening
abstract
A novel version of Békésy audiometry is developed that is suitable for remote administration via an app, and for DSP integration into headphones. To demonstrate the efficacy and usefulness of the approach, experiments were performed with 32 participants with a range of ages and hearing losses. In experiment 1, the developed modified Békésy audiometric approach was performed in the laboratory to show that the approach could rapidly and reliably obtain hearing threshold information from 100-16,000 kHz. In experiment 2, participants used a web app to obtain the same information; this information was used to apply gains on a 10-channel graphic equalizer. Participants then evaluated the quality of music in a direct comparison task. In experiment 3, participants were allowed to adjust the proposed correction levels at three frequency regions. They then compared subjective music quality with the correction. The results show that the proposed approach to include audiometric information is promising for the personalization of music experienced over headphones by deriving equalizations that account for the frequency response of both the individual and the headphone.
Matthew J. Goupell, Marjan Davoodian, Sarah Weinstein, David Gadzinski, Dmitry N. Zotkin, Kaushik Sethunath, Ramani Duraiswami
ICASSP7
2022 Towards Fast And Convenient End-To-End HRTF Personalization
abstract
Incorporating individualized head-related transfer functions (HRTFs) into a high fidelity sound engine can further improve the perceived quality and realism of binaurally-rendered spatial audio. Traditional methods to measure individual HRTFs tend to be cumbersome, expensive and require physical access to the subject. To address these issues, we develop a convolutional neural network model that, given a single photo of an ear, predicts pinna landmarks that can be used to extract anthropometric features commonly used for HRTF personalization, and match to a database of subjects whose HRTFs and pictures are available. We propose and evaluate a system utilizing this model to generate an individualized HRTF using a minimal set of easily obtainable measurements: single photographs of both ears, as well as head and ear scale for matching interaural time difference (ITD). To extend the reach of our database we employ ideas from Kendall shape theory to match ears non-dimensionally, match all ears to right ears, and make corresponding changes to the database HRIRs. We also apply HAT models to the HRIRs to provide better matching.
Bowen Zhi, Dmitry N. Zotkin, Ramani Duraiswami
ICASSP3
2018 Sequential Direction Detection for Sound Scene Analysis
abstract
We introduce a novel algorithm for the decomposition of a broadband soundfield into its component plane-waves. The algorithm, termed Sequential Direction Detection, decomposes the sound-field into L plane waves by recursively minimizing an objective function that determines the plane-wave directions, strengths and the number of plane-waves. The algorithm is described and tested on synthetic and real data. Extensions are discussed.
Nail A. Gumerov, Bowen Zhi, Ramani Duraiswami
ICASSP3
2017 Fast interpolation of bandlimited functions
abstract
The nonuniform fast Fourier transform comprises a set of algorithms which approximately interpolate the usual discrete Fourier transform in the time and/or frequency domain. A nonuniform fast Fourier transform based on the fast multipole method was previously developed but passed over in favor of other approaches [1, 2]. This work extends a recent method for computing a periodized fast multipole method [3] with an adaptive algorithm that reduces the required number of multipole-to-local and local-to-local translations by an order of magnitude. This combination improves the speed and accuracy of the original algorithm, and results in an algorithm that is competitive with other nonuniform fast Fourier transforms. Numerical experiments are carried out comparing our implementation with others, demonstrating its viability.
Samuel F. Potter, Nail A. Gumerov, Ramani Duraiswami
ICASSP3
2017 Incident field recovery for an arbitrary-shaped scatterer
abstract
Any acoustic sensor disturbs the spatial acoustic field to certain extent, and a recorded field is different from a field that would have existed if a sensor were absent. Recovery of the original (incident) field is a fundamental task in spatial audio. For some sensor geometries, the disturbance of the field by the sensor can be characterized analytically and its influence can be undone; however, for arbitrary-shaped sensor numerical methods have to be employed. In the current work, the sensor influence on the field is characterized using numerical (specifically, boundary-element) methods, and a framework to recover the incident field, either in the plane-wave or in the spherical wave function basis, is developed. Field recovery in terms of the spherical basis allows the generation of a higher-order Ambisonics representation of the spatial audio scene. Experimental results using a complex-shaped scatterer are presented.
Dmitry N. Zotkin, Nail A. Gumerov, Ramani Duraiswami
ICASSP3
2015 Head-Mounted Display Visualizations to Support Sound Awareness for the Deaf and Hard of Hearing
abstract
Persons with hearing loss use visual signals such as gestures and lip movement to interpret speech. While hearing aids and cochlear implants can improve sound recognition, they generally do not help the wearer localize sound necessary to leverage these visual cues. In this paper, we design and evaluate visualizations for spatially locating sound on a head-mounted display (HMD). To investigate this design space, we developed eight high-level visual sound feedback dimensions. For each dimension, we created 3-12 example visualizations and evaluated these as a design probe with 24 deaf and hard of hearing participants (Study 1). We then implemented a real-time proof-of-concept HMD prototype and solicited feedback from 4 new participants (Study 2). Study 1 findings reaffirm past work on challenges faced by persons with hearing loss in group conversations, provide support for the general idea of sound awareness visualizations on HMDs, and reveal preferences for specific design options. Although preliminary, Study 2 further contextualizes the design probe and uncovers directions for future work.
Dhruv Jain, Leah Findlater, Jamie Gilkeson, Benjamin Holland, Ramani Duraiswami, Dmitry N. Zotkin, Christian Vogler, Jon Froehlich
CHI5
2014 Gaussian process models for HRTF based 3D sound localization
abstract
The human ability to localize sound-source direction using just two receivers is a complex process of direction inference from spectral cues of sound arriving at the ears. While these cues can be described using the well-known head-related transfer function (HRTF) concept, it is unclear as to how densely HRTF must be sampled and whether a higher-order representation is employed in localization. We propose a class of binaural sound source localization models to answer these two questions. First, using the sound received by two ears, we derive several binaural features that are invariant to the sound source signal. Second, these are implicitly mapped to a high-dimensional reproducing kernel Hilbert space via a Gaussian process regression model for feature-direction tuples. Lastly, the features that are most relevant in the model are found via an efficient forward subset-selection method. Experimental results are shown for HRTFs belonging to the CIPIC database.
Yuancheng Luo, Dmitry N. Zotkin, Ramani Duraiswami
ICASSP3
2013 Fast Near-GRID Gaussian Process Regression
abstract
\emphGaussian process regression (GPR) is a powerful non-linear technique for Bayesian inference and prediction. One drawback is its O(N^3) computational complexity for both prediction and hyperparameter estimation for N input points which has led to much work in sparse GPR methods. In case that the covariance function is expressible as a \emphtensor product kernel (TPK) and the inputs form a multidimensional grid, it was shown that the costs for exact GPR can be reduced to a sub-quadratic function of N. We extend these exact fast algorithms to sparse GPR and remark on a connection to \emphGaussian process latent variable models (GPLVMs). In practice, the inputs may also violate the multidimensional grid constraints so we pose and efficiently solve missing and extra data problems for both exact and sparse grid GPR. We demonstrate our method on synthetic, text scan, and magnetic resonance imaging (MRI) data reconstructions.
Yuancheng Luo, Ramani Duraiswami
AISTATS2
2013 Kernel regression for Head-Related Transfer Function interpolation and spectral extrema extraction
abstract
Head-Related Transfer Function (HRTF) representation and interpolation is an important problem in spatial audio. We present a kernel regression method based on Gaussian process (GP) modeling of the joint spatial-frequency relationship between HRTF measurements and obtain a smooth non-linear representation based on data measured over both arbitrary and structured spherical measurement grids. This representation is further extended to the problem of extracting spectral extrema (notches and peaks). We perform HRTF interpolation and spectral extrema extraction using freely available CIPIC HRTF data. Experimental results are shown.
Yuancheng Luo, Dmitry N. Zotkin, Hal Daumé III, Ramani Duraiswami
ICASSP4
2013 A Symmetric Kernel Partial Least Squares Framework for Speaker Recognition
abstract
I-vectors are concise representations of speaker characteristics. Recent progress in i-vectors related research has utilized their ability to capture speaker and channel variability to develop efficient automatic speaker verification (ASV) systems. Inter-speaker relationships in the i-vector space are non-linear. Accomplishing effective speaker verification requires a good modeling of these non-linearities and can be cast as a machine learning problem. Kernel partial least squares (KPLS) can be used for discriminative training in the i-vector space. However, this framework suffers from training data imbalance and asymmetric scoring. We use “one shot similarity scoring” (OSS) to address this. The resulting ASV system (OSS-KPLS) is tested across several conditions of the NIST SRE 2010 extended core data set and compared against state-of-the-art systems: Joint Factor Analysis (JFA), Probabilistic Linear Discriminant Analysis (PLDA), and Cosine Distance Scoring (CDS) classifiers. Improvements are shown.
Balaji Vasan Srinivasan, Yuancheng Luo, Daniel Garcia-Romero, Dmitry N. Zotkin, Ramani Duraiswami
IEEE Trans. Speech Audio Process.5
2012 The UMD-JHU 2011 speaker recognition system
abstract
In recent years, there have been significant advances in the field of speaker recognition that has resulted in very robust recognition systems. The primary focus of many recent developments have shifted to the problem of recognizing speakers in adverse conditions, e.g in the presence of noise/reverberation. In this paper, we present the UMD-JHU speaker recognition system applied on the NIST 2010 SRE task. The novel aspects of our systems are: 1) Improved performance on trials involving different vocal effort via the use of linear-scale features; 2) Expected improved recognition performance in the presence of reverberation and noise via the use of frequency domain perceptual linear predictor and cortical features; 3) A new discriminative kernel partial least squares (KPLS) framework that complements state-of-the-art back-end systems JFA and PLDA to aid in better overall recognition; and 4) Acceleration of JFA, PLDA and KPLS back-ends via distributed computing. The individual components of the system and the fused system are compared against a baseline JFA system and results reported by SRI and MIT-LL on SRE2010.
Daniel Garcia-Romero, Xinhui Zhou, Dmitry N. Zotkin, Balaji Vasan Srinivasan, Yuancheng Luo, Sriram Ganapathy, Samuel Thomas 0001, Sridhar Krishna Nemala, Garimella S. V. S. Sivaram, Majid Mirbagheri, Sri Harish Reddy Mallidi, Thomas Janu, Padmanabhan Rajan, Nima Mesgarani, Mounya Elhilali, Hynek Hermansky, Shihab A. Shamma, Ramani Duraiswami
ICASSP18
2011 Linear versus mel frequency cepstral coefficients for speaker recognition
abstract
Mel-frequency cepstral coefficients (MFCC) have been dominantly used in speaker recognition as well as in speech recognition. However, based on theories in speech production, some speaker characteristics associated with the structure of the vocal tract, particularly the vocal tract length, are reflected more in the high frequency range of speech. This insight suggests that a linear scale in frequency may provide some advantages in speaker recognition over the mel scale. Based on two state-of-the-art speaker recognition back-end systems (one Joint Factor Analysis system and one Probabilistic Linear Discriminant Analysis system), this study compares the performances between MFCC and LFCC (Linear frequency cepstral coefficients) in the NIST SRE (Speaker Recognition Evaluation) 2010 extended-core task. Our results in SRE10 show that, while they are complementary to each other, LFCC consistently outperforms MFCC, mainly due to its better performance in the female trials. This can be explained by the relatively shorter vocal tract in females and the resulting higher formant frequencies in speech. LFCC benefits more in female speech by better capturing the spectral characteristics in the high frequency region. In addition, our results show some advantage of LFCC over MFCC in reverberant speech. LFCC is as robust as MFCC in the babble noise, but not in the white noise. It is concluded that LFCC should be more widely used, at least for the female trials, by the mainstream of the speaker recognition community.
Xinhui Zhou, Daniel Garcia-Romero, Ramani Duraiswami, Carol Y. Espy-Wilson, Shihab A. Shamma
ASRU3
2011 A partial least squares framework for speaker recognition
abstract
Modern approaches to speaker recognition (verification) operate in a space of "supervectors" created via concatenation of the mean vectors of a Gaussian mixture model (GMM) adapted from a universal background model (UBM). In this space, a number of approaches to model inter-class separability and nuisance attribute variability have been proposed. We develop a method for modeling the variability associated with each class (speaker) by using partial-least-squares - a latent variable modeling technique, which isolates the most informative subspace for each speaker. The method is tested on NIST SRE 2008 data and provides promising results. The method is shown to be noise-robust and to be able to efficiently learn the subspace corresponding to a speaker on training data consisting of multiple utterances.
Balaji Vasan Srinivasan, Dmitry N. Zotkin, Ramani Duraiswami
ICASSP3
2011 Kernel Partial Least Squares for Speaker Recognition
abstract
I-vectors are a concise representation of speaker characteristics. Recent advances in speaker recognition have utilized their ability to capture speaker and channel variability to develop efficient recognition engines. Inter-speaker relationships in the i-vector space are non-linear. Accomplishing effective speaker recognition requires a good modeling of these non-linearities and can be cast as a machine learning problem. In this paper, we propose a kernel partial least squares (kernel PLS, or KPLS) framework for modeling speakers in the i-vectors space. The resulting recognition system is tested across several conditions of the NIST SRE 2010 extended core data set and compared against state-of-the-art systems: Joint Factor Analysis (JFA),
Balaji Vasan Srinivasan, Daniel Garcia-Romero, Dmitry N. Zotkin, Ramani Duraiswami
INTERSPEECH4
2011 Scalable fast multipole methods on distributed heterogeneous architectures
abstract
We fundamentally reconsider implementation of the Fast Multipole Method (FMM) on a computing node with a heterogeneous CPU-GPU architecture with multicore CPU(s) and one or more GPU accelerators, as well as on an interconnected cluster of such nodes. The FMM is a divide-and-conquer algorithm that performs a fast N-body sum using a spatial decomposition and is often used in a time-stepping or iterative loop. Using the observation that the local summation and the analysis-based translation parts of the FMM are independent, we map these respectively to the GPUs and CPUs. Careful analysis of the FMM is performed to distribute work optimally between the multicore CPUs and the GPU accelerators. We first develop a single node version where the CPU part is parallelized using OpenMP and the GPU version via CUDA. New parallel algorithms for creating FMM data structures are presented together with load balancing strategies for the single node and distributed multiple-node versions. Our implementation can perform the N-body sum for 128M particles on 16 nodes in 4.23 seconds, a performance not achieved by others in the literature on such clusters.
Nail A. Gumerov, Ramani Duraiswami
SC3
2010 Automatic matched filter recovery via the audio camera
abstract
The sound reaching the acoustic sensor in a realistic environment contains not only the part arriving directly from the sound source but also a number of environmental reflections. The effect of those on the sound is equivalent to a convolution with the room impulse response and can be undone via deconvolution - a technique known as matched filter processing. However, the filter is usually pre-computed in advance using known room geometry and source/receiver positions, and any deviations from those cause the performance to degrade significantly. In this work, an algorithm is proposed to compute the matched filter automatically using an audio camera - a microphone array based system that provides real-time audio images (essentially plots of steered response power in various directions) of environment. Acoustic sources, as well as their significant reflections, are revealed as peaks in the audio image. The reflections are associated with sound source(s) using an acoustic similarity metric, and an approximate matched filter is computed to align the reflections in time with the direct arrival. Preliminary experimental evaluation of the method is performed. It is shown that in case of two sources the reflections are identified correctly, the time delays recovered agree well with those computed from geometric constraints, and that the output SNR improves when the reflections are added coherently to the signal obtained by beamforming directly at the source.
Adam O'Donovan, Ramani Duraiswami, Dmitry N. Zotkin
ICASSP2
2010 Kernelized Rényi distance for speaker recognition
abstract
Speaker recognition systems classify a test signal as a speaker or an imposter by evaluating a matching score between input and reference signals. We propose a new information theoretic approach for computation of the matching score using the Rényi entropy. The proposed entropic distance, the Kernelized Rényi distance (KRD), is formulated in a non-parametric way and the resulting measure is efficiently evaluated in a parallelized fashion on a graphical processor. The distance is then adapted as a scoring function and its performance compared with other popular scoring approaches in a speaker identification and speaker verification framework.
Balaji Vasan Srinivasan, Ramani Duraiswami, Dmitry N. Zotkin
ICASSP2
2010 Plane-Wave Decomposition of Acoustical Scenes Via Spherical and Cylindrical Microphone Arrays
abstract
Spherical and cylindrical microphone arrays offer a number of attractive properties such as direction-independent acoustic behavior and ability to reconstruct the sound field in the vicinity of the array. Beamforming and scene analysis for such arrays is typically done using sound field representation in terms of orthogonal basis functions (spherical/cylindrical harmonics). In this paper, an alternative sound field representation in terms of plane waves is described, and a method for estimating it directly from measurements at microphones is proposed. It is shown that representing a field as a collection of plane waves arriving from various directions simplifies source localization, beamforming, and spatial audio playback. A comparison of the new method with the well-known spherical harmonics based beamforming algorithm is done, and it is shown that both algorithms can be expressed in the same framework but with weights computed differently. It is also shown that the proposed method can be extended to cylindrical arrays. A number of features important for the design and operation of spherical microphone arrays in real applications are revealed. Results indicate that it is possible to reconstruct the sound scene up to order p with p2 microphones spherical array.
Dmitry N. Zotkin, Ramani Duraiswami, Nail A. Gumerov
IEEE Trans. Speech Audio Process.2
2009 Modal expansion of HRTFs: Continuous representation in frequency-range-angle
abstract
This paper proposes a continuous HRTF representation in both 3D spatial and frequency domains. The method is based on the acoustic reciprocity principle and a modal expansion of the wave equation solution to represent the HRTF variations with different variables in separate basis functions. The derived spatial basis modes can achieve HRTF near-field and far-field representation in one formulation. The HRTF frequency components are expanded using Fourier Spherical Bessel series for compact representation. The proposed model can be used to reconstruct HRTFs at any arbitrary position in space and at any frequency point from a finite number of measurements. Analytical simulated and measured HRTFs from a KEMAR are used to validate the model.
Wen Zhang 0002, Thushara D. Abhayapala, Rodney A. Kennedy, Ramani Duraiswami
ICASSP4
2009 Plane-wave decomposition of a sound scene using a cylindrical microphone array
abstract
The analysis for microphone arrays formed by mounting microphones on a sound-hard spherical or cylindrical baffle is typically performed using a decomposition of the sound field in terms of orthogonal basis functions. An alternative representation in terms of plane waves and a method for obtaining the coefficients of such a representation directly from measurements was proposed recently for the case of a spherical array. It was shown that representing the field as a collection of plane waves arriving from various directions simplifies both source localization and beamforming. In this paper, these results are extended to the case of the cylindrical array. Similarly to the spherical array case, localization and beamforming based on plane-wave decomposition perform as well as the traditional orthogonal function based methods while being numerically more stable. Both simulated and experimental results are presented.
Dmitry N. Zotkin, Ramani Duraiswami
ICASSP2
2009 Efficient subset selection via the kernelized Rényi distance
abstract
With improved sensors, the amount of data available in many vision problems has increased dramatically and allows the use of sophisticated learning algorithms to perform inference on the data. However, since these algorithms scale with data size, pruning the data is sometimes necessary. The pruning procedure must be statistically valid and a representative subset of the data must be selected without introducing selection bias. Information theoretic measures have been used for sampling the data, retaining its original information content. We propose an efficient Rényi entropy based subset selection algorithm. The algorithm is first validated and then applied to two sample applications where machine learning and data pruning are used. In the first application, Gaussian process regression is used to learn object pose. Here it is shown that the algorithm combined with the subset selection is significantly more efficient. In the second application, our subset selection approach is used to replace vector quantization in a standard object recognition algorithm, and improvements are shown.
Balaji Vasan Srinivasan, Ramani Duraiswami
ICCV2
2008 Imaging concert hall acoustics using visual and audio cameras
abstract
Using a developed real time audio camera, that uses the output of a spherical microphone array beamformer steered in all directions to create central projection to create acoustic intensity images, we present a technique to measure the acoustics of rooms and halls. A panoramic mosaiced visual image of the space is also create. Since both the visual and the audio camera images are central projection, registration of the acquired audio and video images can be performed using standard computer vision techniques. We describe the technique, and apply it to the examine the relation between acoustical features and architectural details of the Dekelbaum concert hall at the Clarice Smith Performing Arts Center in College Park, MD.
Adam O'Donovan, Ramani Duraiswami, Dmitry N. Zotkin
ICASSP2
2008 Sound field decomposition using spherical microphone arrays
abstract
Spherical microphone arrays offer a number of attractive properties such as direction-independent acoustic behavior and ability to reconstruct the sound field in the vicinity of the array. Such ability is necessary in applications such as ambisonics and recreating auditory environment over headphones. We compare the performance of two scene reconstruction algorithms - one based on least-squares fitting the observed potentials and another based on computing the far-field signature function directly from the microphone measurements. A number of features important for the design and operation of spherical microphone arrays in real applications are revealed. Results indicate that it is possible to reconstruct the sound scene up to order p with p2microphones.
Dmitry N. Zotkin, Ramani Duraiswami, Nail A. Gumerov
ICASSP2
2008 Automatic online tuning for fast Gaussian summation
abstract
Many machine learning algorithms require the summation of Gaussian kernel functions, an expensive operation if implemented straightforwardly. Several methods have been proposed to reduce the computational complexity of evaluating such sums, including tree and analysis based methods. These achieve varying speedups depending on the bandwidth, dimension, and prescribed error, making the choice between methods difficult for machine learning tasks. We provide an algorithm that combines tree methods with the Improved Fast Gauss Transform (IFGT). As originally proposed the IFGT suffers from two problems: (1) the Taylor series expansion does not perform well for very low bandwidths, and (2) parameter selection is not trivial and can drastically affect performance and ease of use. We address the first problem by employing a tree data structure, resulting in four evaluation methods whose performance varies based on the distribution of sources and targets and input parameters such as desired accuracy and bandwidth. To solve the second problem, we present an online tuning approach that results in a black box method that automatically chooses the evaluation method and its parameters to yield the best performance for the input data, desired accuracy, and bandwidth. In addition, the new IFGT parameter selection approach allows for tighter error bounds. Our approach chooses the fastest method at negligible additional cost, and has superior performance in comparisons with previous approaches.
Vlad I. Morariu, Balaji Vasan Srinivasan, Vikas C. Raykar, Ramani Duraiswami, Larry Davis 0001
NIPS4
2008 Tracking Down Under: Following the Satin Bowerbird
abstract
Socio biologists collect huge volumes of video to study animal behavior (our collaborators work with 30,000 hours of video). The scale of these datasets demands the development of automated video analysis tools. Detecting and tracking animals is a critical first step in this process. However, off-the-shelf methods prove incapable of handling videos characterized by poor quality, drastic illumination changes, non-stationary scenery and foreground objects that become motionless for long stretches of time. We improve on existing approaches by taking advantage of specific aspects of this problem: by using information from the entire video we are able to find animals that become motionless for long intervals of time; we make robust decisions based on regional features; for different parts of the image, we tailor the selection of model features, choosing the features most helpful in differentiating the target animal from the background in that part of the image. We evaluate our method, achieving almost 83% tracking accuracy on a more than 200,000 frame dataset of Satin Bowerbird courtship videos.
Aniruddha Kembhavi, Ryan Farrell, Yuancheng Luo, David Jacobs 0001, Ramani Duraiswami, Larry Davis 0001
WACV5
2008 A Fast Algorithm for Learning a Ranking Function from Large-Scale Data Sets
abstract
We consider the problem of learning the ranking function that maximizes a generalization of the Wilcoxon-Mann-Whitney statistic on the training data. Relying on an $\epsilon$-accurate approximation for the error-function, we reduce the computational complexity of each iteration of a conjugate gradient algorithm for learning ranking functions from O(m2) to O(m2), where m is the number of training samples. Experiments on public benchmarks for ordinal regression and collaborative filtering indicate that the proposed algorithm is as accurate as the best available methods in terms of ranking accuracy, when the algorithms are trained on the same data. However, since it is several orders of magnitude faster than the current state-of-the-art approaches, it is able to leverage much larger training datasets.
Vikas C. Raykar, Ramani Duraiswami, Balaji Krishnapuram
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Microphone Arrays as Generalized Cameras for Integrated Audio Visual Processing
abstract
Combinations of microphones and cameras allow the joint audio visual sensing of a scene. Such arrangements of sensors are common in biological organisms and in applications such as meeting recording and surveillance where both modalities are necessary to provide scene understanding. Microphone arrays provide geometrical information on the source location, and allow the sound sources in the scene to be separated and the noise suppressed, while cameras allow the scene geometry and the location and motion of people and other objects to be estimated. In most previous work the fusion of the audio-visual information occurs at a relatively late stage. In contrast, we take the viewpoint that both cameras and microphone arrays are geometry sensors, and treat the microphone arrays as generalized cameras. We employ computer-vision inspired algorithms to treat the combined system of arrays and cameras. In particular, we consider the geometry introduced by a general microphone array and spherical microphone arrays. The latter show a geometry that is very close to central projection cameras, and we show how standard vision based calibration algorithms can be profitably applied to them. Experiments are presented that demonstrate the usefulness of the considered approach.
Adam O'Donovan, Ramani Duraiswami, Jan Neumann
CVPR2
2007 Multimodal Tracking for Smart Videoconferencing and Video Surveillance
abstract
Many applications require the ability to track the 3-D motion of the subjects. We build a particle filter based framework for multimodal tracking using multiple cameras and multiple microphone arrays. In order to calibrate the resulting system, we propose a method to determine the locations of all microphones using at least five loudspeakers and under assumption that for each loudspeaker there exists a microphone very close to it. We derive the maximum likelihood (ML) estimator, which reduces to the solution of the non-linear least squares problem. We verify the correctness and robustness of the multimodal tracker and of the self-calibration algorithm both with Monte-Carlo simulations and on real data from three experimental setups.
Dmitry N. Zotkin, Vikas C. Raykar, Ramani Duraiswami, Larry Davis 0001
CVPR3
2007 Fast Multipole Accelerated Boundary Elements for Numerical Computation of the Head Related Transfer Function
abstract
The numerical computation of head related transfer functions has been attempted by a number of researchers. However, the cost of the computations has meant that usually only low frequencies can be computed and further the computations take inordinately long times. Because of this, comparisons of the computations with measurements are also difficult. We present a fast multipole based iterative preconditioned Krylov solution of a boundary element formulation of the problem and use a new formulation that enables the reciprocity technique to be accurately employed. This allows the calculation to proceed for higher frequencies and larger discretizations. Preliminary results of the computations and of comparisons with measured HRTFs are presented.
Nail A. Gumerov, Ramani Duraiswami, Dmitry N. Zotkin
ICASSP (1)2
2007 Efficient Conversion of X.Y Surround Sound Content to Binaural Head-Tracked Form for HRTF-Enabled Playback
abstract
Binaural presentation of X.Y sound is usually performed using virtual audio principles - that is, by attempting to virtually reproduce the setup of the X+Y loudspeakers in the reference room configuration. The computational cost of such playback is linear in the number of channels in the X.Y setup. We present a novel scheme that computes, offline, a spatio-temporal representation of the sound field in the listening area and store it as a multipole expansion. During head-tracked playback, the binaural signal is obtained by evaluating the multipole expansion at the ear position corresponding to the current user pose, resulting in a fixed playback cost. The representation is further extended to incorporate individualized HRTFs at no additional cost. Simulation results are presented.
Dmitry N. Zotkin, Ramani Duraiswami, Nail A. Gumerov
ICASSP (1)2
2007 Fast Evaluation of the Room Transfer Function Using Multipole Expansion
abstract
Reverberation in rooms is often simulated with the image method due to Allen and Berkley (1979). This method has an asymptotic complexity that is cubic in terms of the simulated reverberation length. When employed in the frequency domain, it is relatively computationally expensive if there are many receivers in the room or if the source or receiver positions are changing with time. The computational complexity of the image method is due to the repeated summation of the fields generated by a large number of image sources. In this paper, a fast method to perform such summations is presented. The method is based on multipole expansion of the monopole source potential. For offline computation of the room transfer function for N image sources and M receiver points, use of the Allen-Berkley algorithm requires O(NM) operations, whereas use of the proposed method requires only O(N+M) operations, resulting in significantly faster computation of reverberant sound fields. The proposed method also has a considerable speed advantage in situations where the room transfer function must be rapidly updated online in response to source/receiver location changes. Simulation results are presented, and algorithm accuracy, speed, and implementation details are discussed. For problems that require frequency-domain computations, the algorithm is found to generate sound fields identical to the ones obtained with the frequency-domain version of the Allen-Berkley algorithm at a fraction of computational cost
Ramani Duraiswami, Dmitry N. Zotkin, Nail A. Gumerov
IEEE Trans. Speech Audio Process.1
2007 Flexible and Optimal Design of Spherical Microphone Arrays for Beamforming
abstract
This paper describes a methodology for designing a flexible and optimal spherical microphone array for beamforming. Using the approach presented, a spherical microphone array can have very flexible layouts of microphones on the spherical surface, yet optimally approximate a desired beampattern of higher order within a specified robustness constraint. Depending on the specified beampattern order, our approach automatically achieves optimal performances in two cases: when the specified beampattern order is reachable within the robustness constraint we achieve a beamformer with optimal approximation of the desired beampattern; otherwise we achieve a beamformer with maximum directivity, both robustly. For efficient implementation, we also developed an adaptive algorithm for computing the beamformer weights. It converges to the optimal performance quickly while exactly satisfying the specified frequency response and robustness constraint in each step. One application of the method is to allow the building of a real-world system, where microphones may not be placeable on regions, such as near cable outlets and/or a mounting base, while having a minimal effect on the performance. Simulation results are presented
Zhiyun Li, Ramani Duraiswami
IEEE Trans. Speech Audio Process.2
2006 Headphone-Based Reproduction of 3D Auditory Scenes Captured by Spherical/Hemispherical Microphone Arrays
abstract
We propose a method to reproduce 3D auditory scenes captured by spherical microphone arrays over headphones. This algorithm employs expansions of the captured sound and the head related transfer function over the sphere and uses the orthonormality of the spherical harmonics. Using a spherical microphone array, we first record the 3D auditory scene, then the recordings are spatially filtered and reproduced through headphones in the orthogonal beam-space of the head related transfer functions (HRTFs). We use the KEMAR HRTF measurements to verify our algorithm. In experiments, we use a hemispherical array for recording. The reproduction results are posted online.
Zhiyun Li, Ramani Duraiswami
ICASSP (5)2
2006 Frequency Independent Flexible Spherical Beamforming Via Rbf Fitting
abstract
We describe a new method for sound analysis using a spherical microphone array without the use of quadrature over the sphere. Quadrature based solutions are very sensitive to the placement of microphones on the sphere, needing measurements to be made at exactly the quadrature positions. We propose to use fitting with band-limited radial basis functions (RBFs) rather than quadrature. Our approach results in frequency independent beamformer weights for flexibly placed microphone locations. Results are demonstrated using both synthetic and real spherical array data.
Arkady Yerukhimovich, Ramani Duraiswami, Nail A. Gumerov, Dmitry N. Zotkin
ICASSP (5)2
2006 Fast optimal bandwidth selection for kernel density estimation
abstract
We propose a computationally efficient ∊-exact approximation algorithm for univariate Gaussian kernel based density derivative estimation that reduces the computational complexity from O(MN) to linear O(N + M). We apply the procedure to estimate the optimal bandwidth for kernel density estimation. We demonstrate the speedup achieved on this problem using the “solve-the-equation plug-in” method, and on exploratory projection pursuit techniques.
Vikas C. Raykar, Ramani Duraiswami
SDM2
2006 3D Structure Recovery and Unwarping of Surfaces Applicable to Planes
Nail A. Gumerov, Ali Zandifar, Ramani Duraiswami, Larry Davis 0001
Int. J. Comput. Vis.3
2005 Efficient Mean-Shift Tracking via a New Similarity Measure
abstract
The mean shift algorithm has achieved considerable success in object tracking due to its simplicity and robustness. It finds local minima of a similarity measure between the color histograms or kernel density estimates of the model and target image. The most typically used similarity measures are the Bhattacharyya coefficient or the Kullback-Leibler divergence. In practice, these approaches face three difficulties. First, the spatial information of the target is lost when the color histogram is employed, which precludes the application of more elaborate motion models. Second, the classical similarity measures are not very discriminative. Third, the sample-based classical similarity measures require a calculation that is quadratic in the number of samples, making real-time performance difficult. To deal with these difficulties we propose a new, simple-to-compute and more discriminative similarity measure in spatial-feature spaces. The new similarity measure allows the mean shift algorithm to track more general motion models in an integrated way. To reduce the complexity of the computation to linear order we employ the recently proposed improved fast Gauss transform. This leads to a very efficient and robust nonparametric spatial-feature tracking algorithm. The algorithm is tested on several image sequences and shown to achieve robust and reliable frame-rate tracking.
Changjiang Yang, Ramani Duraiswami, Larry Davis 0001
CVPR (1)2
2005 The manifolds of spatial hearing
abstract
We present exploratory studies on learning the non-linear manifold structure, in head related impulse responses (HRIRs). We use the recently popular locally linear embedding technique. The lower dimensional manifold encodes the perceptual information in the HRIRs, namely the direction of the sound source. Based on this, we propose a new method for HRIR interpolation. We also propose that the distance between two HRIRs of an individual be taken as the geodesic distance on the learned manifold.
Ramani Duraiswami, Vikas C. Raykar
ICASSP (3)1
2005 A robust and self-reconfigurable design of spherical microphone array for multi-resolution beamforming
abstract
We describe a robust and self-reconfigurable design of a spherical microphone array for beamforming. Our approach achieves a multi-resolution spherical beamformer with performance that is either optimal in the approximation of desired beampattern or is optimal in the directivity achieved, both robustly. Our implementation converges to the optimal performance quickly while exactly satisfying the specified frequency response and robustness constraint in each iteration step without accumulated round-off errors. The advantage of this design lies in its robustness and self-reconfiguration in microphone array reorganization, such as microphone failure, which is highly desirable in online maintenance and anti-terrorism. Design examples and simulation results are presented.
Zhiyun Li, Ramani Duraiswami
ICASSP (4)2
2005 Approximate expressions for the mean and the covariance of the maximum likelihood estimator for acoustic source localization
abstract
Acoustic source localization using multiple microphones can be formulated as a maximum likelihood estimation problem. The estimator is implicitly defined as the minimum of a certain objective function. As a result, we cannot get explicit expressions for the mean and the covariance of the estimator. We derive approximate expressions for the mean vector and covariance matrix of the estimator using Taylor's series expansion of the implicitly defined estimator. The validity of our expressions is verified by Monte-Carlo simulations. We also study the performance of the estimator for different microphone array configurations.
Vikas C. Raykar, Ramani Duraiswami
ICASSP (3)2
2005 Fast Multiple Object Tracking via a Hierarchical Particle Filter
abstract
A very efficient and robust visual object tracking algorithm based on the particle filter is presented. The method characterizes the tracked objects using color and edge orientation histogram features. While the use of more features and samples can improve the robustness, the computational load required by the particle filter increases. To accelerate the algorithm while retaining robustness we adopt several enhancements in the algorithm. The first is the use of integral images for efficiently computing the color features and edge orientation histograms, which allows a large amount of particles and a better description of the targets. Next, the observation likelihood based on multiple features is computed in a coarse-to-fine manner, which allows the computation to quickly focus on the more promising regions. Quasi-random sampling of the particles allows the filter to achieve a higher convergence rate. The resulting tracking algorithm maintains multiple hypotheses and offers robustness against clutter or short period occlusions. Experimental results demonstrate the efficiency and effectiveness of the algorithm for single and multiple object tracking.
Changjiang Yang, Ramani Duraiswami, Larry Davis 0001
ICCV2
2005 A video-based framework for the analysis of presentations/posters
Ali Zandifar, Ramani Duraiswami, Larry Davis 0001
Int. J. Document Anal. Recognit.2
2005 Speaker Localization Using Excitation Source Information in Speech
abstract
This paper presents the results of simulation and real room studies for localization of a moving speaker using information about the excitation source of speech production. The first step in localization is the estimation of time-delay from speech collected by a pair of microphones. Methods for time-delay estimation generally use spectral features that correspond mostly to the shape of vocal tract during speech production. Spectral features are affected by degradations due to noise and reverberation. This paper proposes a method for localizing a speaker using features that arise from the excitation source during speech production. Experiments were conducted by simulating different noise and reverberation conditions to compare the performance of the time-delay estimation and source localization using the proposed method with the results obtained using the spectrum-based generalized cross correlation (GCC) methods. The results show that the proposed method shows lower number of discrepancies in the estimated time-delays. The bias, variance and the root mean square error (RMSE) of the proposed method is consistently equal or less than the GCC methods. The location of a moving speaker estimated using the time-delays obtained by the proposed method are closer to the actual values, than those obtained by the GCC method.
Vikas C. Raykar, Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Ramani Duraiswami
IEEE Trans. Speech Audio Process.4
2005 Processing of reverberant speech for time-delay estimation
abstract
In this paper, we present a method of extracting the time-delay between speech signals collected at two microphone locations. Time-delay estimation from microphone outputs is the first step for many sound localization algorithms, and also for enhancement of speech. For time-delay estimation, speech signals are normally processed using short-time spectral information (either magnitude or phase or both). The spectral features are affected by degradations in speech caused by noise and reverberation. Features corresponding to the excitation source of the speech production mechanism are robust to such degradations. We show that these source features can be extracted reliably from the speech signal. The time-delay estimate can be obtained using the features extracted even from short segments (50-100 ms) of speech from a pair of microphones. The proposed method for time-delay estimation is found to perform better than the generalized cross-correlation (GCC) approach. A method for enhancement of speech is also proposed using the knowledge of the time-delay and the information of the excitation source.
Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Ramani Duraiswami, Dmitry N. Zotkin
IEEE Trans. Speech Audio Process.3
2004 Structure of Applicable Surfaces from Single Views
Nail A. Gumerov, Ali Zandifar, Ramani Duraiswami, Larry Davis 0001
ECCV (3)3
2004 Interpolation and range extrapolation of HRTFs [head related transfer functions]
abstract
The head related transfer function (HRTF) characterizes the scattering properties of a person's anatomy (especially the pinnae, head and torso), and exhibits considerable person-to-person variability. It is usually measured as a part of a tedious experiment, and this leads to the function being sampled at a few angular locations. When the HRTF is needed at intermediate angles, its value must be interpolated. Further, its range dependence is also neglected, which is invalid for nearby sources. Since the HRTF arises from a scattering process, it can be characterized as a solution of a scattering problem. In this paper, we show that by taking this viewpoint and performing some analysis we can express the HRTF in terms of a series of multipole solutions of the Helmholtz equation. This approach leads to a natural solution to the problem of HRTF interpolation. Furthermore, we show that the range-dependence of the HRTF in the near-field can also be obtained by extrapolation from measurements at one range.
Ramani Duraiswami, Dmitry N. Zotkin, Nail A. Gumerov
ICASSP (4)1
2004 Flexible layout and optimal cancellation of the orthonormality error for spherical microphone arrays
abstract
This paper describes an approach to achieving a flexible layout of microphones on the surface of a spherical microphone array for beamforming. Our approach achieves orthonormality of spherical harmonics to higher order for relatively distributed layouts. This gives great flexibility in microphone layout on the spherical surface. One direct advantage is that it makes it much easier to build a real world system, such as those with cable outlets and a mounting base, with minimal effects on the performance. Simulation results are presented.
Zhiyun Li, Ramani Duraiswami, Elena Grassi, Larry Davis 0001
ICASSP (4)2
2004 Automatic position calibration of multiple microphones
abstract
We describe a method to determine automatically the relative three dimensional positions of multiple microphones using at least five loudspeakers in unknown positions. The only assumption we make is that there is a microphone which is very close to a loudspeaker. In our experimental setup, we attach one microphone to each loudspeaker. We derive the maximum likelihood estimator and the solution turns out to be a non-linear least squares problem. A closed form solution which can be used as the initial guess for the minimization routine is derived. We also derive an approximate expression for the covariance of the estimator using the implicit function theorem. Using this, we analyze the performance of the estimator with respect to the positions of the loudspeakers. The algorithm is validated using both Monte-Carlo simulations and a real-time experimental setup.
Vikas C. Raykar, Ramani Duraiswami
ICASSP (4)2
2004 Multi-level fast multipole method for thin plate spline evaluation
Ali Zandifar, Ser-Nam Lim, Ramani Duraiswami, Nail A. Gumerov, Larry Davis 0001
ICIP3
2004 Recording and Reproducing High Order Surround Auditory Scenes for Mixed and Augmented Reality
abstract
Virtual reality systems are largely based on computer graphics and vision technologies. However, sound also plays an important role in human's interaction with the surrounding environment, especially for the visually impaired people. In this paper, we develop the theory of recording and reproducing real-world surround auditory scenes in high orders using specially designed microphone and loudspeaker arrays. It is complementary to vision-based technologies in creating mixed and augmented realities. Design examples and simulations are presented.
Zhiyun Li, Ramani Duraiswami, Larry Davis 0001
ISMAR2
2004 Efficient Kernel Machines Using the Improved Fast Gauss Transform
abstract
The computation and memory required for kernel machines with N train- ing samples is at least O(N 2). Such a complexity is significant even for moderate size problems and is prohibitive for large datasets. We present an approximation technique based on the improved fast Gauss transform to reduce the computation to O(N ). We also give an error bound for the approximation, and provide experimental results on the UCI datasets.
Changjiang Yang, Ramani Duraiswami, Larry Davis 0001
NIPS2
2004 SoftPOSIT: Simultaneous Pose and Correspondence Determination
Philip David, Daniel DeMenthon, Ramani Duraiswami, Hanan Samet
Int. J. Comput. Vis.3
2004 Accelerated speech source localization via a hierarchical search of steered response power
abstract
Accurate and fast localization of multiple speech sound sources is a problem that is of significant interest in applications such as conferencing systems. Recently, approaches that are based on search for local peaks of the steered response power are becoming popular, despite their known computational expense. Based on the observation that the wavelengths of the sound from a speech source are comparable to the dimensions of the space being searched and that the source is broadband, we have developed an efficient search algorithm. Significant speedups are achieved by using coarse-to-fine strategies in both space and frequency. We present applications of the search algorithm to speed up simple delay-and-sum beamforming and steered response power phase-transform weighted (SRP-PHAT) source localization algorithms. A systematic series of comparisons with previous algorithms are made that show that the technique is much faster, robust, and accurate. The performance of the algorithm can be further improved by using constraints from computer vision.
Dmitry N. Zotkin, Ramani Duraiswami
IEEE Trans. Speech Audio Process.2
2004 Rendering localized spatial audio in a virtual auditory space
abstract
High-quality virtual audio scene rendering is required for emerging virtual and augmented reality applications, perceptual user interfaces, and sonification of data. We describe algorithms for creation of virtual auditory spaces by rendering cues that arise from anatomical scattering, environmental scattering, and dynamical effects. We use a novel way of personalizing the head related transfer functions (HRTFs) from a database, based on anatomical measurements. Details of algorithms for HRTF interpolation, room impulse response creation, HRTF selection from a database, and audio scene presentation are presented. Our system runs in real time on an office PC without specialized DSP hardware.
Dmitry N. Zotkin, Ramani Duraiswami, Larry Davis 0001
IEEE Trans. Multim.2
2003 Simultaneous Pose and Correspondence Determination using Line Feature
abstract
We present a new robust line matching algorithm for solving the model-to-image registration problem. Given a model consisting of 3D lines and a cluttered perspective image of this model, the algorithm simultaneously estimates the pose of the model and the correspondences of model lines to image lines. The algorithm combines softassign for determining correspondences and POSIT for determining pose. Integrating these algorithms into a deterministic annealing procedure allows the correspondence and pose to evolve from initially uncertain values to a joint local optimum. This research extends to line features the SoftPOSIT algorithm proposed recently for point features. Lines detected in images are typically more stable than points and are less likely to be produced by clutter and noise, especially in man-made environments. Experiments on synthetic and real imagery with high levels of clutter, occlusion, and noise demonstrate the robustness of the algorithm.
Philip David, Daniel DeMenthon, Ramani Duraiswami, Hanan Samet
CVPR (2)3
2003 Probabilistic Tracking in Joint Feature-Spatial Spaces
abstract
In this paper, we present a probabilistic framework for tracking regions based on their appearance. We exploit the feature-spatial distribution of a region representing an object as a probabilistic constraint to track that region over time. The tracking is achieved by maximizing a similarity-based objective function over transformation space given a nonparametric representation of the joint feature-spatial distribution. Such a representation imposes a probabilistic constraint on the region feature distribution coupled with the region structure, which yields an appearance tracker that is robust to small local deformations and partial occlusion. We present the approach for the general form of joint feature-spatial distributions and apply it to tracking with different types of image features including row intensity, color and image gradient.
Ahmed M. Elgammal, Ramani Duraiswami, Larry Davis 0001
CVPR (1)2
2003 Pitch and timbre manipulations using cortical representation of sound
abstract
The sound received at the ears is processed by humans using signal processing that separates the signal along intensity, pitch and timbre dimensions. Conventional Fourier-based signal processing, while endowed with fast algorithms, is unable to represent a signal easily along the lines of these attributes. We use a recently proposed cortical representation (Elhilali, M. et al., Speech Communications, 2002) to represent and manipulate sound. We briefly overview algorithms for obtaining, manipulating and inverting cortical representation of a sound and describe algorithms for manipulating signal pitch and timbre separately. The algorithms are first used to create the sound of an instrument between a "guitar" and a "trumpet". Applications to creating maximally separable sounds in auditory user interfaces are discussed.
Dmitry N. Zotkin, Shihab A. Shamma, Powen Ru, Ramani Duraiswami, Larry Davis 0001
ICASSP (5)4
2003 Improved Fast Gauss Transform and Efficient Kernel Density Estimation
abstract
Evaluating sums of multivariate Gaussians is a common computational task in computer vision and pattern recognition, including in the general and powerful kernel density estimation technique. The quadratic computational complexity of the summation is a significant barrier to the scalability of this algorithm to practical applications. The fast Gauss transform (FGT) has successfully accelerated the kernel density estimation to linear running time for low-dimensional problems. Unfortunately, the cost of a direct extension of the FGT to higher-dimensional problems grows exponentially with dimension, making it impractical for dimensions above 3. We develop an improved fast Gauss transform to efficiently estimate sums of Gaussians in higher dimensions, where a new multivariate expansion scheme and an adaptive space subdivision technique dramatically improve the performance. The improved FGT has been applied to the mean shift algorithm achieving linear computational complexity. Experimental results demonstrate the efficiency and effectiveness of our algorithm.
Changjiang Yang, Ramani Duraiswami, Nail A. Gumerov, Larry Davis 0001
ICCV2
2003 Mean-shift analysis using quasiNewton methods
abstract
Mean-shift analysis is a general nonparametric clustering technique based on density estimation for the analysis of complex feature spaces. The algorithm consists of a simple iterative procedure that shifts each of the feature points to the nearest stationary point along the gradient directions of the estimated density function. It has been successfully applied to many applications such as segmentation and tracking. However, despite its promising performance, there are applications for which the algorithm converges too slowly to be practical. We propose and implement an improved version of the mean-shift algorithm using quasiNewton methods to achieve higher convergence rates. Another benefit of our algorithm is its ability to achieve clustering even for very complex and irregular feature-space topography. Experimental results demonstrate the efficiency and effectiveness of our algorithm.
Changjiang Yang, Ramani Duraiswami, Daniel DeMenthon, Larry Davis 0001
ICIP (2)2
2003 Using computer vision to generate customized spatial audio
abstract
Creating high quality virtual spatial audio over headphones requires real-time head tracking, personalized head-related transfer functions (HRTFs) and customized room response models. While there are expensive solutions to address these issues based on costly head trackers, measured personalized HRTFs and room responses, these are not suitable for widespread or easy deployment and use. We report on the development of a system that uses computer vision to produce customizable models for both the HRTF and the room response, and to achieve head-tracking. The system uses relatively inexpensive cameras and widely available personal computers. Computer-vision based anthropometric measurements of the head, torso, and the external ears are used for HRTF customization. For low-frequency HRTF customization we employ a simple head-and-torso model developed recently [V. R. Algazi et al., 2002]. For high frequency customization we employ measured pinna characteristics as an index into a database of HRTFs [D. N. Zotkin et al., 2002]. For head tracking we employ an online implementation of the POSIT algorithm [D. DeMenthon and L. Davis, 1995] along with active markers to compute head pose in real-time. The system provides an enhanced virtual listening experience at low cost.
Ankur Mohan, Ramani Duraiswami, Dmitry N. Zotkin, Daniel DeMenthon, Larry Davis 0001
ICME2
2003 Pitch and timbre manipulations using cortical representation of sound
abstract
The sound receiver at the ears is processed by humans using signal processing that separate the signal along intensity, pitch and timbre dimensions. Conventional Fourier-based signal processing, while endowed with fast algorithms, is unable to easily represent signal along these attributes. In this paper we use a cortical representation to represent the manipulate sound. We briefly overview algorithms for obtaining, manipulating and inverting cortical representation of sound and describe algorithms for manipulating signal pitch and timbre separately. The algorithms are first used to create sound of an instrument between a guitar and a trumpet. Applications to creating maximally separable sounds in auditory user interfaces are discussed.
Dmitry N. Zotkin, Shihab A. Shamma, Powen Ru, Ramani Duraiswami, Larry Davis 0001
ICME4
2003 Tracking a moving speaker using excitation source information
abstract
Microphone arrays are widely used to detect, locate, and track a stationary or moving speaker. The first step is to estimate the time delay, between the speech signals received by a pair of microphones. Conventional methods like generalized crosscorrelation are based on the spectral content of the vocal tract system in the speech signal. The spectral content of the speech signal is affected due to degradations in the speech signal caused by noise and reverberation. However, features corresponding to the excitation source of speech are less affected by such degradations. This paper proposes a novel method to estimate the time delays using the excitation source information in speech. The estimated delays are used to get the position of the moving speaker. The proposed method is compared with the spectrumbased approach using real data from a microphone array setup. 1.
Vikas C. Raykar, Ramani Duraiswami, Bayya Yegnanarayana, S. R. Mahadeva Prasanna
INTERSPEECH2
2003 Efficient Kernel Density Estimation Using the Fast Gauss Transform with Applications to Color Modeling and Tracking
abstract
Many vision algorithms depend on the estimation of a probability density function from observations. Kernel density estimation techniques are quite general and powerful methods for this problem, but have a significant disadvantage in that they are computationally intensive. In this paper, we explore the use of kernel density estimation with the fast Gauss transform (FGT) for problems in vision. The FGT allows the summation of a mixture of ill Gaussians at N evaluation points in O(M+N) time, as opposed to O(MN) time for a naive evaluation and can be used to considerably speed up kernel density estimation. We present applications of the technique to problems from image segmentation and tracking and show that the algorithm allows application of advanced statistical techniques to solve practical vision problems in real-time with today's computers.
Ahmed M. Elgammal, Ramani Duraiswami, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 SoftPOSIT: Simultaneous Pose and Correspondence Determination
Philip David, Daniel DeMenthon, Ramani Duraiswami, Hanan Samet
ECCV (3)3
2002 Numerical study of the influence of the torso on the HRTF
abstract
Understanding and simplified modeling of how humans generate cues from the scattering of sounds from their bodies holds the key to many applications in spatial audio. Scattering from the torso is thought to provide cues for source localization. A simplified two-sphere model has been proposed to understand the influence of the torso, and experimental studies have been performed. The availability of a validated numerical model can enable further parametric study and improved understanding using this model. Here we provide such a numerical model using the multipole method. Our method is validated by comparison with experimental data and commercial boundary element software. We find that most features of this “snowman” HRTF are reproduced by a simple model that places the “head” in the field of the source and the scattered field from the torso.
Nail A. Gumerov, Ramani Duraiswami, Zhihui Tang
ICASSP2
2002 Creation of virtual auditory spaces
abstract
High-quality virtual audio scene rendering is a must for emerging virtual/augmented reality applications and for perceptual user interfaces. We describe algorithms for creation of virtual auditory spaces using measured and non-individualized HRTFs and head tracking. Details of algorithms for HRTF interpolation, room impulse response creation, and audio scene presentation are presented. Tests show that individuals externalize well, and find our interface natural. The system runs in real time with latency of less than 30 ms on an office PC without specialized DSP.
Dmitry N. Zotkin, Ramani Duraiswami, Larry Davis 0001
ICASSP2
2002 A Video Based Interface to Textual Information for the Visually Impaired
abstract
We describe the development of an interface to textual information for the visually impaired that uses video, image processing, optical-character-recognition (OCR) and text-to-speech (TTS). The video provides a sequence of low resolution images in which text must be detected, rectified and converted into high resolution rectangular blocks that are capable of being analyzed via off-the-shelf OCR. To achieve this, various problems related to feature detection, mosaicing, auto-focus, zoom, and systems integration were solved in the development of the system.
Ali Zandifar, Ramani Duraiswami, Antoine Chahine, Larry Davis 0001
ICMI2
2002 Background and foreground modeling using nonparametric kernel density estimation for visual surveillance
abstract
Automatic understanding of events happening at a site is the ultimate goal for many visual surveillance systems. Higher level understanding of events requires that certain lower level computer vision tasks be performed. These may include detection of unusual motion, tracking targets, labeling body parts, and understanding the interactions between people. To achieve many of these tasks, it is necessary to build representations of the appearance of objects in the scene. This paper focuses on two issues related to this problem. First, we construct a statistical representation of the scene background that supports sensitive detection of moving objects in the scene, but is robust to clutter arising out of natural scene variations. Second, we build statistical representations of the foreground regions (moving objects) that support their tracking and support occlusion reasoning. The probability density functions (pdfs) associated with the background and foreground are likely to vary from image to image and will not in general have a known parametric form. We accordingly utilize general nonparametric kernel density estimation techniques for building these statistical representations of the background and the foreground. These techniques estimate the pdf directly from the data without any assumptions about the underlying distributions. Example results from applications are presented.
Ahmed M. Elgammal, Ramani Duraiswami, David Harwood, Larry Davis 0001
Proc. IEEE2
2001 Efficient Non-Parametric Adaptive Color Modeling Using Fast Gauss Transform
abstract
Modeling the color distribution of a homogeneous region is used extensively for object tracking and recognition applications. The color distribution of an object represents a feature that is robust to partial occlusion, scaling and object deformation. A variety of parametric and non-parametric statistical techniques have been used to model color distributions. In this paper we present a non-parametric color modeling approach based on kernel density estimation as well as a computational framework for efficient density estimation. Theoretically, our approach is general since kernel density estimators can converge to any density shape with sufficient samples. Therefore, this approach is suitable to model the color distribution of regions with patterns and mixture of colors. Since kernel density estimation techniques are computationally expensive, the paper introduces the use of the fast Gauss transform for efficient computation of the color densities. We show that this approach can be used successfully for color-based segmentation of body parts as well as segmentation of many people under occlusion.
Ahmed M. Elgammal, Ramani Duraiswami, Larry Davis 0001
CVPR (2)2
2001 Active speech source localization by a dual coarse-to-fine search
abstract
Accurate and fast localization of multiple speech sound sources is a significant problem in videoconferencing systems. Based on the observation that the wavelengths of the sound from a speech source are comparable to the dimensions of the space being searched, and that the source is broadband, we develop an efficient search strategy that finds the source(s) in a given space. The search is made efficient by using coarse-to-fine strategies in both space and frequency. The algorithm is shown to be robust compared to typical delay-based estimators and fast enough for real-time implementation. Its performance can be further improved by using constraints from computer vision.
Ramani Duraiswami, Dmitry N. Zotkin, Larry Davis 0001
ICASSP1
2001 Multimodal localization of a flying bat
abstract
We present a new multimodal system that combines stereoscopic and audio-based source localization to track the position of a flying bat. Also presented are novel algorithms for audio source localization. The bat was allowed to fly in an anechoic room and monitored by two high-speed video cameras. The vocalizations of the bat were simultaneously recorded from six microphones. The data was then processed offline to localize the source and reconstruct the trajectory of the bat. We compare the performance of the localization algorithm with the position data obtained from stereoscopic pictures of the bat. The results confirm that the stereoscopic analysis and the audio localization are in good agreement. This system opens up new possibilities for performing multimodal research, and developing more tightly integrated algorithms.
Kaushik Ghose, Dmitry N. Zotkin, Ramani Duraiswami, Cynthia F. Moss
ICASSP3
2001 Modeling the effect of a nearby boundary on the HRTF
abstract
Understanding and simplified modeling of the head related transfer function (HRTF) holds the key to many applications in spatial audio. We develop an analytical solution to the problem of scattering of sound from a sphere in the vicinity of an infinite plane. Using this solution we study the influence of a nearby scattering rigid surface, on a spherical model for the HRTF.
Nail A. Gumerov, Ramani Duraiswami
ICASSP2
2001 Attentive Toys
Ismail Haritaoglu, Alex Cozzi, David Koons, Myron Flickner, Dmitry N. Zotkin, Ramani Duraiswami, Yaser Yacoob
ICME6
2001 Multimodal Tracking For Smart Videoconferencing
abstract
Many applications require the ability to track the 3-D motion of the subjects. We build a particle filter based framework for multimodal tracking using multiple cameras and multiple microphone arrays. In order to calibrate the resulting system, we propose a method to determine the locations of all microphones using at least five loudspeakers and under assumption that for each loudspeaker there exists a microphone very close to it. We derive the maximum likelihood (ML) estimator, which reduces to the solution of the non-linear least squares problem. We verify the correctness and robustness of the multimodal tracker and of the self-calibration algorithm both with Monte-Carlo simulations and on real data from three experimental setups. 1.
Dmitry N. Zotkin, Ramani Duraiswami, Harsh Nanda, Larry Davis 0001
ICME2
2000 Quasi-Random Sampling for Condensation
Vasanth Philomin, Ramani Duraiswami, Larry Davis 0001
ECCV (2)2
2000 Tracking Humans from a Moving Platform
abstract
Research at the Computer Vision Laboratory at the University of Maryland has focussed on developing algorithms and systems that can look at humans and recognize their activities in near real-time. Our earlier implementation while quite successful, was restricted to applications with a fixed camera. In this paper we present some recent work that removes this restriction. Such systems are required for machine vision from moving platforms such as robots, intelligent vehicles, and unattended large field of regard cameras with a small field of view. Our approach is based on the use of a deformable shape model for humans coupled with a novel variant of the condensation algorithm that uses quasi-random sampling for efficiency. This allows the use of simple motion models which results in algorithm robustness, enabling us to handle unknown camera/human motion with unrestricted camera viewing angles. We present the details of our human tracking algorithms and some examples from pedestrian tracking and automated surveillance.
Larry Davis 0001, Vasanth Philomin, Ramani Duraiswami
ICPR3
2000 An audio-video front-end for multimedia applications
abstract
Applications such as video gaming, virtual reality, multimodal user interfaces and videoconferencing, require systems that can locate and track persons in a room through a combination of visual and audio cues, enhance the sound that they produce, and perform identification. We describe the development of a particular multimodal sensor fusion system that is portable, runs in real time and achieves these objectives. The system employs novel algorithms for acoustical source location, video-based person tracking and overall system control, which are also described.
Dmitry N. Zotkin, Ramani Duraiswami, Larry Davis 0001, Ismail Haritaoglu
SMC2