EDBT 2026 Demo / reviewers in the wild / expert
Dmitry N. Zotkin
dblp:25/3254
· DBLP profile ↗
37ranked-venue papers
14as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 10 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 3 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
6 papers |
Audio and music processing · 99% Virtual and augmented reality · 1% | |
| Human-computer interaction and pervasive computing
1 paper |
Accessibility and assistive technology · 77% Immersive interaction · 23% | |
| Computer networks
1 paper |
Wireless sensing and localization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 87% High-performance computing · 13% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing
speaker recognition |
0.2 | 1 | 2013 | A Symmetric Kernel Partial Least Squares Framework for Speaker Recognition · IEEE Trans. Speech Audio Process. 2013 |
Audio and music processing › speaker recognition
speaker verification |
0.2 | 1 | 2013 | A Symmetric Kernel Partial Least Squares Framework for Speaker Recognition · IEEE Trans. Speech Audio Process. 2013 |
Audio and music processing
microphone array processing |
0.2 | 2 | 2010 | Plane-Wave Decomposition of Acoustical Scenes Via Spherical and Cylindrical Microphone Arrays · IEEE Trans. Speech Audio Process. 2010 Processing of reverberant speech for time-delay estimation · IEEE Trans. Speech Audio Process. 2005 |
Audio and music processing
spatial audio |
0.2 | 2 | 2010 | Plane-Wave Decomposition of Acoustical Scenes Via Spherical and Cylindrical Microphone Arrays · IEEE Trans. Speech Audio Process. 2010 Rendering localized spatial audio in a virtual auditory space · IEEE Trans. Multim. 2004 |
Audio and music processing › spatial audio
sound field reproduction |
0.1 | 1 | 2010 | Plane-Wave Decomposition of Acoustical Scenes Via Spherical and Cylindrical Microphone Arrays · IEEE Trans. Speech Audio Process. 2010 |
Audio and music processing › room acoustics
room acoustics simulation |
0.1 | 1 | 2007 | Fast Evaluation of the Room Transfer Function Using Multipole Expansion · IEEE Trans. Speech Audio Process. 2007 |
Audio and music processing › room acoustics
room transfer function |
0.1 | 1 | 2007 | Fast Evaluation of the Room Transfer Function Using Multipole Expansion · IEEE Trans. Speech Audio Process. 2007 |
Wireless sensing and localization
acoustic source localization |
0.1 | 1 | 2007 | Multimodal Tracking for Smart Videoconferencing and Video Surveillance · CVPR 2007 |
Wireless sensing and localization
self-calibration |
0.1 | 1 | 2007 | Multimodal Tracking for Smart Videoconferencing and Video Surveillance · CVPR 2007 |
Immersive interaction
head-mounted display |
0.1 | 1 | 2015 | Head-Mounted Display Visualizations to Support Sound Awareness for the Deaf and Hard of Hearing · CHI 2015 |
Audio and music processing
speech enhancement |
0.1 | 1 | 2005 | Processing of reverberant speech for time-delay estimation · IEEE Trans. Speech Audio Process. 2005 |
Audio and music processing
speech processing |
0.1 | 1 | 2005 | Processing of reverberant speech for time-delay estimation · IEEE Trans. Speech Audio Process. 2005 |
Audio and music processing › sound source localization
time delay estimation |
0.1 | 1 | 2005 | Processing of reverberant speech for time-delay estimation · IEEE Trans. Speech Audio Process. 2005 |
Audio and music processing › spatial audio › head-related transfer function
head-related transfer function personalization |
0.0 | 1 | 2004 | Rendering localized spatial audio in a virtual auditory space · IEEE Trans. Multim. 2004 |
Audio and music processing
sound source localization |
0.0 | 1 | 2004 | Accelerated speech source localization via a hierarchical search of steered response power · IEEE Trans. Speech Audio Process. 2004 |
Audio and music processing › spatial audio
virtual auditory space |
0.0 | 1 | 2004 | Rendering localized spatial audio in a virtual auditory space · IEEE Trans. Multim. 2004 |
Cloud and datacenter computing › job scheduling › batch scheduling
backfilling |
0.0 | 1 | 1999 | Job-Length Estimation and Performance in Backfilling Schedulers · HPDC 1999 |
Cloud and datacenter computing
job scheduling |
0.0 | 1 | 1999 | Job-Length Estimation and Performance in Backfilling Schedulers · HPDC 1999 |
Methods — techniques the papers use, named apart from their topics
user study · 0.2design probe · 0.2probabilistic linear discriminant analysis · 0.2one shot similarity scoring · 0.2kernel partial least squares · 0.2joint factor analysis · 0.2spherical harmonics · 0.1cylindrical harmonics · 0.1beamforming · 0.1particle filter · 0.1nonlinear least squares · 0.1multipole expansion · 0.1maximum likelihood estimation · 0.1image method · 0.1excitation source feature extraction · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Spatial Audio Rendering Via Differentiable FIR To IIR EstimationabstractThe MPEG-H standard for spatial audio proposes the rendering of multiple auditory objects (up to 16) and ambisonics to create a spatial audio scene. Convolution of these (and their early environmental reflections) with user-specific Head Related Impulse Responses (HRIRs), and a treatment of the late tail of the room reverberation, is the gold standard for creating a spatial audio scene. However, this is expensive both in terms of computational time/battery power for finite impulse response (FIR) convolution, and device memory required to store the HRIRs. If quality could be maintained, an implementation with equivalent infinite impulse response (IIR) filters would mitigate these costs. We propose a novel differentiable optimization approach for determination of a IIR filter cascade from a given FIR filter. This is done via an application specific formulation that yields a convex and differentiable cost function for such conversion. We describe our results for spatial audio rendering of HRIR convolution. We compare our work against a recent neural network based HRIR estimation in terms of accuracy and speed. Finally, we implemented our approach in a real-time setting, suitable for implementation on DSP hardware, and conducted a small user study. Results from human participants were positive. Armin Gerami, Bowen Zhi, Dmitry N. Zotkin, Ramani Duraiswami |
ICASSP | 3 |
| 2023 | Rapid Audiometric Evaluation for Personalized Headphone ListeningabstractA novel version of Békésy audiometry is developed that is suitable for remote administration via an app, and for DSP integration into headphones. To demonstrate the efficacy and usefulness of the approach, experiments were performed with 32 participants with a range of ages and hearing losses. In experiment 1, the developed modified Békésy audiometric approach was performed in the laboratory to show that the approach could rapidly and reliably obtain hearing threshold information from 100-16,000 kHz. In experiment 2, participants used a web app to obtain the same information; this information was used to apply gains on a 10-channel graphic equalizer. Participants then evaluated the quality of music in a direct comparison task. In experiment 3, participants were allowed to adjust the proposed correction levels at three frequency regions. They then compared subjective music quality with the correction. The results show that the proposed approach to include audiometric information is promising for the personalization of music experienced over headphones by deriving equalizations that account for the frequency response of both the individual and the headphone. Matthew J. Goupell, Marjan Davoodian, Sarah Weinstein, David Gadzinski, Dmitry N. Zotkin, Kaushik Sethunath, Ramani Duraiswami |
ICASSP | 5 |
| 2022 | Towards Fast And Convenient End-To-End HRTF PersonalizationabstractIncorporating individualized head-related transfer functions (HRTFs) into a high fidelity sound engine can further improve the perceived quality and realism of binaurally-rendered spatial audio. Traditional methods to measure individual HRTFs tend to be cumbersome, expensive and require physical access to the subject. To address these issues, we develop a convolutional neural network model that, given a single photo of an ear, predicts pinna landmarks that can be used to extract anthropometric features commonly used for HRTF personalization, and match to a database of subjects whose HRTFs and pictures are available. We propose and evaluate a system utilizing this model to generate an individualized HRTF using a minimal set of easily obtainable measurements: single photographs of both ears, as well as head and ear scale for matching interaural time difference (ITD). To extend the reach of our database we employ ideas from Kendall shape theory to match ears non-dimensionally, match all ears to right ears, and make corresponding changes to the database HRIRs. We also apply HAT models to the HRIRs to provide better matching. Bowen Zhi, Dmitry N. Zotkin, Ramani Duraiswami |
ICASSP | 2 |
| 2017 | Incident field recovery for an arbitrary-shaped scattererabstractAny acoustic sensor disturbs the spatial acoustic field to certain extent, and a recorded field is different from a field that would have existed if a sensor were absent. Recovery of the original (incident) field is a fundamental task in spatial audio. For some sensor geometries, the disturbance of the field by the sensor can be characterized analytically and its influence can be undone; however, for arbitrary-shaped sensor numerical methods have to be employed. In the current work, the sensor influence on the field is characterized using numerical (specifically, boundary-element) methods, and a framework to recover the incident field, either in the plane-wave or in the spherical wave function basis, is developed. Field recovery in terms of the spherical basis allows the generation of a higher-order Ambisonics representation of the spatial audio scene. Experimental results using a complex-shaped scatterer are presented. Dmitry N. Zotkin, Nail A. Gumerov, Ramani Duraiswami |
ICASSP | 1 |
| 2015 | Head-Mounted Display Visualizations to Support Sound Awareness for the Deaf and Hard of HearingabstractPersons with hearing loss use visual signals such as gestures and lip movement to interpret speech. While hearing aids and cochlear implants can improve sound recognition, they generally do not help the wearer localize sound necessary to leverage these visual cues. In this paper, we design and evaluate visualizations for spatially locating sound on a head-mounted display (HMD). To investigate this design space, we developed eight high-level visual sound feedback dimensions. For each dimension, we created 3-12 example visualizations and evaluated these as a design probe with 24 deaf and hard of hearing participants (Study 1). We then implemented a real-time proof-of-concept HMD prototype and solicited feedback from 4 new participants (Study 2). Study 1 findings reaffirm past work on challenges faced by persons with hearing loss in group conversations, provide support for the general idea of sound awareness visualizations on HMDs, and reveal preferences for specific design options. Although preliminary, Study 2 further contextualizes the design probe and uncovers directions for future work. Dhruv Jain, Leah Findlater, Jamie Gilkeson, Benjamin Holland, Ramani Duraiswami, Dmitry N. Zotkin, Christian Vogler, Jon Froehlich |
CHI | 6 |
| 2014 | Gaussian process models for HRTF based 3D sound localizationabstractThe human ability to localize sound-source direction using just two receivers is a complex process of direction inference from spectral cues of sound arriving at the ears. While these cues can be described using the well-known head-related transfer function (HRTF) concept, it is unclear as to how densely HRTF must be sampled and whether a higher-order representation is employed in localization. We propose a class of binaural sound source localization models to answer these two questions. First, using the sound received by two ears, we derive several binaural features that are invariant to the sound source signal. Second, these are implicitly mapped to a high-dimensional reproducing kernel Hilbert space via a Gaussian process regression model for feature-direction tuples. Lastly, the features that are most relevant in the model are found via an efficient forward subset-selection method. Experimental results are shown for HRTFs belonging to the CIPIC database. Yuancheng Luo, Dmitry N. Zotkin, Ramani Duraiswami |
ICASSP | 2 |
| 2013 | Kernel regression for Head-Related Transfer Function interpolation and spectral extrema extractionabstractHead-Related Transfer Function (HRTF) representation and interpolation is an important problem in spatial audio. We present a kernel regression method based on Gaussian process (GP) modeling of the joint spatial-frequency relationship between HRTF measurements and obtain a smooth non-linear representation based on data measured over both arbitrary and structured spherical measurement grids. This representation is further extended to the problem of extracting spectral extrema (notches and peaks). We perform HRTF interpolation and spectral extrema extraction using freely available CIPIC HRTF data. Experimental results are shown. Yuancheng Luo, Dmitry N. Zotkin, Hal Daumé III, Ramani Duraiswami |
ICASSP | 2 |
| 2013 | A Symmetric Kernel Partial Least Squares Framework for Speaker RecognitionabstractI-vectors are concise representations of speaker characteristics. Recent progress in i-vectors related research has utilized their ability to capture speaker and channel variability to develop efficient automatic speaker verification (ASV) systems. Inter-speaker relationships in the i-vector space are non-linear. Accomplishing effective speaker verification requires a good modeling of these non-linearities and can be cast as a machine learning problem. Kernel partial least squares (KPLS) can be used for discriminative training in the i-vector space. However, this framework suffers from training data imbalance and asymmetric scoring. We use “one shot similarity scoring” (OSS) to address this. The resulting ASV system (OSS-KPLS) is tested across several conditions of the NIST SRE 2010 extended core data set and compared against state-of-the-art systems: Joint Factor Analysis (JFA), Probabilistic Linear Discriminant Analysis (PLDA), and Cosine Distance Scoring (CDS) classifiers. Improvements are shown. Balaji Vasan Srinivasan, Yuancheng Luo, Daniel Garcia-Romero, Dmitry N. Zotkin, Ramani Duraiswami |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | The UMD-JHU 2011 speaker recognition systemabstractIn recent years, there have been significant advances in the field of speaker recognition that has resulted in very robust recognition systems. The primary focus of many recent developments have shifted to the problem of recognizing speakers in adverse conditions, e.g in the presence of noise/reverberation. In this paper, we present the UMD-JHU speaker recognition system applied on the NIST 2010 SRE task. The novel aspects of our systems are: 1) Improved performance on trials involving different vocal effort via the use of linear-scale features; 2) Expected improved recognition performance in the presence of reverberation and noise via the use of frequency domain perceptual linear predictor and cortical features; 3) A new discriminative kernel partial least squares (KPLS) framework that complements state-of-the-art back-end systems JFA and PLDA to aid in better overall recognition; and 4) Acceleration of JFA, PLDA and KPLS back-ends via distributed computing. The individual components of the system and the fused system are compared against a baseline JFA system and results reported by SRI and MIT-LL on SRE2010. Daniel Garcia-Romero, Xinhui Zhou, Dmitry N. Zotkin, Balaji Vasan Srinivasan, Yuancheng Luo, Sriram Ganapathy, Samuel Thomas 0001, Sridhar Krishna Nemala, Garimella S. V. S. Sivaram, Majid Mirbagheri, Sri Harish Reddy Mallidi, Thomas Janu, Padmanabhan Rajan, Nima Mesgarani, Mounya Elhilali, Hynek Hermansky, Shihab A. Shamma, Ramani Duraiswami |
ICASSP | 3 |
| 2011 | A partial least squares framework for speaker recognitionabstractModern approaches to speaker recognition (verification) operate in a space of "supervectors" created via concatenation of the mean vectors of a Gaussian mixture model (GMM) adapted from a universal background model (UBM). In this space, a number of approaches to model inter-class separability and nuisance attribute variability have been proposed. We develop a method for modeling the variability associated with each class (speaker) by using partial-least-squares - a latent variable modeling technique, which isolates the most informative subspace for each speaker. The method is tested on NIST SRE 2008 data and provides promising results. The method is shown to be noise-robust and to be able to efficiently learn the subspace corresponding to a speaker on training data consisting of multiple utterances. Balaji Vasan Srinivasan, Dmitry N. Zotkin, Ramani Duraiswami |
ICASSP | 2 |
| 2011 | Kernel Partial Least Squares for Speaker RecognitionabstractI-vectors are a concise representation of speaker characteristics. Recent advances in speaker recognition have utilized their ability to capture speaker and channel variability to develop efficient recognition engines. Inter-speaker relationships in the i-vector space are non-linear. Accomplishing effective speaker recognition requires a good modeling of these non-linearities and can be cast as a machine learning problem. In this paper, we propose a kernel partial least squares (kernel PLS, or KPLS) framework for modeling speakers in the i-vectors space. The resulting recognition system is tested across several conditions of the NIST SRE 2010 extended core data set and compared against state-of-the-art systems: Joint Factor Analysis (JFA), Balaji Vasan Srinivasan, Daniel Garcia-Romero, Dmitry N. Zotkin, Ramani Duraiswami |
INTERSPEECH | 3 |
| 2010 | Automatic matched filter recovery via the audio cameraabstractThe sound reaching the acoustic sensor in a realistic environment contains not only the part arriving directly from the sound source but also a number of environmental reflections. The effect of those on the sound is equivalent to a convolution with the room impulse response and can be undone via deconvolution - a technique known as matched filter processing. However, the filter is usually pre-computed in advance using known room geometry and source/receiver positions, and any deviations from those cause the performance to degrade significantly. In this work, an algorithm is proposed to compute the matched filter automatically using an audio camera - a microphone array based system that provides real-time audio images (essentially plots of steered response power in various directions) of environment. Acoustic sources, as well as their significant reflections, are revealed as peaks in the audio image. The reflections are associated with sound source(s) using an acoustic similarity metric, and an approximate matched filter is computed to align the reflections in time with the direct arrival. Preliminary experimental evaluation of the method is performed. It is shown that in case of two sources the reflections are identified correctly, the time delays recovered agree well with those computed from geometric constraints, and that the output SNR improves when the reflections are added coherently to the signal obtained by beamforming directly at the source. Adam O'Donovan, Ramani Duraiswami, Dmitry N. Zotkin |
ICASSP | 3 |
| 2010 | Kernelized Rényi distance for speaker recognitionabstractSpeaker recognition systems classify a test signal as a speaker or an imposter by evaluating a matching score between input and reference signals. We propose a new information theoretic approach for computation of the matching score using the Rényi entropy. The proposed entropic distance, the Kernelized Rényi distance (KRD), is formulated in a non-parametric way and the resulting measure is efficiently evaluated in a parallelized fashion on a graphical processor. The distance is then adapted as a scoring function and its performance compared with other popular scoring approaches in a speaker identification and speaker verification framework. Balaji Vasan Srinivasan, Ramani Duraiswami, Dmitry N. Zotkin |
ICASSP | 3 |
| 2010 | Plane-Wave Decomposition of Acoustical Scenes Via Spherical and Cylindrical Microphone ArraysabstractSpherical and cylindrical microphone arrays offer a number of attractive properties such as direction-independent acoustic behavior and ability to reconstruct the sound field in the vicinity of the array. Beamforming and scene analysis for such arrays is typically done using sound field representation in terms of orthogonal basis functions (spherical/cylindrical harmonics). In this paper, an alternative sound field representation in terms of plane waves is described, and a method for estimating it directly from measurements at microphones is proposed. It is shown that representing a field as a collection of plane waves arriving from various directions simplifies source localization, beamforming, and spatial audio playback. A comparison of the new method with the well-known spherical harmonics based beamforming algorithm is done, and it is shown that both algorithms can be expressed in the same framework but with weights computed differently. It is also shown that the proposed method can be extended to cylindrical arrays. A number of features important for the design and operation of spherical microphone arrays in real applications are revealed. Results indicate that it is possible to reconstruct the sound scene up to order p with p2 microphones spherical array. Dmitry N. Zotkin, Ramani Duraiswami, Nail A. Gumerov |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Plane-wave decomposition of a sound scene using a cylindrical microphone arrayabstractThe analysis for microphone arrays formed by mounting microphones on a sound-hard spherical or cylindrical baffle is typically performed using a decomposition of the sound field in terms of orthogonal basis functions. An alternative representation in terms of plane waves and a method for obtaining the coefficients of such a representation directly from measurements was proposed recently for the case of a spherical array. It was shown that representing the field as a collection of plane waves arriving from various directions simplifies both source localization and beamforming. In this paper, these results are extended to the case of the cylindrical array. Similarly to the spherical array case, localization and beamforming based on plane-wave decomposition perform as well as the traditional orthogonal function based methods while being numerically more stable. Both simulated and experimental results are presented. Dmitry N. Zotkin, Ramani Duraiswami |
ICASSP | 1 |
| 2008 | Imaging concert hall acoustics using visual and audio camerasabstractUsing a developed real time audio camera, that uses the output of a spherical microphone array beamformer steered in all directions to create central projection to create acoustic intensity images, we present a technique to measure the acoustics of rooms and halls. A panoramic mosaiced visual image of the space is also create. Since both the visual and the audio camera images are central projection, registration of the acquired audio and video images can be performed using standard computer vision techniques. We describe the technique, and apply it to the examine the relation between acoustical features and architectural details of the Dekelbaum concert hall at the Clarice Smith Performing Arts Center in College Park, MD. Adam O'Donovan, Ramani Duraiswami, Dmitry N. Zotkin |
ICASSP | 3 |
| 2008 | Sound field decomposition using spherical microphone arraysabstractSpherical microphone arrays offer a number of attractive properties such as direction-independent acoustic behavior and ability to reconstruct the sound field in the vicinity of the array. Such ability is necessary in applications such as ambisonics and recreating auditory environment over headphones. We compare the performance of two scene reconstruction algorithms - one based on least-squares fitting the observed potentials and another based on computing the far-field signature function directly from the microphone measurements. A number of features important for the design and operation of spherical microphone arrays in real applications are revealed. Results indicate that it is possible to reconstruct the sound scene up to order p with p2microphones. Dmitry N. Zotkin, Ramani Duraiswami, Nail A. Gumerov |
ICASSP | 1 |
| 2007 | Multimodal Tracking for Smart Videoconferencing and Video SurveillanceabstractMany applications require the ability to track the 3-D motion of the subjects. We build a particle filter based framework for multimodal tracking using multiple cameras and multiple microphone arrays. In order to calibrate the resulting system, we propose a method to determine the locations of all microphones using at least five loudspeakers and under assumption that for each loudspeaker there exists a microphone very close to it. We derive the maximum likelihood (ML) estimator, which reduces to the solution of the non-linear least squares problem. We verify the correctness and robustness of the multimodal tracker and of the self-calibration algorithm both with Monte-Carlo simulations and on real data from three experimental setups. Dmitry N. Zotkin, Vikas C. Raykar, Ramani Duraiswami, Larry Davis 0001 |
CVPR | 1 |
| 2007 | Fast Multipole Accelerated Boundary Elements for Numerical Computation of the Head Related Transfer FunctionabstractThe numerical computation of head related transfer functions has been attempted by a number of researchers. However, the cost of the computations has meant that usually only low frequencies can be computed and further the computations take inordinately long times. Because of this, comparisons of the computations with measurements are also difficult. We present a fast multipole based iterative preconditioned Krylov solution of a boundary element formulation of the problem and use a new formulation that enables the reciprocity technique to be accurately employed. This allows the calculation to proceed for higher frequencies and larger discretizations. Preliminary results of the computations and of comparisons with measured HRTFs are presented. Nail A. Gumerov, Ramani Duraiswami, Dmitry N. Zotkin |
ICASSP (1) | 3 |
| 2007 | Efficient Conversion of X.Y Surround Sound Content to Binaural Head-Tracked Form for HRTF-Enabled PlaybackabstractBinaural presentation of X.Y sound is usually performed using virtual audio principles - that is, by attempting to virtually reproduce the setup of the X+Y loudspeakers in the reference room configuration. The computational cost of such playback is linear in the number of channels in the X.Y setup. We present a novel scheme that computes, offline, a spatio-temporal representation of the sound field in the listening area and store it as a multipole expansion. During head-tracked playback, the binaural signal is obtained by evaluating the multipole expansion at the ear position corresponding to the current user pose, resulting in a fixed playback cost. The representation is further extended to incorporate individualized HRTFs at no additional cost. Simulation results are presented. Dmitry N. Zotkin, Ramani Duraiswami, Nail A. Gumerov |
ICASSP (1) | 1 |
| 2007 | Fast Evaluation of the Room Transfer Function Using Multipole ExpansionabstractReverberation in rooms is often simulated with the image method due to Allen and Berkley (1979). This method has an asymptotic complexity that is cubic in terms of the simulated reverberation length. When employed in the frequency domain, it is relatively computationally expensive if there are many receivers in the room or if the source or receiver positions are changing with time. The computational complexity of the image method is due to the repeated summation of the fields generated by a large number of image sources. In this paper, a fast method to perform such summations is presented. The method is based on multipole expansion of the monopole source potential. For offline computation of the room transfer function for N image sources and M receiver points, use of the Allen-Berkley algorithm requires O(NM) operations, whereas use of the proposed method requires only O(N+M) operations, resulting in significantly faster computation of reverberant sound fields. The proposed method also has a considerable speed advantage in situations where the room transfer function must be rapidly updated online in response to source/receiver location changes. Simulation results are presented, and algorithm accuracy, speed, and implementation details are discussed. For problems that require frequency-domain computations, the algorithm is found to generate sound fields identical to the ones obtained with the frequency-domain version of the Allen-Berkley algorithm at a fraction of computational cost Ramani Duraiswami, Dmitry N. Zotkin, Nail A. Gumerov |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Frequency Independent Flexible Spherical Beamforming Via Rbf FittingabstractWe describe a new method for sound analysis using a spherical microphone array without the use of quadrature over the sphere. Quadrature based solutions are very sensitive to the placement of microphones on the sphere, needing measurements to be made at exactly the quadrature positions. We propose to use fitting with band-limited radial basis functions (RBFs) rather than quadrature. Our approach results in frequency independent beamformer weights for flexibly placed microphone locations. Results are demonstrated using both synthetic and real spherical array data. Arkady Yerukhimovich, Ramani Duraiswami, Nail A. Gumerov, Dmitry N. Zotkin |
ICASSP (5) | 4 |
| 2005 | Processing of reverberant speech for time-delay estimationabstractIn this paper, we present a method of extracting the time-delay between speech signals collected at two microphone locations. Time-delay estimation from microphone outputs is the first step for many sound localization algorithms, and also for enhancement of speech. For time-delay estimation, speech signals are normally processed using short-time spectral information (either magnitude or phase or both). The spectral features are affected by degradations in speech caused by noise and reverberation. Features corresponding to the excitation source of the speech production mechanism are robust to such degradations. We show that these source features can be extracted reliably from the speech signal. The time-delay estimate can be obtained using the features extracted even from short segments (50-100 ms) of speech from a pair of microphones. The proposed method for time-delay estimation is found to perform better than the generalized cross-correlation (GCC) approach. A method for enhancement of speech is also proposed using the knowledge of the time-delay and the information of the excitation source. Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Ramani Duraiswami, Dmitry N. Zotkin |
IEEE Trans. Speech Audio Process. | 4 |
| 2004 | Interpolation and range extrapolation of HRTFs [head related transfer functions]abstractThe head related transfer function (HRTF) characterizes the scattering properties of a person's anatomy (especially the pinnae, head and torso), and exhibits considerable person-to-person variability. It is usually measured as a part of a tedious experiment, and this leads to the function being sampled at a few angular locations. When the HRTF is needed at intermediate angles, its value must be interpolated. Further, its range dependence is also neglected, which is invalid for nearby sources. Since the HRTF arises from a scattering process, it can be characterized as a solution of a scattering problem. In this paper, we show that by taking this viewpoint and performing some analysis we can express the HRTF in terms of a series of multipole solutions of the Helmholtz equation. This approach leads to a natural solution to the problem of HRTF interpolation. Furthermore, we show that the range-dependence of the HRTF in the near-field can also be obtained by extrapolation from measurements at one range. Ramani Duraiswami, Dmitry N. Zotkin, Nail A. Gumerov |
ICASSP (4) | 2 |
| 2004 | Accelerated speech source localization via a hierarchical search of steered response powerabstractAccurate and fast localization of multiple speech sound sources is a problem that is of significant interest in applications such as conferencing systems. Recently, approaches that are based on search for local peaks of the steered response power are becoming popular, despite their known computational expense. Based on the observation that the wavelengths of the sound from a speech source are comparable to the dimensions of the space being searched and that the source is broadband, we have developed an efficient search algorithm. Significant speedups are achieved by using coarse-to-fine strategies in both space and frequency. We present applications of the search algorithm to speed up simple delay-and-sum beamforming and steered response power phase-transform weighted (SRP-PHAT) source localization algorithms. A systematic series of comparisons with previous algorithms are made that show that the technique is much faster, robust, and accurate. The performance of the algorithm can be further improved by using constraints from computer vision. Dmitry N. Zotkin, Ramani Duraiswami |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Rendering localized spatial audio in a virtual auditory spaceabstractHigh-quality virtual audio scene rendering is required for emerging virtual and augmented reality applications, perceptual user interfaces, and sonification of data. We describe algorithms for creation of virtual auditory spaces by rendering cues that arise from anatomical scattering, environmental scattering, and dynamical effects. We use a novel way of personalizing the head related transfer functions (HRTFs) from a database, based on anatomical measurements. Details of algorithms for HRTF interpolation, room impulse response creation, HRTF selection from a database, and audio scene presentation are presented. Our system runs in real time on an office PC without specialized DSP hardware. Dmitry N. Zotkin, Ramani Duraiswami, Larry Davis 0001 |
IEEE Trans. Multim. | 1 |
| 2003 | Pitch and timbre manipulations using cortical representation of soundabstractThe sound received at the ears is processed by humans using signal processing that separates the signal along intensity, pitch and timbre dimensions. Conventional Fourier-based signal processing, while endowed with fast algorithms, is unable to represent a signal easily along the lines of these attributes. We use a recently proposed cortical representation (Elhilali, M. et al., Speech Communications, 2002) to represent and manipulate sound. We briefly overview algorithms for obtaining, manipulating and inverting cortical representation of a sound and describe algorithms for manipulating signal pitch and timbre separately. The algorithms are first used to create the sound of an instrument between a "guitar" and a "trumpet". Applications to creating maximally separable sounds in auditory user interfaces are discussed. Dmitry N. Zotkin, Shihab A. Shamma, Powen Ru, Ramani Duraiswami, Larry Davis 0001 |
ICASSP (5) | 1 |
| 2003 | Using computer vision to generate customized spatial audioabstractCreating high quality virtual spatial audio over headphones requires real-time head tracking, personalized head-related transfer functions (HRTFs) and customized room response models. While there are expensive solutions to address these issues based on costly head trackers, measured personalized HRTFs and room responses, these are not suitable for widespread or easy deployment and use. We report on the development of a system that uses computer vision to produce customizable models for both the HRTF and the room response, and to achieve head-tracking. The system uses relatively inexpensive cameras and widely available personal computers. Computer-vision based anthropometric measurements of the head, torso, and the external ears are used for HRTF customization. For low-frequency HRTF customization we employ a simple head-and-torso model developed recently [V. R. Algazi et al., 2002]. For high frequency customization we employ measured pinna characteristics as an index into a database of HRTFs [D. N. Zotkin et al., 2002]. For head tracking we employ an online implementation of the POSIT algorithm [D. DeMenthon and L. Davis, 1995] along with active markers to compute head pose in real-time. The system provides an enhanced virtual listening experience at low cost. Ankur Mohan, Ramani Duraiswami, Dmitry N. Zotkin, Daniel DeMenthon, Larry Davis 0001 |
ICME | 3 |
| 2003 | Pitch and timbre manipulations using cortical representation of soundabstractThe sound receiver at the ears is processed by humans using signal processing that separate the signal along intensity, pitch and timbre dimensions. Conventional Fourier-based signal processing, while endowed with fast algorithms, is unable to easily represent signal along these attributes. In this paper we use a cortical representation to represent the manipulate sound. We briefly overview algorithms for obtaining, manipulating and inverting cortical representation of sound and describe algorithms for manipulating signal pitch and timbre separately. The algorithms are first used to create sound of an instrument between a guitar and a trumpet. Applications to creating maximally separable sounds in auditory user interfaces are discussed. Dmitry N. Zotkin, Shihab A. Shamma, Powen Ru, Ramani Duraiswami, Larry Davis 0001 |
ICME | 1 |
| 2002 | Creation of virtual auditory spacesabstractHigh-quality virtual audio scene rendering is a must for emerging virtual/augmented reality applications and for perceptual user interfaces. We describe algorithms for creation of virtual auditory spaces using measured and non-individualized HRTFs and head tracking. Details of algorithms for HRTF interpolation, room impulse response creation, and audio scene presentation are presented. Tests show that individuals externalize well, and find our interface natural. The system runs in real time with latency of less than 30 ms on an office PC without specialized DSP. Dmitry N. Zotkin, Ramani Duraiswami, Larry Davis 0001 |
ICASSP | 1 |
| 2001 | Active speech source localization by a dual coarse-to-fine searchabstractAccurate and fast localization of multiple speech sound sources is a significant problem in videoconferencing systems. Based on the observation that the wavelengths of the sound from a speech source are comparable to the dimensions of the space being searched, and that the source is broadband, we develop an efficient search strategy that finds the source(s) in a given space. The search is made efficient by using coarse-to-fine strategies in both space and frequency. The algorithm is shown to be robust compared to typical delay-based estimators and fast enough for real-time implementation. Its performance can be further improved by using constraints from computer vision. Ramani Duraiswami, Dmitry N. Zotkin, Larry Davis 0001 |
ICASSP | 2 |
| 2001 | Multimodal localization of a flying batabstractWe present a new multimodal system that combines stereoscopic and audio-based source localization to track the position of a flying bat. Also presented are novel algorithms for audio source localization. The bat was allowed to fly in an anechoic room and monitored by two high-speed video cameras. The vocalizations of the bat were simultaneously recorded from six microphones. The data was then processed offline to localize the source and reconstruct the trajectory of the bat. We compare the performance of the localization algorithm with the position data obtained from stereoscopic pictures of the bat. The results confirm that the stereoscopic analysis and the audio localization are in good agreement. This system opens up new possibilities for performing multimodal research, and developing more tightly integrated algorithms. Kaushik Ghose, Dmitry N. Zotkin, Ramani Duraiswami, Cynthia F. Moss |
ICASSP | 2 |
| 2001 | Attentive Toys
Ismail Haritaoglu, Alex Cozzi, David Koons, Myron Flickner, Dmitry N. Zotkin, Ramani Duraiswami, Yaser Yacoob |
ICME | 5 |
| 2001 | Multimodal Tracking For Smart VideoconferencingabstractMany applications require the ability to track the 3-D motion of the subjects. We build a particle filter based framework for multimodal tracking using multiple cameras and multiple microphone arrays. In order to calibrate the resulting system, we propose a method to determine the locations of all microphones using at least five loudspeakers and under assumption that for each loudspeaker there exists a microphone very close to it. We derive the maximum likelihood (ML) estimator, which reduces to the solution of the non-linear least squares problem. We verify the correctness and robustness of the multimodal tracker and of the self-calibration algorithm both with Monte-Carlo simulations and on real data from three experimental setups. 1. Dmitry N. Zotkin, Ramani Duraiswami, Harsh Nanda, Larry Davis 0001 |
ICME | 1 |
| 2000 | An audio-video front-end for multimedia applicationsabstractApplications such as video gaming, virtual reality, multimodal user interfaces and videoconferencing, require systems that can locate and track persons in a room through a combination of visual and audio cues, enhance the sound that they produce, and perform identification. We describe the development of a particular multimodal sensor fusion system that is portable, runs in real time and achieves these objectives. The system employs novel algorithms for acoustical source location, video-based person tracking and overall system control, which are also described. Dmitry N. Zotkin, Ramani Duraiswami, Larry Davis 0001, Ismail Haritaoglu |
SMC | 1 |
| 1999 | Job-Length Estimation and Performance in Backfilling SchedulersabstractBackfilling is a simple and effective way of improving the utilization of space-sharing schedulers. Simple first-come-first-served approaches are ineffective because large jobs can fragment the available resources. Backfilling schedulers address this problem by allowing jobs to move ahead in the queue, provided that they will not delay subsequent jobs. Previous research has shown that inaccurate estimates of execution times can lead to better backfilling schedules. We characterize this effect on several workloads, and show that average slowdowns can be effectively reduced by systematically lengthening estimated execution times. Further, we show that the average job slowdown metric can be addressed directly by sorting jobs by increasing execution time. Finally, we modify our sorting scheduler to ensure that incoming jobs can be given hard guarantees. The resulting scheduler guarantees to avoid starvation, and performs significantly better than previous backfilling schedulers. Dmitry N. Zotkin, Peter J. Keleher |
HPDC | 1 |
| 1998 | Pictorial query trees for query specification in image databasesabstractA technique that enables specifying complex queries in image databases using pictorial query trees is presented. The leaves of a pictorial query tree correspond to individual pictorial queries that specify which objects should appear in the target images as well as how many occurrences of each object are required. In addition, the minimum required certainty of matching between query-image objects and database-image objects, as well as spatial constraints that specify bounds on the distance between objects and the relative direction between them are also specified. Internal nodes in the query tree represent logical operations (AND, OR, XOR) and their negations on the set of pictorial queries (or subtrees) represented by its children. The syntax of query trees is described. Algorithms for processing individual pictorial queries and for parsing and computing the overall result of a pictorial query tree are outlined. Aya Soffer, Hanan Samet, Dmitry N. Zotkin |
ICPR | 3 |