VLDB 2026 Research / reviewers in the wild / expert
Demba Ba 0001
dblp:96/4772 · also Demba E. Ba
· DBLP profile ↗
27ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0002-1139-1030ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision ModelsabstractSparse Autoencoders (SAEs) have emerged as a powerful framework for machine learning interpretability, enabling the unsupervised decomposition of model representations into a dictionary of abstract, human-interpretable concepts. However, we reveal a fundamental limitation: SAEs exhibit severe instability, as identical models trained on similar datasets can produce sharply different dictionaries, undermining their reliability as an interpretability tool. To address this issue, we draw inspiration from the Archetypal Analysis framework introduced by Cutler & Breiman (1994) and present Archetypal SAEs (A-SAE), wherein dictionary atoms are constrained to the data’s convex hull. This geometric anchoring significantly enhances the stability and plausibility of inferred dictionaries, and their mildly relaxed variants RA-SAEs further match state-of-the-art reconstruction abilities. To rigorously assess dictionary quality learned by SAEs, we introduce two new benchmarks that test (i) plausibility, if dictionaries recover “true” classification directions and (ii) identifiability, if dictionaries disentangle synthetic concept mixtures. Across all evaluations, RA-SAEs consistently yield more structured representations while uncovering novel, semantically meaningful concepts in large-scale vision models. Thomas Fel, Ekdeep Singh Lubana, Jacob S. Prince, Matthew Kowal, Victor Boutin, Isabel Papadimitriou, Binxu Wang, Martin Wattenberg, Demba Ba 0001, Talia Konkle |
ICML | 9 |
| 2025 | From Flat to Hierarchical: Extracting Sparse Representations with Matching PursuitabstractMotivated by the hypothesis that neural network representations encode abstract, interpretable features as linearly accessible, approximately orthogonal directions, sparse autoencoders (SAEs) have become a popular tool in interpretability literature. However, recent work has demonstrated phenomenology of model representations that lies outside the scope of this hypothesis, showing signatures of hierarchical, nonlinear, and multi-dimensional features. This raises the question: do SAEs represent features that possess structure at odds with their motivating hypothesis? If not, does avoiding this mismatch help identify said features and gain further insights into neural network representations? To answer these questions, we take a construction-based approach and re-contextualize the popular matching pursuit (MP) algorithm from sparse coding to design MP-SAE—an SAE that unrolls its encoder into a sequence of residual-guided steps, allowing it to capture hierarchical and nonlinearly accessible features. Comparing this architecture with existing SAEs on a mixture of synthetic and natural data settings, we show: (i) hierarchical concepts induce conditionally orthogonal features, which existing SAEs are unable to faithfully capture, and (ii) the nonlinear encoding step of MP-SAE recovers highly meaningful features, helping us unravel shared structure in the seemingly dichotomous representation spaces of different modalities in a vision-language model, hence demonstrating the assumption that useful features are solely linearly accessible is insufficient. We also show that the sequential encoder principle of MP-SAE affords an additional benefit of adaptive sparsity at inference time, which may be of independent interest. Overall, we argue our results provide credence to the idea that interpretability should begin with the phenomenology of representations, with methods emerging from assumptions that fit it. Valérie Costa, Thomas Fel, Ekdeep Singh Lubana, Bahareh Tolooshams, Demba Ba 0001 |
NeurIPS | 5 |
| 2025 | Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept GeometryabstractSparse Autoencoders (SAEs) are widely used to interpret neural networks by identifying meaningful concepts from their representations. However, do SAEs truly uncover all concepts a model relies on, or are they inherently biased toward certain kinds of concepts? We introduce a unified framework that recasts SAEs as solutions to a bilevel optimization problem, revealing a fundamental challenge: each SAE imposes structural assumptions about how concepts are encoded in model representations, which in turn shapes what it can and cannot detect. This means different SAEs are not interchangeable -- switching architectures can expose entirely new concepts or obscure existing ones. To systematically probe this effect, we evaluate SAEs across a spectrum of settings: from controlled toy models that isolate key variables, to semi-synthetic experiments on real model activations and finally to large-scale, naturalistic datasets. Across this progression, we examine two fundamental properties that real-world concepts often exhibit: heterogeneity in intrinsic dimensionality (some concepts are inherently low-dimensional, others are not) and nonlinear separability. We show that SAEs fail to recover concepts when these properties are ignored, and we design a new SAE that explicitly incorporates both, enabling the discovery of previously hidden concepts and reinforcing our theoretical insights. Our findings challenge the idea of a universal SAE and underscores the need for architecture-specific choices in model interpretability. Overall, we argue an SAE does not just reveal concepts -- it determines what can be seen at all. Sai Sumedh R. Hindupur, Ekdeep Singh Lubana, Thomas Fel, Demba Ba 0001 |
NeurIPS | 4 |
| 2024 | An Efficient Algorithm For Clustered Multi-Task Compressive SensingabstractThis paper considers clustered multi-task compressive sensing, a hierarchical model that solves multiple compressive sensing tasks by finding clusters of tasks that leverage shared information to mutually improve signal reconstruction. The existing inference algorithm for this model is computationally expensive and does not scale well in high dimensions. The main bottleneck involves repeated matrix inversion and log-determinant computation for multiple large covariance matrices. We propose a new algorithm that substantially accelerates model inference by avoiding the need to explicitly compute these covariance matrices. Our approach combines Monte Carlo sampling with iterative linear solvers. Our experiments reveal that compared to the existing baseline, our algorithm can be up to thousands of times faster and an order of magnitude more memory-efficient. Alexander Lin, Demba Ba 0001 |
ICASSP | 2 |
| 2023 | Learning Silhouettes with Group Sparse AutoencodersabstractSparse coding has been extensively used in neuroscience to model brain-like computation by drawing analogues between neurons’ firing activity and the nonzero elements of sparse vectors. Contemporary deep learning architectures have been used to model neural activity, inspired by signal processing algorithms; however sparse coding architectures are not able to explain the higher-order categorization that has been empirically observed at the neural level. In this work, we propose a novel model-based architecture, termed group-sparse autoencoder, that produces sparse activity patterns in line with neural modeling, but showcases a higher-level order in its activation maps. We evaluate a dense model of our architecture on MNIST and CIFAR-10 and show that it learns dictionaries that resemble silhouettes of the given class, while its activations have a significantly higher level order compared to sparse architectures. Source code is available at: https://github.com/manosth/silhouette-learning. Emmanouil Theodosis, Demba Ba 0001 |
ICASSP | 2 |
| 2023 | Probabilistic Unrolling: Scalable, Inverse-Free Maximum Likelihood Estimation for Latent Gaussian ModelsabstractLatent Gaussian models have a rich history in statistics and machine learning, with applications ranging from factor analysis to compressed sensing to time series analysis. The classical method for maximizing the likelihood of these models is the expectation-maximization (EM) algorithm. For problems with high-dimensional latent variables and large datasets, EM scales poorly because it needs to invert as many large covariance matrices as the number of data points. We introduce probabilistic unrolling, a method that combines Monte Carlo sampling with iterative linear solvers to circumvent matrix inversion. Our theoretical analyses reveal that unrolling and backpropagation through the iterations of the solver can accelerate gradient estimation for maximum likelihood estimation. In experiments on simulated and real data, we demonstrate that probabilistic unrolling learns latent Gaussian models up to an order of magnitude faster than gradient EM, with minimal losses in model performance. Alexander Lin, Bahareh Tolooshams, Yves F. Atchadé, Demba Ba 0001 |
ICML | 4 |
| 2022 | Mixture Model Auto-Encoders: Deep Clustering Through Dictionary LearningabstractState-of-the-art approaches for clustering high-dimensional data utilize deep auto-encoder architectures. Many of these networks require a large number of parameters and suffer from a lack of interpretability, due to the black-box nature of the auto-encoders. We introduce Mixture Model Auto-Encoders (MixMate), a novel architecture that clusters data by performing inference on a generative model. Built on ideas from sparse dictionary learning and mixture models, MixMate comprises several auto-encoders, each tasked with reconstructing data in a distinct cluster, while enforcing sparsity in the latent space. Through experiments on various image datasets, we show that MixMate achieves competitive performance versus state-of-the-art deep clustering algorithms, while using orders of magnitude fewer parameters. Alexander Lin, Andrew H. Song, Demba Ba 0001 |
ICASSP | 3 |
| 2022 | High-Dimensional Sparse Bayesian Learning without Covariance MatricesabstractSparse Bayesian learning (SBL) is a powerful framework for tackling the sparse coding problem. However, the most popular inference algorithms for SBL become too expensive for high-dimensional settings, due to the need to store and compute a large covariance matrix. We introduce a new inference scheme that avoids explicit construction of the covariance matrix by solving multiple linear systems in parallel to obtain the posterior moments for SBL. Our approach couples a little-known diagonal estimation result from numerical linear algebra with the conjugate gradient algorithm. On several simulations, our method scales better than existing approaches in computation time and memory, especially for structured dictionaries capable of fast matrix-vector multiplication. Alexander Lin, Andrew H. Song, Berkin Bilgic, Demba Ba 0001 |
ICASSP | 4 |
| 2022 | Gaussian Process Convolutional Dictionary LearningabstractConvolutional dictionary learning (CDL), the problem of estimating shift-invariant templates from data, is typically conducted in the absence of a prior/structure on the templates. In data-scarce or low signal-to-noise ratio (SNR) regimes, learned templates overfit the data and lack smoothness, which can affect the predictive performance of downstream tasks. To address this limitation, we propose GPCDL, a convolutional dictionary learning framework that enforces priors on templates using Gaussian Processes (GPs). With the focus on smoothness, we show theoretically that imposing a GP prior is equivalent to Wiener filtering the learned templates, thereby suppressing high-frequency components and promoting smoothness. We show that the algorithm is a simple extension of the classical iteratively reweighted least squares algorithm, independent of the choice of GP kernels. This property allows one to experiment flexibly with different smoothness assumptions. Through simulation, we show that GPCDL learns smooth dictionaries with better accuracy than the unregularized alternative across a range of SNRs. Through an application to neural spiking data, we show that GPCDL learns a more accurate and visually-interpretable smooth dictionary, leading to superior predictive performance compared to non-regularized CDL, as well as parametric alternatives. Andrew H. Song, Bahareh Tolooshams, Demba Ba 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Unfolding Neural Networks for Compressive Multichannel Blind DeconvolutionabstractWe propose a learned-structured unfolding neural network for the problem of compressive sparse multichannel blind-deconvolution. In this problem, each channel’s measurements are given as convolution of a common source signal and sparse filter. Unlike prior works where the compression is achieved either through random projections or by applying a fixed structured compression matrix, this paper proposes to learn the compression matrix from data. Given the full measurements, the proposed network is trained in an unsupervised fashion to learn the source and estimate sparse filters. Then, given the estimated source, we learn a structured compression operator while optimizing for signal reconstruction and sparse filter recovery. The efficient structure of the compression allows its practical hardware implementation. The proposed neural network is an autoencoder constructed based on an unfolding approach: upon training, the encoder maps the compressed measurements into an estimate of sparse filters using the compression operator and the source, and the linear convolutional decoder reconstructs the full measurements. We demonstrate that our method is superior to classical structured compressive sparse multichannel blind-deconvolution methods in terms of accuracy and speed of sparse filter recovery. Bahareh Tolooshams, Satish Mulleti, Demba Ba 0001, Yonina C. Eldar |
ICASSP | 3 |
| 2021 | PLSO: A generative framework for decomposing nonstationary time-series into piecewise stationary oscillatory componentsabstractTo capture the slowly time-varying spectral content of real-world time-series, a common paradigm is to partition the data into approximately stationary intervals and perform inference in the time-frequency domain. However, this approach lacks a corresponding nonstationary time-domain generative model for the entire data and thus, time-domain inference occurs in each interval separately. This results in distortion/discontinuity around interval boundaries and can consequently lead to erroneous inferences based on any quantities derived from the posterior, such as the phase. To address these shortcomings, we propose the Piecewise Locally Stationary Oscillation (PLSO) model for decomposing time-series data with slowly time-varying spectra into several oscillatory, piecewise-stationary processes. PLSO, as a nonstationary time-domain generative model, enables inference on the entire time-series without boundary effects and simultaneously provides a characterization of its time-varying spectral properties. We also propose a novel two-stage inference algorithm that combines Kalman theory and an accelerated proximal gradient algorithm. We demonstrate these points through experiments on simulated data and real neural data from the rat and the human brain. Andrew H. Song, Demba Ba 0001, Emery N. Brown |
UAI | 2 |
| 2021 | Deep Residual Autoencoders for Expectation Maximization-Inspired Dictionary LearningabstractWe introduce a neural-network architecture, termed the constrained recurrent sparse autoencoder (CRsAE), that solves convolutional dictionary learning problems, thus establishing a link between dictionary learning and neural networks. Specifically, we leverage the interpretation of the alternating-minimization algorithm for dictionary learning as an approximate expectation-maximization algorithm to develop autoencoders that enable the simultaneous training of the dictionary and regularization parameter (ReLU bias). The forward pass of the encoder approximates the sufficient statistics of the E-step as the solution to a sparse coding problem, using an iterative proximal gradient algorithm called FISTA. The encoder can be interpreted either as a recurrent neural network or as a deep residual network, with two-sided ReLU nonlinearities in both cases. The M-step is implemented via a two-stage backpropagation. The first stage relies on a linear decoder applied to the encoder and a norm-squared loss. It parallels the dictionary update step in dictionary learning. The second stage updates the regularization parameter by applying a loss function to the encoder that includes a prior on the parameter motivated by Bayesian statistics. We demonstrate in an image-denoising task that CRsAE learns Gabor-like filters and that the EM-inspired approach for learning biases is superior to the conventional approach. In an application to recordings of electrical activity from the brain, we demonstrate that CRsAE learns realistic spike templates and speeds up the process of identifying spike times by 900× compared with algorithms based on convex optimization. Bahareh Tolooshams, Sourav Dey, Demba Ba 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Convolutional dictionary learning based auto-encoders for natural exponential-family distributionsabstractWe introduce a class of auto-encoder neural networks tailored to data from the natural exponential family (e.g., count data). The architectures are inspired by the problem of learning the filters in a convolutional generative model with sparsity constraints, often referred to as convolutional dictionary learning (CDL). Our work is the first to combine ideas from convolutional generative models and deep learning for data that are naturally modeled with a non-Gaussian distribution (e.g., binomial and Poisson). This perspective provides us with a scalable and flexible framework that can be re-purposed for a wide range of tasks and assumptions on the generative model. Specifically, the iterative optimization procedure for solving CDL, an unsupervised task, is mapped to an unfolded and constrained neural network, with iterative adjustments to the inputs to account for the generative distribution. We also show that the framework can easily be extended for discriminative training, appropriate for a supervised task. We 1) demonstrate that fitting the generative model to learn, in an unsupervised fashion, the latent stimulus that underlies neural spiking data leads to better goodness-of-fit compared to other baselines, 2) show competitive performance compared to state-of-the-art algorithms for supervised Poisson image denoising, with significantly fewer parameters, and 3) characterize the gradient dynamics of the shallow binomial auto-encoder. Bahareh Tolooshams, Andrew H. Song, Simona Temereanca, Demba Ba 0001 |
ICML | 4 |
| 2019 | Clustering Time Series with Nonlinear Dynamics: A Bayesian Non-Parametric and Particle-Based ApproachabstractWe propose a general statistical framework for clustering multiple time series that exhibit nonlinear dynamics into an a-priori-unknown number of sub-groups. Our motivation comes from neuroscience, where an important problem is to identify, within a large assembly of neurons, subsets that respond similarly to a stimulus or contingency. Upon modeling the multiple time series as the output of a Dirichlet process mixture of nonlinear state-space models, we derive a Metropolis-within-Gibbs algorithm for full Bayesian inference that alternates between sampling cluster assignments and sampling parameter values that form the basis of the clustering. The Metropolis step employs recent innovations in particle-based methods. We apply the framework to clustering time series acquired from the prefrontal cortex of mice in an experiment designed to characterize the neural underpinnings of fear. Alexander Lin, Yingzhuo Zhang, Jeremy Heng, Stephen A. Allsop, Kay Tye, Pierre E. Jacob, Demba Ba 0001 |
AISTATS | 7 |
| 2018 | Wavelet Shrinkage and Thresholding Based Robust Classification for Brain-Computer InterfaceabstractA macaque monkey is trained to perform two different kinds of tasks, memory aided and visually aided. In each task, the monkey saccades to eight possible target locations. A classifier is proposed for direction decoding and task decoding based on local field potentials (LFP) collected from the prefrontal cortex. The LFP time-series data is modeled in a nonparametric regression framework, as a function corrupted by Gaussian noise. It is shown that if the function belongs to Besov bodies, then the proposed wavelet shrinkage and thresholding based classifier is robust and consistent. The classifier is then applied to the LFP data to achieve high decoding performance. The proposed classifier is also quite general and can be applied for the classification of other types of time-series data as well, not necessarily brain data. Taposh Banerjee, John S. Choi, Bijan Pesaran, Demba Ba 0001, Vahid Tarokh |
ICASSP | 4 |
| 2018 | Estimating a Separably Markov Random Field from Binary ObservationsabstractA fundamental problem in neuroscience is to characterize the dynamics of spiking from the neurons in a circuit that is involved in learning about a stimulus or a contingency. A key limitation of current methods to analyze neural spiking data is the need to collapse neural activity over time or trials, which may cause the loss of information pertinent to understanding the function of a neuron or circuit. We introduce a new method that can determine not only the trial-to-trial dynamics that accompany the learning of a contingency by a neuron, but also the latency of this learning with respect to the onset of a conditioned stimulus. The backbone of the method is a separable two-dimensional (2D) random field (RF) model of neural spike rasters, in which the joint conditional intensity function of a neuron over time and trials depends on two latent Markovian state sequences that evolve separately but in parallel. Classical tools to estimate state-space models cannot be applied readily to our 2D separable RF model. We develop efficient statistical and computational tools to estimate the parameters of the separable 2D RF model. We apply these to data collected from neurons in the prefrontal cortex in an experiment designed to characterize the neural underpinnings of the associative learning of fear in mice. Overall, the separable 2D RF model provides a detailed, interpretable characterization of the dynamics of neural spiking that accompany the learning of a contingency. Yingzhuo Zhang, Noa Malem-Shinitski, Stephen A. Allsop, Kay Tye, Demba Ba 0001 |
Neural Comput. | 5 |
| 2018 | A Multitaper Frequency-Domain Bootstrap MethodabstractSpectral properties of the electroencephalogram (EEG) are commonly analyzed to characterize the brain's oscillatory properties in basic science and clinical neuroscience studies. The spectrum is a function that describes power as a function of frequency. To date inference procedures for spectra have focused on constructing confidence intervals at single frequencies using large sample-based analytic procedures or jackknife techniques. These procedures perform well when the frequencies of interest are chosen before the analysis. When these frequencies are chosen after some of the data have been analyzed, the validity of these conditional inferences is not addressed. If power at more than one frequency is investigated, corrections for multiple comparisons must also be incorporated. To develop a statistical inference approach that considers the spectrum as a function defined across frequencies, we combine multitaper spectral methods with a frequency-domain bootstrap (FDB) procedure. The multitaper method is optimal for minimizing the bias-variance tradeoff in spectral estimation. The FDB makes it possible to conduct Monte Carlo based inferences for any part of the spectrum by drawing random samples that respect the dependence structure in the EEG time series. We show that our multitaper FDB procedure performs well in simulation studies and in analyses comparing EEG recordings of children from two different age groups receiving general anesthesia. Seong-Eun Kim, Demba Ba 0001, Emery N. Brown |
IEEE Signal Process. Lett. | 2 |
| 2014 | Likelihood Methods for Point Processes with RefractorinessabstractLikelihood-based encoding models founded on point processes have received significant attention in the literature because of their ability to reveal the information encoded by spiking neural populations. We propose an approximation to the likelihood of a point-process model of neurons that holds under assumptions about the continuous time process that are physiologically reasonable for neural spike trains: the presence of a refractory period, the predictability of the conditional intensity function, and its integrability. These are properties that apply to a large class of point processes arising in applications other than neuroscience. The proposed approach has several advantages over conventional ones. In particular, one can use standard fitting procedures for generalized linear models based on iteratively reweighted least squares while improving the accuracy of the approximation to the likelihood and reducing bias in the estimation of the parameters of the underlying continuous-time model. As a result, the proposed approach can use a larger bin size to achieve the same accuracy as conventional approaches would with a smaller bin size. This is particularly important when analyzing neural data with high mean and instantaneous firing rates. We demonstrate these claims on simulated and real neural spiking activity. By allowing a substantive increase in the required bin size, our algorithm has the potential to lower the barrier to the use of point-process methods in an increasing number of applications. Luca Citi, Demba Ba 0001, Emery N. Brown, Riccardo Barbieri |
Neural Comput. | 2 |
| 2012 | Exact and Stable Recovery of Sequences of Signals with Sparse Increments via Differential _1-MinimizationabstractWe consider the problem of recovering a sequence of vectors, $(x_k)_{k=0}^K$, for which the increments $x_k-x_{k-1}$ are $S_k$-sparse (with $S_k$ typically smaller than $S_1$), based on linear measurements $(y_k = A_k x_k + e_k)_{k=1}^K$, where $A_k$ and $e_k$ denote the measurement matrix and noise, respectively. Assuming each $A_k$ obeys the restricted isometry property (RIP) of a certain order---depending only on $S_k$---we show that in the absence of noise a convex program, which minimizes the weighted sum of the $\ell_1$-norm of successive differences subject to the linear measurement constraints, recovers the sequence $(x_k)_{k=1}^K$ \emph{exactly}. This is an interesting result because this convex program is equivalent to a standard compressive sensing problem with a highly-structured aggregate measurement matrix which does not satisfy the RIP requirements in the standard sense, and yet we can achieve exact recovery. In the presence of bounded noise, we propose a quadratically-constrained convex program for recovery and derive bounds on the reconstruction error of the sequence. We supplement our theoretical analysis with simulations and an application to real video data. These further support the validity of the proposed approach for acquisition and recovery of signals with time-varying sparsity. Demba Ba 0001, Behtash Babadi, Patrick L. Purdon, Emery N. Brown |
NIPS | 1 |
| 2012 | Geometrically Constrained Room Modeling With Compact Microphone ArraysabstractThe geometry of an acoustic environment can be an important information in many audio signal processing applications. To estimate such a geometry, previous work has relied on large microphone arrays, multiple test sources, moving sources or the assumption of a 2-D room. In this paper, we lift these requirements and present a novel method that uses a compact microphone array to estimate a 3-D room geometry, delivering effective estimates with low-cost hardware. Our approach first probes the environment with a known test signal emitted by a loudspeaker co-located with the array, from which the room impulse responses (RIRs) are estimated. It then uses an ℓ1-regularized least-squares minimization to fit synthetically generated reflections to the RIRs, producing a sparse set of reflections. By enforcing structural constraints derived from the image model, these are classified into first-, second-, and third-order reflections, thereby deriving the room geometry. Using this method, we detect walls using off-the-shelf teleconferencing hardware with a typical range resolution of about 1 cm. We present results using simulations and data from real environments. Flavio P. Ribeiro, Dinei A. F. Florêncio, Demba Ba 0001, Cha Zhang |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | L1 regularized room modeling with compact microphone arraysabstractAcoustic room modeling has several applications. Recent results using large microphone arrays show good performance, and are helpful in many applications. For example, when designing a better acoustic treatment for a concert hall, these large arrays can be used to help map the acoustic environment and aid in the design. However, in real-time applications - including de-reverberation, sound source localization, speech enhancement and 3D audio - it is desirable to model the room with existing small arrays and existing loudspeakers. In this paper we propose a novel room modeling algorithm, which uses a constrained room model and ℓ1-regularized least-squares to achieve good estimation of room geometry. We present experimental results on both real and synthetic data. Demba Ba 0001, Flavio P. Ribeiro, Cha Zhang, Dinei A. F. Florêncio |
ICASSP | 1 |
| 2010 | Turning enemies into friends: Using reflections to improve sound source localizationabstractSound Source Localization (SSL) based on microphone arrays has numerous applications, and has received significant research attention. Common to all published research is the observation that the accuracy of SSL degrades with reverberation. Indeed, early (strong) reflections can have amplitudes similar to the direct signal, and will often interfere with the estimation. In this paper, we show that reverberation is not the enemy, and can be used to improve estimation. More specifically, we are able to use early reflections to significantly improve range and elevation estimation. The process requires two steps: during setup, a loudspeaker integrated with the array emits a probing sound, which is used to obtain estimates of the ceiling height, as well as the locations of the walls. In a second step (e.g., during a meeting), the device incorporates this knowledge into a maximum likelihood SSL algorithm. Experimental results on both real and synthetic data show huge improvements in range estimation accuracy. Flavio P. Ribeiro, Demba Ba 0001, Cha Zhang, Dinei A. F. Florêncio |
ICME | 2 |
| 2010 | Using Reverberation to Improve Range and Elevation Discrimination for Small Array Sound Source LocalizationabstractSound source localization (SSL) is an essential task in many applications involving speech capture and enhancement. As such, speaker localization with microphone arrays has received significant research attention. Nevertheless, existing SSL algorithms for small arrays still have two significant limitations: lack of range resolution, and accuracy degradation with increasing reverberation. The latter is natural and expected, given that strong reflections can have amplitudes similar to that of the direct signal, but different directions of arrival. Therefore, correctly modeling the room and compensating for the reflections should reduce the degradation due to reverberation. In this paper, we show a stronger result. If modeled correctly, early reflections can be used to provide more information about the source location than would have been available in an anechoic scenario. The modeling not only compensates for the reverberation, but also significantly increases resolution for range and elevation. Thus, we show that under certain conditions and limitations, reverberation can be used to improve SSL performance. Prior attempts to compensate for reverberation tried to model the room impulse response (RIR). However, RIRs change quickly with speaker position, and are nearly impossible to track accurately. Instead, we build a 3-D model of the room, which we use to predict early reflections, which are then incorporated into the SSL estimation. Simulation results with real and synthetic data show that even a simplistic room model is sufficient to produce significant improvements in range and elevation estimation, tasks which would be very difficult when relying only on direct path signal components. Flavio P. Ribeiro, Cha Zhang, Dinei A. F. Florêncio, Demba Ba 0001 |
IEEE Trans. Speech Audio Process. | 4 |
| 2008 | Maximum Likelihood Sound Source Localization and Beamforming for Directional Microphone Arrays in Distributed MeetingsabstractIn distributed meeting applications, microphone arrays have been widely used to capture superior speech sound and perform speaker localization through sound source localization (SSL) and beamforming. This paper presents a unified maximum likelihood framework of these two techniques, and demonstrates how such a framework can be adapted to create efficient SSL and beamforming algorithms for reverberant rooms and unknown directional patterns of microphones. The proposed method is closely related to steered response power-based algorithms, which are known to work extremely well in real-world environments. We demonstrate the effectiveness of the proposed method on challenging synthetic and real-world datasets, including over six hours of recorded meetings. Cha Zhang, Dinei A. F. Florêncio, Demba Ba 0001, Zhengyou Zhang |
IEEE Trans. Multim. | 3 |
| 2007 | Enhanced MVDR Beamforming for Arrays of Directional MicrophonesabstractMicrophone arrays based on the minimum variance distortionless response (MVDR) beamformer are among the most popular for speech enhancement applications. The original MVDR is excessively sensitive to source location and microphone gains. Previous research has made MVDR practical by successfully increasing the robustness of MVDR to source location, and MVDR-based microphone arrays are already commercially available. Nevertheless, MVDR performance is still weak in cases where microphone gain variations are too large, e.g., for circular arrays of directional microphones. In this paper we propose an improved MVDR beamformer which takes into account the effect of sensors (e.g. microphones) with arbitrary, potentially directional responses. Specifically, we form estimates of the relative magnitude responses of the sensors based on the data received at the array and include those in the original formulation of the MVDR beamforming problem. Experimental results on real-world audio data show an average 2.4 dB improvement over conventional MVDR beamforming, which does not account for the magnitude responses of the sensors. Demba Ba 0001, Dinei A. F. Florêncio, Cha Zhang |
ICME | 1 |
| 2007 | Integer Polar Coordinates for CompressionabstractThis paper introduces a family of integer-to-integer approximations to the Cartesian-to-polar coordinate transformation and analyzes its application to lossy compression. A high-rate analysis is provided for an encoder that first uniformly scalar quantizes, then transforms to "integer polar coordinates," and finally separately entropy codes angle and radius. For sources separable in polar coordinates, the performance (at high rate) is shown to match that of entropy-constrained unconstrained polar quantization - where the angular quantization is allowed to depend on the radius. Thus, for sources separable in polar coordinates but not separable in rectangular coordinates - including certain Gaussian scale mixtures - the proposed system performs better than any transform code. Furthermore, unlike unconstrained polar quantization, integer polar coordinates are appropriate for lossless compression of integer-valued vectors. Combination of integer polar coordinates with integer-to-integer transform coding is also discussed. Demba Ba 0001, Vivek K. Goyal |
ISIT | 1 |
| 2006 | Nonlinear Transform Coding: Polar Coordinates RevisitedabstractSummary form only given. We designed a family of integer-to-integer (i2i) approximations to the Cartesian-to-polar transformation and analyzed its behavior for high-rate transform coding. Denoting (ordinary, continuous) polar coordinates by (r, 0), our precise high-rate analysis relates the performance to the differential entropies of r2and 0, which are often easy to evaluate. One may thus predict when there is an improvement over linear transform coding. The analysis matches our simulations for coding of Gaussian scale mixtures and other polar-separable sources. The advantage over the best linear transform coder can be large. Our hope is to extend the polar-coordinate results to a general theory for nonlinear transform coding based on i2i implementations of arbitrary nonlinear transformations Demba Ba 0001, Vivek K. Goyal |
DCC | 1 |