EDBT 2026 Demo / reviewers in the wild / expert
Ronen Talmon
dblp:54/7051
· DBLP profile ↗
53ranked-venue papers
9as first author
21since 2021 · last 2025
0000-0002-6838-1423ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Finsler Multi-Dimensional Scaling: Manifold Learning for Asymmetric Dimensionality Reduction and EmbeddingabstractDimensionality reduction is a fundamental task that aims to simplify complex data by reducing its feature dimensionality while preserving essential patterns, with core applications in data analysis and visualisation. To preserve the underlying data structure, multi-dimensional scaling (MDS) methods focus on preserving pairwise dissimilarities, such as distances. They optimise the embedding to have pairwise distances as close as possible to the data dissimilarities. However, the current standard is limited to embedding data in Riemannian manifolds. Motivated by the lack of asymmetry in the Riemannian metric of the embedding space, this paper extends the MDS problem to a natural asymmetric generalisation of Riemannian manifolds called Finsler manifolds. Inspired by Euclidean space, we define a canonical Finsler space for embedding asymmetric data. Due to its simplicity with respect to geodesics, data representation in this space is both intuitive and simple to analyse. We demonstrate that our generalisation benefits from the same theoretical convergence guarantees. We reveal the effectiveness of our Finsler embedding across various types of non-symmetric data, highlighting its value in applications such as data visualisation, dimensionality reduction, directed graph embedding, and link prediction. Thomas Dagès, Simon Weber 0002, Ya-Wei Eileen Lin, Ronen Talmon, Daniel Cremers, Michael Lindenbaum, Alfred M. Bruckstein, Ron Kimmel |
CVPR | 4 |
| 2025 | Hyperbolic Distance Based on EMD and Diffusion for Hyperspectral ImagingabstractIn this paper, we introduce EMD-Based Hyperbolic Diffusion Distance (EMD-HDD), a new method for constructing a meaningful distance metric for hierarchical data with latent hierarchical structure. Our method relies on hyperbolic geometry, diffusion geometry, and the Earth Mover’s Distance (EMD). Specifically, our method embeds data points into a product manifold of hyperbolic spaces, allowing us to recover the hidden hierarchical structure encoded by the mutual relationships between features. We demonstrate the effectiveness of EMD-HDD through experiments on five hyperspectral imaging datasets, showcasing its capability to capture and reveal the intrinsic hierarchical structures inherent in such data. Elad Lavi, Amir Bourvine, Ya-Wei Eileen Lin, Ronen Talmon |
ICASSP | 4 |
| 2025 | RTF Estimation Using Riemannian Geometry for Speech Enhancement in the Presence of InterferencesabstractWe address the problem of multichannel audio signal enhancement in reverberant environments with interfering sources. We propose an approach that leverages the Riemannian geometry of the spatial correlation matrices of the received signals to estimate the relative transfer function (RTF) of the desired source. Specifically, we compute the spatial correlation matrices in short-time segments, and subsequently, their Riemannian mean, which preserves shared spatial components while attenuating unshared ones. This enables an effective rejection of intermittent interference, leading to accurate RTF estimation. We experimentally show that when the proposed RTF estimation is incorporated into the Minimum Variance Distortionless Response (MVDR) beamformer, it enhances the desired signal, outperforming the MVDR beamformer that is based on standard (Euclidean) RTF estimation. These favorable experimental results are demonstrated in challenging acoustic environments including multiple strong interfering sources, noise, and reverberations. Or Ronai, Yuval Sitton, Amitay Bar, Ronen Talmon |
ICASSP | 4 |
| 2025 | Tree-Wasserstein Distance for High Dimensional Data with a Latent Feature HierarchyabstractFinding meaningful distances between high-dimensional data samples is an important scientific task. To this end, we propose a new tree-Wasserstein distance (TWD) for high-dimensional data with two key aspects. First, our TWD is specifically designed for data with a latent feature hierarchy, i.e., the features lie in a hierarchical space, in contrast to the usual focus on embedding samples in hyperbolic space. Second, while the conventional use of TWD is to speed up the computation of the Wasserstein distance, we use its inherent tree as a means to learn the latent feature hierarchy. The key idea of our method is to embed the features into a multi-scale hyperbolic space using diffusion geometry and then present a new tree decoding method by establishing analogies between the hyperbolic embedding and trees. We show that our TWD computed based on data observations provably recovers the TWD defined with the latent feature hierarchy and that its computation is efficient and scalable. We showcase the usefulness of the proposed TWD in applications to word-document and single-cell RNA-sequencing datasets, demonstrating its advantages over existing TWDs and methods based on pre-trained models. Ya-Wei Eileen Lin, Ronald R. Coifman, Gal Mishne, Ronen Talmon |
ICLR | 4 |
| 2025 | Supervised and Semi-Supervised Diffusion Maps with Label-Driven DiffusionabstractIn this paper, we introduce Supervised Diffusion Maps (SDM) and Semi-Supervised Diffusion Maps (SSDM), which transform the well-known unsupervised dimensionality reduction algorithm, Diffusion Maps, into supervised and semi-supervised learning tools. The proposed methods, SDM and SSDM, are based on our new approach that treats the labels as a second view of the data. This unique framework allows us to incorporate ideas from multi-view learning. Specifically, we propose constructing two affinity kernels corresponding to the data and the labels. We then propose a multiplicative interpolation scheme of the two kernels, whose purpose is twofold. First, our scheme extracts the common structure underlying the data and the labels by defining a diffusion process driven by the data and the labels. This label-driven diffusion produces an embedding that emphasizes the properties relevant to the label-related task. Second, the proposed interpolation scheme balances the influence of the two kernels. We show on multiple benchmark datasets that the embedding learned by SDM and SSDM is more effective in downstream regression and classification tasks than existing unsupervised, supervised, and semi-supervised nonlinear dimension reduction methods. Harel Mendelman, Ronen Talmon |
ICLR | 2 |
| 2025 | Spectral Graph Coarsening Using Inner Product Preservation and the Grassmann ManifoldabstractWe propose a novel functorial graph coarsening method that preserves inner products between node features, a property often overlooked by existing approaches focusing primarily on structural fidelity.
By treating node features as functions on the graph and preserving their inner products, our method retains both structural and feature relationships, facilitating substantial benefits for downstream tasks.
To formalize this, we introduce the Inner Product Error (IPE), which quantifies how the inner products between node features are preserved.
Leveraging the underlying geometry of the problem on the Grassmann manifold, we formulate an optimization objective that minimizes the IPE, also for unseen smooth functions. We show that minimizing the IPE improves standard coarsening metrics, and illustrate our method’s properties through visual examples that highlight its clustering ability. Empirical results on benchmarks for graph coarsening and node classification show that our approach outperforms existing state-of-the-art methods. Ronen Talmon |
NeurIPS | 2 |
| 2025 | Joint Hierarchical Representation Learning of Samples and Features via Informed Tree-Wasserstein DistanceabstractHigh-dimensional data often exhibit hierarchical structures in both modes: samples and features. Yet, most existing approaches for hierarchical representation learning consider only one mode at a time.
In this work, we propose an unsupervised method for jointly learning hierarchical representations of samples and features via Tree-Wasserstein Distance (TWD).
Our method alternates between the two data modes. It first constructs a tree for one mode, then computes a TWD for the other mode based on that tree, and finally uses the resulting TWD to build the second mode’s tree. By repeatedly alternating through these steps, the method gradually refines both trees and the corresponding TWDs, capturing meaningful hierarchical representations of the data.
We provide a theoretical analysis showing that our method converges.
We show that our method can be integrated into hyperbolic graph convolutional networks as a pre-processing technique, improving performance in link prediction and node classification tasks.
In addition, our method outperforms baselines in sparse approximation and unsupervised Wasserstein distance learning tasks on word-document and single-cell RNA-sequencing datasets. Ya-Wei Eileen Lin, Ronald R. Coifman, Gal Mishne, Ronen Talmon |
NeurIPS | 4 |
| 2025 | Learning Shared Representations from Unpaired DataabstractLearning shared representations is a primary area of multimodal representation learning. The current approaches to achieve a shared embedding space rely heavily on paired samples from each modality, which are significantly harder to obtain than unpaired ones. In this work, we demonstrate that shared representations can be learned almost exclusively from unpaired data. Our arguments are grounded in the spectral embeddings of the random walk matrices constructed independently from each unimodal representation. Empirical results in computer vision and natural language processing domains support its potential, revealing the effectiveness of unpaired data in capturing meaningful cross-modal relations, demonstrating high capabilities in retrieval tasks, generation, arithmetics, zero-shot, and cross-domain classification. This work, to the best of our knowledge, is the first to demonstrate these capabilities almost exclusively from unpaired samples, giving rise to a cross-modal embedding that could be viewed as universal, i.e., independent of the specific modalities of the data. Our project page: https://shaham-lab.github.io/SUE_page. Amitai Yacobi, Nir Ben-Ari, Ronen Talmon, Uri Shaham 0001 |
NeurIPS | 3 |
| 2024 | The Expected Loss of Preconditioned Langevin Dynamics Reveals the Hessian RankabstractLangevin dynamics (LD) is widely used for sampling from distributions and for optimization. In this work, we derive a closed-form expression for the expected loss of preconditioned LD near stationary points of the objective function. We use the fact that at the vicinity of such points, LD reduces to an Ornstein–Uhlenbeck process, which is amenable to convenient mathematical treatment. Our analysis reveals that when the preconditioning matrix satisfies a particular relation with respect to the noise covariance, LD's expected loss becomes proportional to the rank of the objective's Hessian. We illustrate the applicability of this result in the context of neural networks, where the Hessian rank has been shown to capture the complexity of the predictor function but is usually computationally hard to probe. Finally, we use our analysis to compare SGD-like and Adam-like preconditioners and identify the regimes under which each of them leads to a lower expected loss. Amitay Bar, Rotem Mulayoff, Tomer Michaeli, Ronen Talmon |
AAAI | 4 |
| 2024 | Hyperbolic Diffusion Procrustes Analysis for Intrinsic Representation of Hierarchical Data SetsabstractIn this paper, we present Hyperbolic Diffusion Procrustes Analysis (HDPA), a new method for informative representation of hierarchical datasets based on hyperbolic geometry, diffusion geometry, and Procrustes analysis. Our method jointly embeds multiple datasets in a product manifold of hyperbolic spaces, where the data's hidden common hierarchical structure is provably recovered. In addition, our method generates an intrinsic embedding that accommodates the joint representation of multiple datasets with different features, acquired by different equipment, at different sites, or under different environmental conditions. Experimental results demonstrate the efficacy of HDPA on three biomedical datasets comprising heterogeneous gene expression and mass cytometry data. Ya-Wei Eileen Lin, Yuval Kluger, Ronen Talmon |
ICASSP | 3 |
| 2024 | Direct Position Determination by Covariance-Fitting on the Riemannian Manifold of Hermitian Positive Definite MatricesabstractDirect Position Determination (DPD) is the state-of-the-art solution for emitter localization using multiple phased arrays. This paper shows that DPD can be recast as a covariance-fitting (CF) problem that minimizes the Euclidean distance between a sample covariance matrix ${\mathbf{\hat R}}$ and its location-dependent model R. By showing equivalence to existing DPD methods, this CF viewpoint highlights that the geometry of the Hermitian Positive Definite (HPD) covariance matrices R and ${\mathbf{\hat R}}$ is simply overlooked. Based on this critical observation, we propose a new CF approach for DPD that specifically exploits the Riemannian geometry of HPD matrices for measuring the distance between R and ${\mathbf{\hat R}}$. Experimental results showcase that the proposed Riemannian CF approach for DPD leads to a significant improvement in localization accuracy. Joseph S. Picard, Amitay Bar, Ronen Talmon |
ICASSP | 3 |
| 2024 | Equivariant Machine Learning on Graphs with Nonlinear Spectral FiltersabstractEquivariant machine learning is an approach for designing deep learning models that respect the symmetries of the problem, with the aim of reducing model complexity and improving generalization.
In this paper, we focus on an extension of shift equivariance, which is the basis of convolution networks on images, to general graphs. Unlike images, graphs do not have a natural notion of domain translation.
Therefore, we consider the graph functional shifts as the symmetry group: the unitary operators that commute with the graph shift operator.
Notably, such symmetries operate in the signal space rather than directly in the spatial space.
We remark that each linear filter layer of a standard spectral graph neural network (GNN) commutes with graph functional shifts, but the activation function breaks this symmetry. Instead, we propose nonlinear spectral filters (NLSFs) that are fully equivariant to graph functional shifts and show that they have universal approximation properties.
The proposed NLSFs are based on a new form of spectral domain that is transferable between graphs.
We demonstrate the superior performance of NLSFs over existing spectral GNNs in node and graph classification benchmarks. Ya-Wei Eileen Lin, Ronen Talmon, Ron Levie |
NeurIPS | 2 |
| 2024 | Graph signal interpolation and extrapolation over manifold of Gaussian mixture
Itay Zach, Tsvi G. Dvorkind, Ronen Talmon |
Signal Process. | 3 |
| 2023 | Few-Sample Feature Selection via Feature Manifold LearningabstractIn this paper, we present a new method for few-sample supervised feature selection (FS). Our method first learns the manifold of the feature space of each class using kernels capturing multi-feature associations. Then, based on Riemannian geometry, a composite kernel is computed, extracting the differences between the learned feature associations. Finally, a FS score based on spectral analysis is proposed. Considering multi-feature associations makes our method multivariate by design. This in turn allows for the extraction of the hidden manifold underlying the features and avoids overfitting, facilitating few-sample FS. We showcase the efficacy of our method on illustrative examples and several benchmarks, where our method demonstrates higher accuracy in selecting the informative features compared to competing methods. In addition, we show that our FS leads to improved classification and better generalization when applied to test data. Tal Shnitzer, Yuval Kluger, Ronen Talmon |
ICML | 4 |
| 2023 | Hyperbolic Diffusion Embedding and Distance for Hierarchical Representation LearningabstractFinding meaningful representations and distances of hierarchical data is important in many fields. This paper presents a new method for hierarchical data embedding and distance. Our method relies on combining diffusion geometry, a central approach to manifold learning, and hyperbolic geometry. Specifically, using diffusion geometry, we build multi-scale densities on the data, aimed to reveal their hierarchical structure, and then embed them into a product of hyperbolic spaces. We show theoretically that our embedding and distance recover the underlying hierarchical structure. In addition, we demonstrate the efficacy of the proposed method and its advantages compared to existing methods on graph embedding benchmarks and hierarchical datasets. Ya-Wei Eileen Lin, Ronald R. Coifman, Gal Mishne, Ronen Talmon |
ICML | 4 |
| 2022 | Discovery of Single Independent Latent VariableabstractLatent variable discovery is a central problem in data analysis with a broad range of applications in applied science.In this work, we consider data given as an invertible mixture of two statistically independent components, and assume that one of the components is observed while the other is hidden. Our goal is to recover the hidden component.For this purpose, we propose an autoencoder equipped with a discriminator.Unlike the standard nonlinear ICA problem, which was shown to be non-identifiable, in the special case of ICA we consider here, we show that our approach can recover the component of interest up to entropy-preserving transformation.We demonstrate the performance of the proposed approach in several tasks, including image synthesis, voice cloning, and fetal ECG extraction. Uri Shaham 0001, Jonathan Svirsky, Ori Katz, Ronen Talmon |
NeurIPS | 4 |
| 2022 | Deep Isometric MapsabstractIsometric feature mapping is an established time-honored algorithm in manifold learning and non-linear dimensionality reduction. Its prominence can be attributed to the output of a coherent global low-dimensional representation of data by preserving intrinsic distances. In order to enable an efficient and more applicable isometric feature mapping, a diverse set of sophisticated advancements have been proposed to the original algorithm to incorporate important factors like sparsity of computation, conformality, topological constraints and spectral geometry. However, a significant shortcoming of most approaches is the dependence on large-scale dense-spectral decompositions and the inability to generalize to points far away from the sampling of the manifold. In this paper, we explore an unsupervised deep learning approach for computing distance-preserving maps for non-linear dimensionality reduction. We demonstrate that our framework is general enough to incorporate all previous advancements and show a significantly improved local and non-local generalization of the isometric mapping. Our approach involves training with only a few landmark points and avoids the need for population of dense matrices as well as computing their spectral decomposition. Gautam Pai 0001, Alexander M. Bronstein, Ronen Talmon, Ron Kimmel |
Image Vis. Comput. | 3 |
| 2021 | Aligning Sets of Temporal Signals with Riemannian Geometry and Koopman OperatorabstractIn this paper, we consider the problem of aligning data sets of short temporal signals without any a-priori known correspondence. We present a method combining Koopman operator theory and the Riemannian geometry of symmetric positive-definite (SPD) matrices. First, by taking a Koopman operator theory standpoint, we build feature matrices of the signals using dynamic mode decomposition (DMD). Second, we align these features using parallel transport of SPD matrices, built from the DMD feature matrices. We showcase the performance of the proposed method on simulated observations of a mechanical system and on two real-world applications: sleep stage identification and pre-epileptic seizure prediction. Ohad Rahamim, Ronen Talmon |
ICASSP | 2 |
| 2021 | Hyperbolic Procrustes Analysis Using Riemannian GeometryabstractLabel-free alignment between datasets collected at different times, locations, or by different instruments is a fundamental scientific task. Hyperbolic spaces have recently provided a fruitful foundation for the development of informative representations of hierarchical data. Here, we take a purely geometric approach for label-free alignment of hierarchical datasets and introduce hyperbolic Procrustes analysis (HPA). HPA consists of new implementations of the three prototypical Procrustes analysis components: translation, scaling, and rotation, based on the Riemannian geometry of the Lorentz model of hyperbolic space. We analyze the proposed components, highlighting their useful properties for alignment. The efficacy of HPA, its theoretical properties, stability and computational efficiency are demonstrated in simulations. In addition, we showcase its performance on three batch correction tasks involving gene expression and mass cytometry data. Specifically, we demonstrate high-quality unsupervised batch effect removal from data acquired at different sites and with different technologies that outperforms recent methods for label-free alignment in hyperbolic spaces. Ya-Wei Eileen Lin, Yuval Kluger, Ronen Talmon |
NeurIPS | 3 |
| 2021 | Joint Geometric and Topological Analysis of Hierarchical Datasets
Lior Aloni, Omer Bobrowski, Ronen Talmon |
ECML/PKDD (3) | 3 |
| 2021 | Graph of graphs analysis for multiplexed data with application to imaging mass cytometryabstractImaging Mass Cytometry (IMC) combines laser ablation and mass spectrometry to quantitate metal-conjugated primary antibodies incubated in intact tumor tissue slides. This strategy allows spatially-resolved multiplexing of dozens of simultaneous protein targets with 1μm resolution. Each slide is a spatial assay consisting of high-dimensional multivariate observations (m-dimensional feature space) collected at different spatial positions and capturing data from a single biological sample or even representative spots from multiple samples when using tissue microarrays. Often, each of these spatial assays could be characterized by several regions of interest (ROIs). To extract meaningful information from the multi-dimensional observations recorded at different ROIs across different assays, we propose to analyze such datasets using a two-step graph-based approach. We first construct for each ROI a graph representing the interactions between the m covariates and compute an m dimensional vector characterizing the steady state distribution among features. We then use all these m-dimensional vectors to construct a graph between the ROIs from all assays. This second graph is subjected to a nonlinear dimension reduction analysis, retrieving the intrinsic geometric representation of the ROIs. Such a representation provides the foundation for efficient and accurate organization of the different ROIs that correlates with their phenotypes. Theoretically, we show that when the ROIs have a particular bi-modal distribution, the new representation gives rise to a better distinction between the two modalities compared to the maximum a posteriori (MAP) estimator. We applied our method to predict the sensitivity to PD-1 axis blockers treatment of lung cancer subjects based on IMC data, achieving 97.3% average accuracy on two IMC datasets. This serves as empirical evidence that the graph of graphs approach enables us to integrate multiple ROIs and the intra-relationships between the features at each ROI, giving rise to an informative representation that is strongly associated with the phenotypic state of the entire image. Ya-Wei Eileen Lin, Tal Shnitzer, Ronen Talmon, Franz Villarroel-Espindola, Shruti Desai, Kurt A. Schalper, Yuval Kluger |
PLoS Comput. Biol. | 3 |
| 2020 | Option Discovery in the Absence of Rewards with Manifold AnalysisabstractOptions have been shown to be an effective tool in reinforcement learning, facilitating improved exploration and learning. In this paper, we present an approach based on spectral graph theory and derive an algorithm that systematically discovers options without access to a specific reward or task assignment. As opposed to the common practice used in previous methods, our algorithm makes full use of the spectrum of the graph Laplacian. Incorporating modes associated with higher graph frequencies unravels domain subtleties, which are shown to be useful for option discovery. Using geometric and manifold-based analysis, we present a theoretical justification for the algorithm. In addition, we showcase its performance in several domains, demonstrating clear improvements compared to competing methods. Amitay Bar, Ronen Talmon, Ron Meir |
ICML | 2 |
| 2020 | Global and Local Simplex Representations for Multichannel Source SeparationabstractThe problem of blind audio source separation (BASS) in noisy and reverberant conditions is addressed by a novel approach, termed Global and LOcal Simplex Separation (GLOSS), which integrates full- and narrow-band simplex representations. We show that the eigenvectors of the correlation matrix between time frames in a certain frequency band form a simplex that organizes the frames according to the speaker activities in the corresponding band. We propose to build two simplex representations: one global based on a broad frequency band and one local based on a narrow band. In turn, the two representations are combined to determine the dominant speaker in each time-frequency (TF) bin. Using the identified dominating speakers, a spectral mask is computed and is utilized for extracting each of the speakers using spatial beamforming followed by spectral postfiltering. The performance of the proposed algorithm is demonstrated using real-life recordings in various noisy and reverberant conditions. Bracha Laufer-Goldshtein, Ronen Talmon, Sharon Gannot |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Localization of an Unknown Number of Speakers in Adverse Acoustic Conditions Using Reliability Information and DiarizationabstractThis paper investigates localization of an arbitrary number of simultaneously active speakers in an acoustic enclosure. We propose an algorithm capable of estimating the number of speakers, using reliability information to obtain robust estimation results in adverse acoustic scenarios and estimating individual probability distributions describing the position of each speaker using convex geometry tools. To this end, we start from an established algorithm for localization of acoustic sources based on the EM algorithm. There, the estimation of the number of sources as well as the handling of reverberation has not been addressed sufficiently. We show improvement in the localization of a higher number of sources and in the robustness in adverse conditions including interference from competing speakers, reverberation and noise. Andreas Brendel, Bracha Laufer-Goldshtein, Sharon Gannot, Ronen Talmon, Walter Kellermann |
ICASSP | 4 |
| 2019 | Domain Adaptation Using Riemannian Geometry of Spd MatricesabstractIn this paper, we propose a new unsupervised domain adaptation method based on the Riemannian geometry of Symmetric Positive-Definite (SPD) matrices. The proposed domain adaptation is based on parallel transport (PT) and moments alignment. We show that this method facilitates meaningful comparisons between data points from different domains, while preserving the inherent internal structure of each domain. Experimental results demonstrate the adaptation of high-dimensional noisy electrophysiological signals collected from different subjects. Gal Maman, Or Yair, Danny Eytan, Ronen Talmon |
ICASSP | 4 |
| 2019 | DIMAL: Deep Isometric Manifold Learning Using Sparse Geodesic SamplingabstractThis paper explores a fully unsupervised deep learning approach for computing distance-preserving maps that generate low-dimensional embeddings for a certain class of manifolds. We use the Siamese configuration to train a neural network to solve the problem of least squares multidimensional scaling for generating maps that approximately preserve geodesic distances. By training with only a few landmarks, we show a significantly improved local and nonlocal generalization of the isometric mapping as compared to analogous non-parametric counterparts. Importantly, the combination of a deep-learning framework with a multidimensional scaling objective enables a numerical analysis of network architectures to aid in understanding their representation power. This provides a geometric perspective to the generalizability of deep learning. Gautam Pai 0001, Ronen Talmon, Alexander M. Bronstein, Ron Kimmel |
WACV | 2 |
| 2019 | Intrinsic Isometric Manifold Learning with Application to LocalizationabstractData living on manifolds commonly appear in many applications. Often this results from observing an inherently latent low-dimensional system via higher-dimensional measurements. We show that, under certain conditions, it is possible to construct an intrinsic and isometric data representation for such data which respects an underlying latent intrinsic geometry. Namely, we view the observed data only as a proxy and learn the structure of a latent unobserved intrinsic manifold, whereas common practice is to learn the manifold of the observed data. For this purpose, we build a new metric and propose a method for its robust estimation by assuming mild statistical priors and by using artificial neural networks as a mechanism for metric regularization and parametrization. We show a successful application to unsupervised indoor localization in ad hoc sensor networks. Specifically, we show that our proposed method facilitates accurate localization of a moving agent from imaging data it acquires. Importantly, our method is applied in the same way to two different imaging modalities, thereby demonstrating its intrinsic and modality-invariant capabilities. Ariel Schwartz, Ronen Talmon |
SIAM J. Imaging Sci. | 2 |
| 2018 | Statistical Tomography of Microscopic LifeabstractWe achieve tomography of 3D volumetric natural objects, where each projected 2D image corresponds to a different specimen. Each specimen has unknown random 3D orientation, location, and scale. This imaging scenario is relevant to microscopic and mesoscopic organisms, aerosols and hydrosols viewed naturally by a microscope. In-class scale variation inhibits prior single-particle reconstruction methods. We thus generalize tomographic recovery to account for all degrees of freedom of a similarity transformation. This enables geometric self-calibration in imaging of transparent objects. We make the computational load manageable and reach good quality reconstruction in a short time. This enables extraction of statistics that are important for a scientific study of specimen populations, specifically size distribution parameters. We apply the method to study of plankton. Aviad Levis, Yoav Y. Schechner, Ronen Talmon |
CVPR | 3 |
| 2018 | Multi-View Source Localization Based on Power RatiosabstractDespite attracting significant research efforts, the problem of source localization in noisy and reverberant environments remains challenging. Novel learning-based methods attempt to solve the problem by modelling the acoustic environment from the observed data. Typically, appropriate feature vectors are defined, and then used for constructing a model, which maps the extracted features to the corresponding source positions. In this paper, we focus on localizing a source using a distributed network with several arrays of unidirectional microphones. We introduce new feature vectors, which utilize the special characteristic of unidirectional microphones, receiving different parts of the reverberated speech. The new features are computed locally for each array, using the power-ratios between its measured signals, and are used to construct a local model, representing the unique view point of each array. The models of the different arrays, conveying distinct and complementing structures, are merged by a Multi-View Gaussian Process (MVGP), mapping the new features to their corresponding source positions. Based on this unifying model, a Bayesian estimator is derived, exploiting the relations conveyed by the covariance terms of the MVGP. The resulting localizer is shown to be robust to noise and reverberation, utilizing a computationally efficient feature extraction. Bracha Laufer-Goldshtein, Ronen Talmon, Israel Cohen, Sharon Gannot |
ICASSP | 2 |
| 2018 | Multimodal latent variable analysis
Vardan Papyan, Ronen Talmon |
Signal Process. | 2 |
| 2018 | A Hybrid Approach for Speaker Tracking Based on TDOA and Data-Driven ModelsabstractThe problem of speaker tracking in noisy and reverberant enclosures is addressed in this paper. We present a hybrid algorithm, combining traditional tracking schemes with a new learning-based approach. A state-space representation, consisting of a propagation and observation models, is learned from signals measured by several distributed microphone pairs. The proposed representation is based on two data modalities corresponding to high-dimensional acoustic features representing the full reverberant acoustic channels as well as low-dimensional time difference of arrival (TDOA) estimates. The state-space representation is accompanied by a statistical model based on a Gaussian process used to relate the variations of the acoustic channels to the physical variations of the associated source positions, thereby forming a data-driven propagation model for the source movement. In the observation model, the source positions are nonlinearly mapped to the associated TDOA readings. The obtained propagation and observation models establish the basis for employing an extended Kalman filter. The simulation results demonstrate the robustness of the proposed method in noisy and reverberant conditions. Bracha Laufer-Goldshtein, Ronen Talmon, Sharon Gannot |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Alternating diffusion maps for dementia severity assessmentabstractIn this paper we address the detection of Alzheimer's disease based solely on EEG recordings. We assume that the state of Alzheimer's disease can be described by a latent manifold, captured by the EEG sensors and apply alternating diffusion to reveal this common underlying manifold from multiple EEG sensors. We show that based on a small number of EEG electrodes, a new representation can be obtained, which allows a clear distinction between healthy subjects and Alzheimer patients in different disease stages. Tal Shnitzer, Maya Rapaport, Noga Cohen, Natalya Yarovinsky, Ronen Talmon, Judith Aharon-Peretz |
ICASSP | 5 |
| 2017 | Dynamical system classification with diffusion embedding for ECG-based person identification
Jeremias Sulam, Yaniv Romano, Ronen Talmon |
Signal Process. | 3 |
| 2017 | Multimodal Kernel Method for Activity Detection of Sound SourcesabstractWe consider the problem of acoustic scene analysis of multiple sound sources. In our setting, the sound sources are measured by a single microphone, and a particular source of interest is also captured by a video camera during a short time interval. The goal in this paper is to detect the activity of the source of interest even when the video data are missing, while ignoring the other sound sources. To address this problem, we propose a kernel-based algorithm that incorporates the audio-visual data by a combination of affinity kernels, constructed separately from the audio and the video data. We introduce a distance measure between data points that is associated with the source of interest, while reducing the effect of the other (interfering) sources. Using this distance, we devise a measure for the presence of the source of interest, which is naturally extended to time intervals, in which only the audio signal is available. Experimental results demonstrate the improved performance of the proposed algorithm compared to competing approaches implying the significance of the video signal in the analysis of complex acoustic scenes. David Dov, Ronen Talmon, Israel Cohen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Semi-Supervised Source Localization on Multiple Manifolds With Distributed MicrophonesabstractThe problem of single-source localization with ad hoc microphone networks in noisy and reverberant enclosures is addressed in this paper. A training set is formed by prerecorded measurements collected in advance and consists of a limited number of labelled measurements, attached with corresponding positions, and a larger number of unlabelled measurements from unknown locations. Further information about the enclosure characteristics or the microphone positions is not required. We propose a Bayesian inference approach for estimating a function that maps measurement-based features to the corresponding positions. The signals measured by the microphones represent different viewpoints, which are combined in a unified statistical framework. For this purpose, the mapping function is modelled by a Gaussian process with a covariance function that encapsulates both the connections between pairs of microphones and the relations among the samples in the training set. The parameters of the process are estimated by optimizing a maximum likelihood criterion. In addition, a recursive adaptation mechanism is derived, where the new streaming measurements are used to update the model. Performance is demonstrated for both simulated data and real-life recordings in a variety of reverberation and noise levels. Bracha Laufer-Goldshtein, Ronen Talmon, Sharon Gannot |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Improving resolution in supervised patch-based target detectionabstractRecently, a supervised graph-based target detection method was proposed based on a new affinity measure between a set of target training patches and a test image. In this paper, we propose a new high-resolution detection score, which enhances the performance of the previous method by utilizing the known locations of the targets in the training images. We show that our new score is more reliable and spatially accurate, not only improving the detection resolution of true targets, but also reducing the number of false alarms. The method is successfully tested on side-scan sonar images of sea-mines, demonstrating an improved true detection rate. Our approach is general and can improve the detection resolution of the target in other patch-based detection algorithms for various signals and applications. Ron Amit, Gal Mishne, Ronen Talmon |
ICASSP | 3 |
| 2016 | Manifold-based Bayesian inference for semi-supervised source localizationabstractSound source localization is addressed by a novel Bayesian approach using a data-driven geometric model. The goal is to recover the target function that attaches each acoustic sample, formed by the measured signals, with its corresponding position. The estimation is derived by maximizing the posterior probability of the target function, computed on the basis of acoustic samples from known locations (labelled data) as well as acoustic samples from unknown locations (unlabelled data). To form the posterior probability we use a manifold-based prior, which relies on the geometric structure of the manifold from which the acoustic samples are drawn. The proposed method is shown to be analogous to a recently presented semi-supervised localization approach based on manifold regularization. Simulation results demonstrate the robustness of the method in noisy and reverberant environments. Bracha Laufer-Goldshtein, Ronen Talmon, Sharon Gannot |
ICASSP | 2 |
| 2016 | Kernel Method for Voice Activity Detection in the Presence of TransientsabstractVoice activity detection in the presence of transient interferences is a challenging problem since transients are often detected incorrectly as speech by existing detectors. In this paper, we deviate from traditional approaches and take a geometric standpoint, in which the key element in obtaining an accurate voice activity detection is finding a metric that appropriately distinguishes between speech and transients. For example, speech and transients may often appear similar through the Euclidean distance when represented, e.g., by the Mel-frequency cepstral coefficients, thereby resulting in incorrect speech detection. To address this challenge, we propose to use a metric based on the statistics of the signal in short temporal windows and justify its use by modeling speech and transients by their latent generating variables. These latent variables may be related to physical constraints controlling the generation of the signal, and, as such, they accurately represent the content of the signal - speech or transient. We show that the Euclidean distance between the latent variables is approximated by the proposed metric. Then, by incorporating this metric into a kernel-based manifold learning method, we devise a measure of voice activity and show it leads to improved detection scores compared with competing detectors. David Dov, Ronen Talmon, Israel Cohen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Semi-Supervised Sound Source Localization Based on Manifold RegularizationabstractConventional speaker localization algorithms, based merely on the received microphone signals, are often sensitive to adverse conditions, such as: high reverberation or low signal-to-noise ratio (SNR). In some scenarios, e.g., in meeting rooms or cars, it can be assumed that the source position is confined to a predefined area, and the acoustic parameters of the environment are approximately fixed. Such scenarios give rise to the assumption that the acoustic samples from the region of interest have a distinct geometrical structure. In this paper, we show that the high-dimensional acoustic samples indeed lie on a low-dimensional manifold and can be embedded into a low-dimensional space. Motivated by this result, we propose a semi-supervised source localization algorithm based on two-microphone measurements, which recovers the inverse mapping between the acoustic samples and their corresponding locations. The idea is to use an optimization framework based on manifold regularization, that involves smoothness constraints of possible solutions with respect to the manifold. The proposed algorithm, termed manifold regularization for localization, is adapted while new unlabelled measurements (from unknown source locations) are accumulated during runtime. Experimental results show superior localization performance when compared with a recently presented algorithm based on a manifold learning approach and with the generalized cross-correlation algorithm as a baseline. The algorithm achieves 2° accuracy in typical noisy and reverberant environments (reverberation time between 200 and 800 ms and SNR between 5 and 20 dB). Bracha Laufer-Goldshtein, Ronen Talmon, Sharon Gannot |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Alternating diffusion for common manifold learning with application to sleep stage assessmentabstractIn this paper, we address the problem of multimodal signal processing and present a manifold learning method to extract the common source of variability from multiple measurements. This method is based on alternating-diffusion and is particularly adapted to time series. We show that the common source of variability is extracted from multiple sensors as if it were the only source of variability, extracted by a standard manifold learning method from a single sensor, without the influence of the sensor-specific variables. In addition, we present application to sleep stage assessment. We demonstrate that, indeed, through alternating-diffusion, the sleep information hidden inside multimodal respiratory signals can be better captured compared to single-modal methods. Roy R. Lederman, Ronen Talmon, Hau-Tieng Wu, Yu-Lun Lo, Ronald R. Coifman |
ICASSP | 2 |
| 2015 | Multivariate time-series analysis and diffusion maps
Wenzhao Lian, Ronen Talmon, Hitten Zaveri, Lawrence Carin, Ronald R. Coifman |
Signal Process. | 2 |
| 2015 | Audio-Visual Voice Activity Detection Using Diffusion MapsabstractThe performance of traditional voice activity detectors significantly deteriorates in the presence of highly nonstationary noise and transient interferences. One solution is to incorporate a video signal which is invariant to the acoustic environment. Although several voice activity detectors based on the video signal were recently presented, merely few detectors which are based on both the audio and the video signals exist in the literature to date. In this paper, we present an audio-visual voice activity detector and show that the incorporation of both audio and video signals is highly beneficial for voice activity detection. The algorithm is based on a supervised learning procedure, and a labeled training data set is considered. The algorithm comprises a feature extraction procedure, where the features are designed to separate speech from nonspeech frames. Diffusion maps is applied separately and similarly to the features of each modality and builds a low dimensional representation. Using the new representation, we propose a measure for voice activity which is based on a supervised learning procedure and the variability between adjacent frames in time. The measures of the two modalities are merged to provide voice activity detection based on both the audio and the video signals. Experimental results demonstrate the improved performance of the proposed algorithm compared to state-of-the-art detectors. David Dov, Ronen Talmon, Israel Cohen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Graph-Based Supervised Automatic Target DetectionabstractIn this paper, we propose a detection method based on data-driven target modeling, which implicitly handles variations in the target appearance. Given a training set of images of the target, our approach constructs models based on local neighborhoods within the training set. We present a new metric using these models and show that, by controlling the notion of locality within the training set, this metric is invariant to perturbations in the appearance of the target. Using this metric in a supervised graph framework, we construct a low-dimensional embedding of test images. Then, a detection score based on the embedding determines the presence of a target in each image. The method is applied to a data set of side-scan sonar images and achieves impressive results in the detection of sea mines. The proposed framework is general and can be applied to different target detection problems in a broad range of signals. Gal Mishne, Ronen Talmon, Israel Cohen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2013 | Blind reverberation time estimation by intrinsic modeling of reverberant speechabstractThe reverberation time (RT) is a very important measure that quantifies the acoustic properties of a room and provides information about the quality and intelligibility of speech recorded in that room. Moreover, information about the RT can be used to improve the performance of automatic speech recognition systems and speech dereverberation algorithms. In a recent study, it has been shown that existing methods for blind estimation of the RT are highly sensitive to additive noise. In this paper, a novel method is proposed to blindly estimate the RT based on the decay rate distribution. Firstly, a data-driven representation of the underlying decay rates of several training rooms is obtained via the eigenvalue decomposition of a specially-tailored kernel. Secondly, the representation is extended to a room under test and used to estimate its decay rate (and hence its RT). The presented results show that the proposed method outperforms a competing method and is significantly more robust to noise. Ronen Talmon, Emanuël A. P. Habets |
ICASSP | 1 |
| 2013 | Single-Channel Transient Interference Suppression With Diffusion MapsabstractA transient is an abrupt or impulsive sound followed by decaying oscillations, e.g., keyboard typing and door knocking. Such sounds often arise as interference in everyday applications, e.g., hearing aids, hands-free accessories, mobile phones, and conference-room devices. In this paper, we present an algorithm for single-channel transient interference suppression. The main component of the proposed algorithm is the estimation of the spectral variance of the interference. We propose a statistical model of the transient interference and combine it with non-local filtering. We exploit the unique spectral structure of the transients along with their impulsive temporal nature to distinct them from speech. A particular attention is given to handling both short- and long-duration transients. Experimental results show that the proposed algorithm enables significant transient suppression for a variety of transient types. Ronen Talmon, Israel Cohen, Sharon Gannot |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Supervised system identification based on local PCA modelsabstractWe propose a supervised system identification method for recovering an acoustic impulse response in a reverberant room. Unlike most existing methods, our algorithm is based on prior information given in the form of a training set of known impulse responses acquired in a controlled environment. By relying on the prior information, we train local Principal Component Analysis (PCA) models of impulse responses corresponding to several different regions in the room. We propose to crudely localize the respective source position, and subsequently, based on the appropriate local model, recover the impulse response. In order to approximate the source location, we introduce a specially-tailored distance measure which is based on an affinity between the trained local models. Experimental results in simulated noisy and reverberant environments demonstrate significant improvements over existing methods. Tomer Koren, Ronen Talmon, Israel Cohen |
ICASSP | 2 |
| 2012 | Supervised Graph-Based Processing for Sequential Transient Interference SuppressionabstractIn this paper, we present a supervised graph-based framework for sequential processing and employ it to the problem of transient interference suppression. Transients typically consist of an initial peak followed by decaying short-duration oscillations. Such sounds, e.g., keyboard typing and door knocking, often arise as an interference in everyday applications: hearing aids, hands-free accessories, mobile phones, and conference-room devices. We describe a graph construction using a noisy speech signal and training recordings of typical transients. The main idea is to capture the transient interference structure, which may emerge from the construction of the graph. The graph parametrization is then viewed as a data-driven model of the transients and utilized to define a filter that extracts the transients from noisy speech measurements. Unlike previous transient interference suppression studies, in this work the graph is constructed in advance from training recordings. Then, the graph is extended to newly acquired measurements, providing a sequential filtering framework of noisy speech. Ronen Talmon, Israel Cohen, Sharon Gannot, Ronald R. Coifman |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Clustering and suppression of transient noise in speech signals using diffusion mapsabstractRecently we have presented a novel approach for transient noise reduction that relies on non-local (NL) filtering. In this paper, we modify and extend our approach to support clustering and suppression of a few transient noise types simultaneously, by introducing two novel concepts. We observe that voiced speech spectral components are slowly varying compared to transient noise. Thus, by applying an algorithm for noise power spectral density (PSD) estimation, configured to track faster variations than pseudo-stationary noise, the PSD of speech components may be estimated. In addition, we utilize diffusion maps to embed the measurements into a new do main. We obtain a new representation which enables clustering of different transient noise types. The new representation is incorporated into a NL filter as a better affinity metric for averaging over transient instances. Experimental results show that the proposed algorithm enables clustering and suppression of multiple transient interferences. Ronen Talmon, Israel Cohen, Sharon Gannot |
ICASSP | 1 |
| 2011 | Transient Noise Reduction Using Nonlocal Diffusion FiltersabstractEnhancement of speech signals for hands-free communication systems has attracted significant research efforts in the last few decades. Still, many aspects and applications remain open and require further research. One of the important open problems is the single-channel transient noise reduction. In this paper, we present a novel approach for transient noise reduction that relies on non-local (NL) neighborhood filters. In particular, we propose an algorithm for the enhancement of a speech signal contaminated by repeating transient noise events. We assume that the time duration of each reoccurring transient event is relatively short compared to speech phonemes and model the speech source as an auto-regressive (AR) process. The proposed algorithm consists of two stages. In the first stage, we estimate the power spectral density (PSD) of the transient noise by employing a NL neighborhood filter. In the second stage, we utilize the optimally modified log spectral amplitude (OM-LSA) estimator for denoising the speech using the noise PSD estimate from the first stage. Based on a statistical model for the measurements and diffusion interpretation of NL filtering, we obtain further insight into the algorithm behavior. In particular, for given transient noise, we determine whether estimation of the noise PSD is feasible using our approach, how to properly set the algorithm parameters, and what is the expected performance of the algorithm. Experimental study shows good results in enhancing speech signals contaminated by transient noise, such as typical household noises, construction sounds, keyboard typing, and metronome clacks. Ronen Talmon, Israel Cohen, Sharon Gannot |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Speech enhancement in transient noise environment using diffusion filteringabstractRecently, we have presented a transient noise reduction algorithm for speech signals that relies on non-local diffusion filtering. By exploiting the repetitive nature of transient noises we proposed a simple and efficient algorithm, which enabled suppression of various noise types. In this paper, we incorporate a modified diffusion operator in order to obtain a more robust algorithm and further enhancement of the speech. We demonstrate the performance of the modified algorithm and compare it with a competing solution. We show that the proposed algorithm enables improved suppression of various transient interferences without any further computational burden. Ronen Talmon, Israel Cohen, Sharon Gannot |
ICASSP | 1 |
| 2009 | Multichannel speech enhancement using convolutive transfer function approximation in reverberant environmentsabstractRecently, we have presented a transfer-function generalized sidelobe canceler (TF-GSC) beamformer in the short time Fourier transform domain, which relies on a convolutive transfer function approximation of relative transfer functions between distinct sensors. In this paper, we combine a delay-and-sum beamformer with the TF-GSC structure in order to suppress the speech signal reflections captured at the sensors in reverberant environments. We demonstrate the performance of the proposed beamformer and compare it with the TF-GSC. We show that the proposed algorithm enables suppression of reverberations and further noise reduction compared with the TF-GSC beamformer. Ronen Talmon, Israel Cohen, Sharon Gannot |
ICASSP | 1 |
| 2009 | Relative Transfer Function Identification Using Convolutive Transfer Function ApproximationabstractIn this paper, we present a relative transfer function (RTF) identification method for speech sources in reverberant environments. The proposed method is based on the convolutive transfer function (CTF) approximation, which enables to represent a linear convolution in the time domain as a linear convolution in the short-time Fourier transform (STFT) domain. Unlike the restrictive and commonly used multiplicative transfer function (MTF) approximation, which becomes more accurate when the length of a time frame increases relative to the length of the impulse response, the CTF approximation enables representation of long impulse responses using short time frames. We develop an unbiased RTF estimator that exploits the nonstationarity and presence probability of the speech signal and derive an analytic expression for the estimator variance. Experimental results show that the proposed method is advantageous compared to common RTF identification methods in various acoustic environments, especially when identifying long RTFs typical to real rooms. Ronen Talmon, Israel Cohen, Sharon Gannot |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Convolutive Transfer Function Generalized Sidelobe CancelerabstractIn this paper, we propose a convolutive transfer function generalized sidelobe canceler (CTF-GSC), which is an adaptive beamformer designed for multichannel speech enhancement in reverberant environments. Using a complete system representation in the short-time Fourier transform (STFT) domain, we formulate a constrained minimization problem of total output noise power subject to the constraint that the signal component of the output is the desired signal, up to some prespecified filter. Then, we employ the general sidelobe canceler (GSC) structure to transform the problem into an equivalent unconstrained form by decoupling the constraint and the minimization. The CTF-GSC is obtained by applying a convolutive transfer function (CTF) approximation on the GSC scheme, which is a more accurate and a less restrictive than a multiplicative transfer function (MTF) approximation. Experimental results demonstrate that the proposed beamformer outperforms the transfer function GSC (TF-GSC) in reverberant environments and achieves both improved noise reduction and reduced speech distortion. Ronen Talmon, Israel Cohen, Sharon Gannot |
IEEE Trans. Speech Audio Process. | 1 |