Gal Mishne

dblp:125/3214 · DBLP profile ↗
← Back
31ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0002-5287-3626ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 2 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unsupervised Feature Selection Through Group Discovery
abstract
Unsupervised feature selection (FS) is essential for high-dimensional learning tasks where labels are not available. It helps reduce noise, improve generalization, and enhance interpretability. However, most existing unsupervised FS methods evaluate features in isolation, even though informative signals often emerge from groups of related features. For example, adjacent pixels, functionally connected brain regions, or correlated financial indicators tend to act together, making independent evaluation suboptimal. Although some methods attempt to capture group structure, they typically rely on predefined partitions or label supervision, limiting their applicability. We propose GroupFS, an end-to-end, fully differentiable framework that jointly discovers latent feature groups and selects the most informative groups among them, without relying on fixed a priori groups or label supervision. GroupFS enforces Laplacian smoothness on both feature and sample graphs and applies a group sparsity regularizer to learn a compact, structured representation. Across nine benchmarks spanning images, tabular data, and biological datasets, GroupFS consistently outperforms state-of-the-art unsupervised FS in clustering and selects groups of features that align with meaningful patterns.
Shira Lifshitz, Ofir Lindenbaum, Gal Mishne, Ron Meir, Hadas Benisty
AAAI3
2026 Fast and accessible morphology-free functional fluorescence imaging analysis
abstract
Optical calcium imaging is a powerful tool for recording neural activity across a wide range of spatial scales, from dendrites and spines to whole-brain imaging through two-photon and widefield microscopy. Traditional methods for analyzing functional calcium imaging data rely heavily on spatial features, such as the compact shapes of somas, to extract regions of interest and their associated temporal traces. This spatial dependency can introduce biases in time trace estimation and limit the applicability of these methods across different neuronal morphologies and imaging scales. To address these limitations, the Graph Filtered Temporal Dictionary Learning (GraFT) uses a graph-based approach to identify neural components based on shared temporal activity rather than spatial proximity, enhancing generalizability across diverse datasets. Here we present significant advancements to the GraFT algorithm, including the integration of a more efficient solver for the L1 least absolute shrinkage and selection operator (LASSO) problem and the application of compressive sensing techniques to reduce computational complexity. By employing random projections to reduce data dimensionality, we achieve substantial speedups while maintaining analytical accuracy. These advancements significantly accelerate the GraFT algorithm, making it more scalable for larger and more complex datasets. Moreover, to increase accessibility, we developed a graphical user interface to facilitate running and analyzing the outputs of GraFT. Finally, we demonstrate the utility of GraFT to imaging data beyond meso-scale imaging, including vascular and axonal imaging.
Alejandro Estrada Berlanga, Gabrielle Y. Kang, Amanda Kwok, Thomas Broggini, Jennifer Lawlor, Kishore V. Kuchibhotla, David Kleinfeld, Gal Mishne, Adam S. Charles
PLoS Comput. Biol.8
2025 RnGCam: High-Speed Video from Rolling & Global Shutter Measurements
abstract
Compressive video capture encodes a short high-speed video into a single measurement using a low-speed sensor, then computationally reconstructs the original video. Prior implementations rely on expensive hardware and are restricted to imaging sparse scenes with empty backgrounds. We propose RnGCam, a system that fuses measurements from low-speed consumer-grade rolling-shutter (RS) and global-shutter (GS) sensors into video at kHz frame rates. The RS sensor is combined with a pseudorandom optic, called a diffuser, which spatially multiplexes scene information. The GS sensor is coupled with a conventional lens. The RS-diffuser provides low spatial detail and high temporal detail, complementing the GS-lens system's high spatial detail and low temporal detail. We propose a reconstruction method using implicit neural representations (INR) to fuse the measurements into a high-speed video. Our INR method separately models the static and dynamic scene components, while explicitly regularizing dynamics. In simulation, we show that our approach significantly outperforms previous RS compressive video methods, as well as state-of-the-art frame interpolators. We validate our approach in a dual-camera hardware setup, which generates 230 frames of video at 4,800 frames per second for dense scenes, using hardware that costs $10 \times$ less than previous compressive video systems.
Kevin Tandi, Chinmay Talegaonkar, Gal Mishne, Nicholas Antipa
ICCV4
2025 Tree-Wasserstein Distance for High Dimensional Data with a Latent Feature Hierarchy
abstract
Finding meaningful distances between high-dimensional data samples is an important scientific task. To this end, we propose a new tree-Wasserstein distance (TWD) for high-dimensional data with two key aspects. First, our TWD is specifically designed for data with a latent feature hierarchy, i.e., the features lie in a hierarchical space, in contrast to the usual focus on embedding samples in hyperbolic space. Second, while the conventional use of TWD is to speed up the computation of the Wasserstein distance, we use its inherent tree as a means to learn the latent feature hierarchy. The key idea of our method is to embed the features into a multi-scale hyperbolic space using diffusion geometry and then present a new tree decoding method by establishing analogies between the hyperbolic embedding and trees. We show that our TWD computed based on data observations provably recovers the TWD defined with the latent feature hierarchy and that its computation is efficient and scalable. We showcase the usefulness of the proposed TWD in applications to word-document and single-cell RNA-sequencing datasets, demonstrating its advantages over existing TWDs and methods based on pre-trained models.
Ya-Wei Eileen Lin, Ronald R. Coifman, Gal Mishne, Ronen Talmon
ICLR3
2025 Elucidating Flow Matching ODE Dynamics via Data Geometry and Denoisers
abstract
Flow matching (FM) models extend ODE sampler based diffusion models into a general framework, significantly reducing sampling steps through learned vector fields. However, the theoretical understanding of FM models, particularly how their sample trajectories interact with underlying data geometry, remains underexplored. A rigorous theoretical analysis of FM ODE is essential for sample quality, stability, and broader applicability. In this paper, we advance the theory of FM models through a comprehensive analysis of sample trajectories. Central to our theory is the discovery that the denoiser, a key component of FM models, guides ODE dynamics through attracting and absorbing behaviors that adapt to the data geometry. We identify and analyze the three stages of ODE evolution: in the initial and intermediate stages, trajectories move toward the mean and local clusters of the data. At the terminal stage, we rigorously establish the convergence of FM ODE under weak assumptions, addressing scenarios where the data lie on a low-dimensional submanifold—cases that previous results could not handle. Our terminal stage analysis offers insights into the memorization phenomenon and establishes equivariance properties of FM ODEs. These findings bridge critical gaps in understanding flow matching models, with practical implications for optimizing sampling strategies and architectures guided by the intrinsic geometry of data.
Zhengchao Wan, Gal Mishne, Yusu Wang 0001
ICML3
2025 Word-Level Error Analysis in Decoding Systems: From Speech Recognition to Brain-Computer Interfaces
Jingya Huang, Aashish N. Patel, Sowmya Manojna Narasimha, Gal Mishne, Vikash Gilja
INTERSPEECH4
2025 Explaining GNN Explanations with Edge Gradients
abstract
In recent years, the remarkable success of graph neural networks (GNNs) on graph-structured data has prompted a surge of methods for explaining GNN predictions. However, the state-of-the-art for GNN explainability remains in flux. Different comparisons find mixed results for different methods, with many explainers struggling on more complex GNN architectures and tasks. This presents an urgent need for a more careful theoretical analysis of competing GNN explanation methods. In this work we take a closer look at GNN explanations in two different settings: input-level explanations, which produce explanatory subgraphs of the input graph, and layerwise explanations, which produce explanatory subgraphs of the computation graph. We establish the first theoretical connections between the popular perturbation-based and classical gradient-based methods, as well as point out connections between other recently proposed methods. At the input level, we demonstrate conditions under which GNNExplainer can be approximated by a simple heuristic based on the sign of the edge gradients. In the layerwise setting, we point out that edge gradients are equivalent to occlusion search for linear GNNs. Finally, we demonstrate how our theoretical results manifest in practice with experiments on both synthetic and real datasets.
Jesse He, Akbar Rafiey, Gal Mishne, Yusu Wang 0001
KDD (2)3
2025 Joint Hierarchical Representation Learning of Samples and Features via Informed Tree-Wasserstein Distance
abstract
High-dimensional data often exhibit hierarchical structures in both modes: samples and features. Yet, most existing approaches for hierarchical representation learning consider only one mode at a time. In this work, we propose an unsupervised method for jointly learning hierarchical representations of samples and features via Tree-Wasserstein Distance (TWD). Our method alternates between the two data modes. It first constructs a tree for one mode, then computes a TWD for the other mode based on that tree, and finally uses the resulting TWD to build the second mode’s tree. By repeatedly alternating through these steps, the method gradually refines both trees and the corresponding TWDs, capturing meaningful hierarchical representations of the data. We provide a theoretical analysis showing that our method converges. We show that our method can be integrated into hyperbolic graph convolutional networks as a pre-processing technique, improving performance in link prediction and node classification tasks. In addition, our method outperforms baselines in sparse approximation and unsupervised Wasserstein distance learning tasks on word-document and single-cell RNA-sequencing datasets.
Ya-Wei Eileen Lin, Ronald R. Coifman, Gal Mishne, Ronen Talmon
NeurIPS3
2024 Learning Cartesian Product Graphs with Laplacian Constraints
abstract
Graph Laplacian learning, also known as network topology inference, is a problem of great interest to multiple communities. In Gaussian graphical models (GM), graph learning amounts to endowing covariance selection with the Laplacian structure. In graph signal processing (GSP), it is essential to infer the unobserved graph from the outputs of a filtering system. In this paper, we study the problem of learning Cartesian product graphs under Laplacian constraints. The Cartesian graph product is a natural way for modeling higher-order conditional dependencies and is also the key for generalizing GSP to multi-way tensors. We establish statistical consistency for the penalized maximum likelihood estimation (MLE) of a Cartesian product Laplacian, and propose an efficient algorithm to solve the problem. We also extend our method for efficient joint graph learning and imputation in the presence of structural missing values. Experiments on synthetic and real-world datasets demonstrate that our method is superior to previous GSP and GM methods.
Changhao Shi, Gal Mishne
AISTATS2
2024 Comparing Graph Transformers via Positional Encodings
abstract
The distinguishing power of graph transformers is tied to the choice of positional encoding: features used to augment the base transformer with information about the graph. There are two primary types of positional encoding: absolute positional encodings (APEs) and relative positional encodings (RPEs). APEs assign features to each node and are given as input to the transformer. RPEs instead assign a feature to each pair of nodes, e.g., shortest-path distance, and are used to augment the attention block. A priori, it is unclear which method is better for maximizing the power of the resulting graph transformer. In this paper, we aim to understand the relationship between these different types of positional encodings. Interestingly, we show that graph transformers using APEs and RPEs are equivalent in their ability to distinguish non-isomorphic graphs. In particular, we demonstrate how to interchange APEs and RPEs while maintaining their distinguishing power in terms of graph transformers. However, in the case of graphs with node features, we show that RPEs may have an advantage over APEs. Based on our theoretical results, we provide a study of different APEs and RPEs—including the shortest-path and resistance distance and the recently introduced stable and expressive positional encoding (SPE)—and compare their distinguishing power in terms of transformers. We believe our work will help navigate the vast number of positional encoding choices and provide guidance on the future design of positional encodings for graph transformers.
Mitchell Black 0002, Zhengchao Wan, Gal Mishne, Amir Nayyeri, Yusu Wang 0001
ICML3
2024 SiBBlInGS: Similarity-driven Building-Block Inference using Graphs across States
abstract
Time series data across scientific domains are often collected under distinct states (e.g., tasks), wherein latent processes (e.g., biological factors) create complex inter- and intra-state variability. A key approach to capture this complexity is to uncover fundamental interpretable units within the data, Building Blocks (BBs), which modulate their activity and adjust their structure across observations. Existing methods for identifying BBs in multi-way data often overlook inter- vs. intra-state variability, produce uninterpretable components, or do not align with properties of real-world data, such as missing samples and sessions of different duration. Here, we present a framework for Similarity-driven Building Block Inference using Graphs across States (SiBBlInGS). SiBBlInGS offers a graph-based dictionary learning approach for discovering sparse BBs along with their temporal traces, based on co-activity patterns and inter- vs. intra-state relationships. Moreover, SiBBlInGS captures per-trial temporal variability and controlled cross-state structural BB adaptations, identifies state-specific vs. state-invariant components, and accommodates variability in the number and duration of observed sessions across states. We demonstrate SiBBlInGS's ability to reveal insights into complex phenomena as well as its robustness to noise and missing samples through several synthetic and real-world examples, including web search and neural data.
Noga Mudrik, Gal Mishne, Adam S. Charles
ICML2
2024 Contextual Feature Selection with Conditional Stochastic Gates
abstract
Feature selection is a crucial tool in machine learning and is widely applied across various scientific disciplines. Traditional supervised methods generally identify a universal set of informative features for the entire population. However, feature relevance often varies with context, while the context itself may not directly affect the outcome variable. Here, we propose a novel architecture for contextual feature selection where the subset of selected features is conditioned on the value of *context variables*. Our new approach, Conditional Stochastic Gates (c-STG), models the importance of features using conditional Bernoulli variables whose parameters are predicted based on contextual variables. We introduce a hypernetwork that maps context variables to feature selection parameters to learn the context-dependent gates along with a prediction model. We further present a theoretical analysis of our model, indicating that it can improve performance and flexibility over population-level methods in complex feature selection settings. Finally, we conduct an extensive benchmark using simulated and real-world datasets across multiple domains demonstrating that c-STG can lead to improved feature selection capabilities while enhancing prediction accuracy and interpretability.
Ram Dyuthi Sristi, Ofir Lindenbaum, Shira Lifshitz, Maria Lavzin, Jackie Schiller, Gal Mishne, Hadas Benisty
ICML6
2024 Continuous Partitioning for Graph-Based Semi-Supervised Learning
abstract
Laplace learning algorithms for graph-based semi-supervised learning have been shown to produce degenerate predictions at low label rates and in imbalanced class regimes, particularly near class boundaries. We propose CutSSL: a framework for graph-based semi-supervised learning based on continuous nonconvex quadratic programming, which provably obtains \emph{integer} solutions. Our framework is naturally motivated by an \emph{exact} quadratic relaxation of a cardinality-constrained minimum-cut graph partitioning problem. Furthermore, we show our formulation is related to an optimization problem whose approximate solution is the mean-shifted Laplace learning heuristic, thus providing new insight into the performance of this heuristic. We demonstrate that CutSSL significantly surpasses the current state-of-the-art on k-nearest neighbor graphs and large real-world graph benchmarks across a variety of label rates, class imbalance, and label imbalance regimes. Our implementation is available on Colab\footnote{\url{https://colab.research.google.com/drive/1tGU5rxE1N5d0KGcNzlvZ0BgRc7_vob7b?usp=sharing}}.
Chester Holtz, Pengwen Chen, Zhengchao Wan, Chung-Kuan Cheng, Gal Mishne
NeurIPS5
2023 Implicit Graphon Neural Representation
abstract
Graphons are general and powerful models for generating graphs of varying size. In this paper, we propose to directly model graphons using neural networks, obtaining Implicit Graphon Neural Representation (IGNR). Existing work in modeling and reconstructing graphons often approximates a target graphon by a fixed resolution piece-wise constant representation. Our IGNR has the benefit that it can represent graphons up to arbitrary resolutions, and enables natural and efficient generation of arbitrary sized graphs with desired structure once the model is learned. Furthermore, we allow the input graph data to be unaligned and have different sizes by leveraging the Gromov-Wasserstein distance. We first demonstrate the effectiveness of our model by showing its superior performance on a graphon learning task. We then propose an extension of IGNR that can be incorporated into an auto-encoder framework, and demonstrate its good performance under a more general setting of graphon learning. We also show that our model is suitable for graph representation learning and graph generation.
Xinyue Xia, Gal Mishne, Yusu Wang 0001
AISTATS2
2023 Hyperbolic Diffusion Embedding and Distance for Hierarchical Representation Learning
abstract
Finding meaningful representations and distances of hierarchical data is important in many fields. This paper presents a new method for hierarchical data embedding and distance. Our method relies on combining diffusion geometry, a central approach to manifold learning, and hyperbolic geometry. Specifically, using diffusion geometry, we build multi-scale densities on the data, aimed to reveal their hierarchical structure, and then embed them into a product of hyperbolic spaces. We show theoretically that our embedding and distance recover the underlying hierarchical structure. In addition, we demonstrate the efficacy of the proposed method and its advantages compared to existing methods on graph embedding benchmarks and hierarchical datasets.
Ya-Wei Eileen Lin, Ronald R. Coifman, Gal Mishne, Ronen Talmon
ICML3
2023 The Numerical Stability of Hyperbolic Representation Learning
abstract
The hyperbolic space is widely used for representing hierarchical datasets due to its ability to embed trees with small distortion. However, this property comes at a price of numerical instability such that training hyperbolic learning models will sometimes lead to catastrophic NaN problems, encountering unrepresentable values in floating point arithmetic. In this work, we analyze the limitations of two popular models for the hyperbolic space, namely, the Poincaré ball and the Lorentz model. We find that, under the 64-bit arithmetic system, the Poincaré ball has a relatively larger capacity than the Lorentz model for correctly representing points. However, the Lorentz model is superior to the Poincaré ball from the perspective of optimization, which we theoretically validate. To address these limitations, we identify one Euclidean parametrization of the hyperbolic space which can alleviate these issues. We further extend this Euclidean parametrization to hyperbolic hyperplanes and demonstrate its effectiveness in improving the performance of hyperbolic SVM.
Gal Mishne, Zhengchao Wan, Yusu Wang 0001, Sheng Yang 0004
ICML1
2022 DiSC: Differential Spectral Clustering of Features
abstract
Selecting subsets of features that differentiate between two conditions is a key task in a broad range of scientific domains. In many applications, the features of interest form clusters with similar effects on the data at hand. To recover such clusters we develop DiSC, a data-driven approach for detecting groups of features that differentiate between conditions. For each condition, we construct a graph whose nodes correspond to the features and whose weights are functions of the similarity between them for that condition. We then apply a spectral approach to compute subsets of nodes whose connectivity pattern differs significantly between the condition-specific feature graphs. On the theoretical front, we analyze our approach with a toy example based on the stochastic block model. We evaluate DiSC on a variety of datasets, including MNIST, hyperspectral imaging, simulated scRNA-seq and task fMRI, and demonstrate that DiSC uncovers features that better differentiate between conditions compared to competing methods.
Ram Dyuthi Sristi, Gal Mishne, Ariel Jaffe
NeurIPS2
2022 GraFT: Graph Filtered Temporal Dictionary Learning for Functional Neural Imaging
abstract
Optical imaging of calcium signals in the brain has enabled researchers to observe the activity of hundreds-to-thousands of individual neurons simultaneously. Current methods predominantly use morphological information, typically focusing on expected shapes of cell bodies, to better identify neurons in the field-of-view. The explicit shape constraints limit the applicability of automated cell identification to other important imaging scales with more complex morphologies, e.g., dendritic or widefield imaging. Specifically, fluorescing components may be broken up, incompletely found, or merged in ways that do not accurately describe the underlying neural activity. Here we present Graph Filtered Temporal Dictionary (GraFT), a new approach that frames the problem of isolating independent fluorescing components as a dictionary learning problem. Specifically, we focus on the time-traces-the main quantity used in scientific discovery-and learn a time trace dictionary with the spatial maps acting as the presence coefficients encoding which pixels the time-traces are active in. Furthermore, we present a novel graph filtering model which redefines connectivity between pixels in terms of their shared temporal activity, rather than spatial proximity. This model greatly eases the ability of our method to handle data with complex non-local spatial structure. We demonstrate important properties of our method, such as robustness to morphology, simultaneously detecting different neuronal types, and implicitly inferring number of neurons, on both synthetic data and real data examples. Specifically, we demonstrate applications of our method to calcium imaging both at the dendritic, somatic, and widefield scales.
Adam S. Charles, Nathan Cermak, Rifqi O. Affan, Benjamin B. Scott, Jackie Schiller, Gal Mishne
IEEE Trans. Image Process.6
2021 Online Adversarial Purification based on Self-supervised Learning
Changhao Shi, Chester Holtz, Gal Mishne
ICLR3
2021 Learning Disentangled Behavior Embeddings
abstract
To understand the relationship between behavior and neural activity, experiments in neuroscience often include an animal performing a repeated behavior such as a motor task. Recent progress in computer vision and deep learning has shown great potential in the automated analysis of behavior by leveraging large and high-quality video datasets. In this paper, we design Disentangled Behavior Embedding (DBE) to learn robust behavioral embeddings from unlabeled, multi-view, high-resolution behavioral videos across different animals and multiple sessions. We further combine DBE with a stochastic temporal model to propose Variational Disentangled Behavior Embedding (VDBE), an end-to-end approach that learns meaningful discrete behavior representations and generates interpretable behavioral videos. Our models learn consistent behavior representations by explicitly disentangling the dynamic behavioral factors (pose) from time-invariant, non-behavioral nuisance factors (context) in a deep autoencoder, and exploit the temporal structures of pose dynamics. Compared to competing approaches, DBE and VDBE enjoy superior performance on downstream tasks such as fine-grained behavioral motif generation and behavior decoding.
Changhao Shi, Sivan Schwartz, Shahar Levy, Shay Achvat, Maisan Abboud, Amir Ghanayim, Jackie Schiller, Gal Mishne
NeurIPS8
2021 COBRAC: a fast implementation of convex biclustering with compression
abstract
SUMMARY: Biclustering is a generalization of clustering used to identify simultaneous grouping patterns in observations (rows) and features (columns) of a data matrix. Recently, the biclustering task has been formulated as a convex optimization problem. While this convex recasting of the problem has attractive properties, existing algorithms do not scale well. To address this problem and make convex biclustering a practical tool for analyzing larger data, we propose an implementation of fast convex biclustering called COBRAC to reduce the computing time by iteratively compressing problem size along with the solution path. We apply COBRAC to several gene expression datasets to demonstrate its effectiveness and efficiency. Besides the standalone version for COBRAC, we also developed a related online web server for online calculation and visualization of the downloadable interactive results. AVAILABILITY AND IMPLEMENTATION: The source code and test data are available at https://github.com/haidyi/cvxbiclustr or https://zenodo.org/record/4620218. The web server is available at https://cvxbiclustr.ericchi.com. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Haidong Yi, Gal Mishne, Eric C. Chi
Bioinform.3
2021 LDLE: Low Distortion Local Eigenmaps
abstract
We present Low Distortion Local Eigenmaps (LDLE), a manifold learning technique which constructs a set of low distortion local views of a data set in lower dimension and registers them to obtain a global embedding. The local views are constructed using the global eigenvectors of the graph Laplacian and are registered using Procrustes analysis. The choice of these eigenvectors may vary across the regions. In contrast to existing techniques, LDLE can embed closed and non-orientable manifolds into their intrinsic dimension by tearing them apart. It also provides gluing instruction on the boundary of the torn embedding to help identify the topology of the original manifold. Our experimental results will show that LDLE largely preserved distances up to a constant scale while other techniques produced higher distortion. We also demonstrate that LDLE produces high quality embeddings even when the data is noisy or sparse.
Dhruv Kohli, Alexander Cloninger, Gal Mishne
J. Mach. Learn. Res.3
2020 Poincaré Embedding Reveals Edge-Based Functional Networks of the Brain
Gal Mishne, Dustin Scheinost
MICCAI (7)2
2020 Spectral Embedding Norm: Looking Deep into the Spectrum of the Graph Laplacian
abstract
The extraction of clusters from a dataset which includes multiple clusters and a significant background component is a nontrivial task of practical importance. In image analysis this manifests for example in anomaly detection and target detection. The traditional spectral clustering algorithm, which relies on the leading $K$ eigenvectors to detect $K$ clusters, fails in such cases. In this paper we propose the spectral embedding norm which sums the squared values of the first $I$ normalized eigenvectors, where $I$ can be significantly larger than $K$. We prove that this quantity can be used to separate clusters from the background in unbalanced settings, including extreme cases such as outlier detection. The performance of the algorithm is not sensitive to the choice of $I$, and we demonstrate its application on synthetic and real-world remote sensing and neuroimaging datasets.
Xiuyuan Cheng, Gal Mishne
SIAM J. Imaging Sci.2
2019 Learning Spatially-correlated Temporal Dictionaries for Calcium Imaging
abstract
Calcium imaging has become a fundamental neural imaging technique, aiming to recover the individual activity of hundreds of neurons in a cortical region. Current methods (mostly matrix factorization) are aimed at detecting neurons in the field-of-view and then inferring the corresponding time-traces. In this paper, we reverse the modeling and instead aim to minimize the spatial inference, while focusing on finding the set of temporal traces present in the data. We reframe the problem in a dictionary learning setting, where the dictionary contains the time-traces and the sparse coefficient are spatial maps. We adapt dictionary learning to calcium imaging by introducing constraints on the norms and correlations of the time-traces, and incorporating a hierarchical spatial filtering model that correlates the time-trace usage over the field-of-view. We demonstrate on synthetic and real data that our solution has advantages regarding initialization, implicitly inferring number of neurons and simultaneously detecting different neuronal types.
Gal Mishne, Adam S. Charles
ICASSP1
2019 Co-manifold learning with missing data
abstract
Representation learning is typically applied to only one mode of a data matrix, either its rows or columns. Yet in many applications, there is an underlying geometry to both the rows and the columns. We propose utilizing this coupled structure to perform co-manifold learning: uncovering the underlying geometry of both the rows and the columns of a given matrix, where we focus on a missing data setting. Our unsupervised approach consists of three components. We first solve a family of optimization problems to estimate a complete matrix at multiple scales of smoothness. We then use this collection of smooth matrix estimates to compute pairwise distances on the rows and columns based on a new multi-scale metric that implicitly introduces a coupling between the rows and the columns. Finally, we construct row and column representations from these multi-scale metrics. We demonstrate that our approach outperforms competing methods in both data visualization and clustering.
Gal Mishne, Eric C. Chi, Ronald R. Coifman
ICML1
2019 Visualizing the PHATE of Neural Networks
abstract
Understanding why and how certain neural networks outperform others is key to guiding future development of network architectures and optimization methods. To this end, we introduce a novel visualization algorithm that reveals the internal geometry of such networks: Multislice PHATE (M-PHATE), the first method designed explicitly to visualize how a neural network's hidden representations of data evolve throughout the course of training. We demonstrate that our visualization provides intuitive, detailed summaries of the learning dynamics beyond simple global measures (i.e., validation loss and accuracy), without the need to access validation data. Furthermore, M-PHATE better captures both the dynamics and community structure of the hidden units as compared to visualization based on standard dimensionality reduction methods (e.g., ISOMAP, t-SNE). We demonstrate M-PHATE with two vignettes: continual learning and generalization. In the former, the M-PHATE visualizations display the mechanism of "catastrophic forgetting" which is a major challenge for learning in task-switching contexts. In the latter, our visualizations reveal how increased heterogeneity among hidden units correlates with improved generalization performance. An implementation of M-PHATE, along with scripts to reproduce the figures in this paper, is available at https://github.com/scottgigante/M-PHATE.
Scott Gigante, Adam S. Charles, Smita Krishnaswamy, Gal Mishne
NeurIPS4
2017 Iterative diffusion-based anomaly detection
abstract
Diffusion maps, when applied to large datasets, are typically constructed by a process of sampling and out-of-sample function extension. However, the performance of anomaly detection in large data when using diffusion maps is sensitive to the chosen samples. In this paper we propose an iterative data-driven approach to improve the sample set and diffusion maps representation. By updating the sample set with suspicious points detected in the previous iteration, the constructed diffusion maps better separate the anomaly from the normal points in each iteration. Experimental results in side-scan sonar images demonstrate the improvement gained by our iterative sampling compared to random sampling and other competing detection algorithms.
Gal Mishne, Israel Cohen
ICASSP1
2016 Improving resolution in supervised patch-based target detection
abstract
Recently, a supervised graph-based target detection method was proposed based on a new affinity measure between a set of target training patches and a test image. In this paper, we propose a new high-resolution detection score, which enhances the performance of the previous method by utilizing the known locations of the targets in the training images. We show that our new score is more reliable and spatially accurate, not only improving the detection resolution of true targets, but also reducing the number of false alarms. The method is successfully tested on side-scan sonar images of sea-mines, demonstrating an improved true detection rate. Our approach is general and can improve the detection resolution of the target in other patch-based detection algorithms for various signals and applications.
Ron Amit, Gal Mishne, Ronen Talmon
ICASSP2
2015 Graph-Based Supervised Automatic Target Detection
abstract
In this paper, we propose a detection method based on data-driven target modeling, which implicitly handles variations in the target appearance. Given a training set of images of the target, our approach constructs models based on local neighborhoods within the training set. We present a new metric using these models and show that, by controlling the notion of locality within the training set, this metric is invariant to perturbations in the appearance of the target. Using this metric in a supervised graph framework, we construct a low-dimensional embedding of test images. Then, a detection score based on the embedding determines the presence of a target in each image. The method is applied to a data set of side-scan sonar images and achieves impressive results in the detection of sea mines. The proposed framework is general and can be applied to different target detection problems in a broad range of signals.
Gal Mishne, Ronen Talmon, Israel Cohen
IEEE Trans. Geosci. Remote. Sens.1
2014 Multiscale anomaly detection using diffusion maps and saliency score
abstract
Recently, we presented a multiscale approach to anomaly detection in images, combining diffusion maps for dimensionality reduction and a nearest-neighbor-based anomaly score in the reduced dimension. When applying diffusion maps to images, usually a process of sampling and out-of-sample extension is used, which has limitations in regards to anomaly detection. To overcome the limitations, a multiscale approach was proposed, which drives the sampling process to ensure separability of the anomaly from the background clutter. In this paper, we propose a new anomaly score used in the diffusion map space, which shows increased performance. We show that this algorithm enables improved detection when tested on side-scan sonar images of sea-mines and compare it with competing algorithms.
Gal Mishne, Israel Cohen
ICASSP1