Pablo Sprechmann

dblp:72/4107 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Reinforcement learning · 67% Deep learning architectures and training · 14% Representation and self-supervised learning · 6%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 87% Medical and health informatics · 13%
Computer graphics and multimedia
2 papers
Image and video processing · 84% Audio and music processing · 16%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
deep reinforcement learning
0.822020
Agent57: Outperforming the Atari Human Benchmark · ICML 2020
Fast deep reinforcement learning using online adjustments from the past · NeurIPS 2018
Machine learning › Reinforcement learning › exploration › novelty-based exploration
count-based exploration
0.812024
Unlocking the Power of Representations in Long-term Novelty-based Exploration · ICLR 2024
Machine learning › Reinforcement learning
exploration
0.812024
Unlocking the Power of Representations in Long-term Novelty-based Exploration · ICLR 2024
Machine learning › Reinforcement learning › exploration
novelty-based exploration
0.812024
Unlocking the Power of Representations in Long-term Novelty-based Exploration · ICLR 2024
Machine learning › Reinforcement learning › exploration
directed exploration
0.412020
Never Give Up: Learning Directed Exploration Strategies · ICLR 2020
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.412020
Agent57: Outperforming the Atari Human Benchmark · ICML 2020
Machine learning › Reinforcement learning › exploration
exploration strategies
0.412020
Never Give Up: Learning Directed Exploration Strategies · ICLR 2020
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.412020
Never Give Up: Learning Directed Exploration Strategies · ICLR 2020
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay
0.312018
Fast deep reinforcement learning using online adjustments from the past · NeurIPS 2018
Robotics › Motion planning and robot control › robot control › adaptive control
parameter adaptation
0.312018
Memory-based Parameter Adaptation · ICLR (Poster) 2018
Machine learning › Reinforcement learning
value function estimation
0.312018
Fast deep reinforcement learning using online adjustments from the past · NeurIPS 2018
Machine learning › Deep learning architectures and training
convolutional neural network
0.312017
Accelerating Eulerian Fluid Simulation With Convolutional Networks · ICML 2017
Computational science and engineering
computational fluid dynamics
0.312017
Accelerating Eulerian Fluid Simulation With Convolutional Networks · ICML 2017
Machine learning › Deep learning architectures and training › autoencoder
adversarial autoencoder
0.212016
Disentangling factors of variation in deep representation using adversarial training · NIPS 2016
Machine learning › Deep learning architectures and training
autoencoder
0.212016
Disentangling factors of variation in deep representation using adversarial training · NIPS 2016
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.212016
Disentangling factors of variation in deep representation using adversarial training · NIPS 2016
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state representation
0.212024
Unlocking the Power of Representations in Long-term Novelty-based Exploration · ICLR 2024
Machine learning › Optimization for machine learning
bilevel optimization
0.212013
Supervised Sparse Analysis and Synthesis Operators · NIPS 2013
Machine learning › Deep learning architectures and training
feedforward neural network
0.212013
Supervised Sparse Analysis and Synthesis Operators · NIPS 2013
Machine learning › Graph learning
graph inference
0.212013
Robust Multimodal Graph Matching: Sparse Coding Meets Graph Matching · NIPS 2013
Machine learning › Graph learning
graph matching
0.212013
Robust Multimodal Graph Matching: Sparse Coding Meets Graph Matching · NIPS 2013
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding
0.212013
Supervised Sparse Analysis and Synthesis Operators · NIPS 2013
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning
0.112010
Classification and clustering via dictionary learning with structured incoherence and shared features · CVPR 2010
Data mining
clustering
0.112010
Classification and clustering via dictionary learning with structured incoherence and shared features · CVPR 2010
Data mining › clustering › high-dimensional clustering
subspace clustering
0.112010
Classification and clustering via dictionary learning with structured incoherence and shared features · CVPR 2010
Image and video processing
image segmentation
0.112007
Ultrasound Image Segmentation With Shape Priors: Application to Automatic Cattle Rib-Eye Area Estimation · IEEE Trans. Image Process. 2007
Image and video processing › image segmentation › deformable model segmentation
shape-prior segmentation
0.112007
Ultrasound Image Segmentation With Shape Priors: Application to Automatic Cattle Rib-Eye Area Estimation · IEEE Trans. Image Process. 2007
Image and video processing › image segmentation › medical image segmentation
ultrasound image segmentation
0.112007
Ultrasound Image Segmentation With Shape Priors: Application to Automatic Cattle Rib-Eye Area Estimation · IEEE Trans. Image Process. 2007
Image and video processing
image reconstruction
0.012013
Supervised Sparse Analysis and Synthesis Operators · NIPS 2013
Audio and music processing
music analysis
0.012013
Supervised Sparse Analysis and Synthesis Operators · NIPS 2013

Methods — techniques the papers use, named apart from their topics

masked transformer · 0.8inverse dynamics loss · 0.8clustering-based online density estimation · 0.8sparse coding · 0.6neural network policy parameterization · 0.4deep reinforcement learning · 0.4adaptive policy selection · 0.4prioritised sweeping · 0.3memory-based parameter adaptation · 0.3ephemeral value adjustments · 0.3unsupervised learning · 0.3operator splitting · 0.3convolutional network · 0.3convolutional operator learning · 0.2bilevel optimization · 0.2shape space · 0.1geodesic distance · 0.1curve fitting · 0.1
YearPublicationVenuePosition
2024 Unlocking the Power of Representations in Long-term Novelty-based Exploration
abstract
We introduce Robust Exploration via Clustering-based Online Density Estimation (RECODE), a non-parametric method for novelty-based exploration that estimates visitation counts for clusters of states based on their similarity in a chosen embedding space. By adapting classical clustering to the nonstationary setting of Deep RL, RECODE can efficiently track state visitation counts over thousands of episodes. We further propose a novel generalization of the inverse dynamics loss, which leverages masked transformer architectures for multi-step prediction; which in conjunction with \DETOCS achieves a new state-of-the-art in a suite of challenging 3D-exploration tasks in DM-Hard-8. RECODE also sets new state-of-the-art in hard exploration Atari games, and is the first agent to reach the end screen in "Pitfall!"
Alaa Saade, Steven Kapturowski, Daniele Calandriello, Charles Blundell, Pablo Sprechmann, Leopoldo Sarra, Oliver Groth, Michal Valko, Bilal Piot
ICLR5
2021 Game Plan: What AI can do for Football, and What Football can do for AI
abstract
The rapid progress in artificial intelligence (AI) and machine learning has opened unprecedented analytics possibilities in various team and individual sports, including baseball, basketball, and tennis. More recently, AI techniques have been applied to football, due to a huge increase in data collection by professional teams, increased computational power, and advances in machine learning, with the goal of better addressing new scientific challenges involved in the analysis of both individual players’ and coordinated teams’ behaviors. The research challenges associated with predictive and prescriptive football analytics require new developments and progress at the intersection of statistical learning, game theory, and computer vision. In this paper, we provide an overarching perspective highlighting how the combination of these fields, in particular, forms a unique microcosm for AI research, while offering mutual benefits for professional teams, spectators, and broadcasters in the years to come. We illustrate that this duality makes football analytics a game changer of tremendous value, in terms of not only changing the game of football itself, but also in terms of what this domain can mean for the field of AI. We review the state-of-the-art and exemplify the types of analysis enabled by combining the aforementioned fields, including illustrative examples of counterfactual analysis using predictive models, and the combination of game-theoretic analysis of penalty kicks with statistical learning of player attributes. We conclude by highlighting envisioned downstream impacts, including possibilities for extensions to other sports (real and virtual).
Karl Tuyls, Shayegan Omidshafiei, Paul Muller, Zhe Wang 0055, Jerome T. Connor, Daniel Hennes, Ian Graham, William Spearman, Tim Waskett, Dafydd Steele, Pauline Luc, Adrià Recasens, Alexandre Galashov, Gregory Thornton, Romuald Elie, Pablo Sprechmann, Pol Moreno, Kris Cao, Marta Garnelo, Praneet Dutta, Michal Valko, Nicolas Heess, Alex Bridgland, Julien Pérolat, Bart De Vylder, S. M. Ali Eslami, Mark Rowland 0001, Andrew Jaegle, Rémi Munos, Trevor Back, Razia Ahamed, Simon Bouton, Nathalie Beauguerlange, Jackson Broshear, Thore Graepel, Demis Hassabis
J. Artif. Intell. Res.16
2020 Never Give Up: Learning Directed Exploration Strategies
Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martín Arjovsky, Alexander Pritzel, Andrew Bolt, Charles Blundell
ICLR2
2020 Agent57: Outperforming the Atari Human Benchmark
abstract
Atari games have been a long-standing benchmark in the reinforcement learning (RL) community for the past decade. This benchmark was proposed to test general competency of RL algorithms. Previous work has achieved good average performance by doing outstandingly well on many games of the set, but very poorly in several of the most challenging games. We propose Agent57, the first deep RL agent that outperforms the standard human benchmark on all 57 Atari games. To achieve this result, we train a neural network which parameterizes a family of policies ranging from very exploratory to purely exploitative. We propose an adaptive mechanism to choose which policy to prioritize throughout the training process. Additionally, we utilize a novel parameterization of the architecture that allows for more consistent and stable learning.
Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Guo, Charles Blundell
ICML4
2018 Memory-based Parameter Adaptation
Pablo Sprechmann, Siddhant M. Jayakumar, Jack W. Rae, Alexander Pritzel, Adrià Puigdomènech Badia, Benigno Uria, Oriol Vinyals, Demis Hassabis, Razvan Pascanu, Charles Blundell
ICLR (Poster)1
2018 Fast deep reinforcement learning using online adjustments from the past
abstract
We propose Ephemeral Value Adjusments (EVA): a means of allowing deep reinforcement learning agents to rapidly adapt to experience in their replay buffer. EVA shifts the value predicted by a neural network with an estimate of the value function found by prioritised sweeping over experience tuples from the replay buffer near the current state. EVA combines a number of recent ideas around combining episodic memory-like structures into reinforcement learning agents: slot-based storage, content-based retrieval, and memory-based planning. We show that EVA is performant on a demonstration task and Atari games.
Steven Hansen 0001, Alexander Pritzel, Pablo Sprechmann, André Barreto 0001, Charles Blundell
NeurIPS3
2017 Accelerating Eulerian Fluid Simulation With Convolutional Networks
abstract
Efficient simulation of the Navier-Stokes equations for fluid flow is a long standing problem in applied mathematics, for which state-of-the-art methods require large compute resources. In this work, we propose a data-driven approach that leverages the approximation power of deep-learning with the precision of standard solvers to obtain fast and highly realistic simulations. Our method solves the incompressible Euler equations using the standard operator splitting method, in which a large sparse linear system with many free parameters must be solved. We use a Convolutional Network with a highly tailored architecture, trained using a novel unsupervised learning framework to solve the linear system. We present real-time 2D and 3D simulations that outperform recently proposed data-driven methods; the obtained results are realistic and show good generalization properties.
Jonathan Tompson, Kristofer Schlachter, Pablo Sprechmann, Ken Perlin
ICML3
2016 Disentangling factors of variation in deep representation using adversarial training
abstract
We propose a deep generative model for learning to distill the hidden factors of variation within a set of labeled observations into two complementary codes. One code describes the factors of variation relevant to solving a specified task. The other code describes the remaining factors of variation that are irrelevant to solving this task. The only available source of supervision during the training process comes from our ability to distinguish among different observations belonging to the same category. Concrete examples include multiple images of the same object from different viewpoints, or multiple speech samples from the same speaker. In both of these instances, the factors of variation irrelevant to classification are implicitly expressed by intra-class variabilities, such as the relative position of an object in an image, or the linguistic content of an utterance. Most existing approaches for solving this problem rely heavily on having access to pairs of observations only sharing a single factor of variation, e.g. different objects observed in the exact same conditions. This assumption is often not encountered in realistic settings where data acquisition is not controlled and labels for the uninformative components are not available. In this work, we propose to overcome this limitation by augmenting deep convolutional autoencoders with a form of adversarial training. Both factors of variation are implicitly captured in the organization of the learned embedding space, and can be used for solving single-image analogies. Experimental results on synthetic and real datasets show that the proposed method is capable of disentangling the influences of style and content factors using a flexible representation, as well as generalizing to unseen styles or content classes.
Michaël Mathieu, Junbo Jake Zhao, Pablo Sprechmann, Aditya Ramesh, Yann LeCun
NIPS3
2015 Source separation with scattering Non-Negative Matrix Factorization
abstract
This paper presents a single-channel source separation method that extends the ideas of Nonnegative Matrix Factorization (NMF). We interpret the approach of audio demixing via NMF as a cascade of a pooled analysis operator, given for example by the magnitude spectrogram, and a synthesis operators given by the matrix decomposition. Instead of imposing the temporal consistency of the decomposition through sophisticated structured penalties in the synthesis stage, we propose to change the analysis operator for a deep scattering representation, where signals are represented at several time resolutions. This new signal representation is invariant to smooth changes in the signal, consistent with its temporal dynamics. We evaluate the proposed approach in a speech separation task obtaining promising results.
Joan Bruna, Pablo Sprechmann, Yann LeCun
ICASSP2
2015 Multi-temporal foreground detection in videos
abstract
A common task in video processing is the binary separation of a video's content into either background or moving foreground. However, many situations require a foreground analysis with a finer temporal granularity, in particular for objects or people which remain immobile for a certain period of time. We propose an efficient method which detects foreground at different timescales, by exploiting the desirable theoretical and practical properties of Robust Principal Component Analysis. Our algorithm can be used in a variety of scenarios such as detecting people who have fallen in a video, or analysing the fluidity of road traffic, while avoiding costly computations needed for nearest neighbours searches or optical flow analysis. Finally, our algorithm has the useful ability to perform motion analysis without explicitly requiring computationally expensive motion estimation.
Mariano Tepper, Alasdair Newson, Pablo Sprechmann, Guillermo Sapiro
ICIP3
2015 Learning Efficient Sparse and Low Rank Models
abstract
Parsimony, including sparsity and low rank, has been shown to successfully model data in numerous machine learning and signal processing tasks. Traditionally, such modeling approaches rely on an iterative algorithm that minimizes an objective function with parsimony-promoting terms. The inherently sequential structure and data-dependent complexity and latency of iterative optimization constitute a major limitation in many applications requiring real-time performance or involving large-scale data. Another limitation encountered by these modeling techniques is the difficulty of their inclusion in discriminative learning scenarios. In this work, we propose to move the emphasis from the model to the pursuit algorithm, and develop a process-centric view of parsimonious modeling, in which a learned deterministic fixed-complexity pursuit process is used in lieu of iterative optimization. We show a principled way to construct learnable pursuit process architectures for structured sparse and robust low rank models, derived from the iteration of proximal descent algorithms. These architectures learn to approximate the exact parsimonious representation at a fraction of the complexity of the standard optimization methods. We also show that appropriate training regimes allow to naturally extend parsimonious models to discriminative settings. State-of-the-art results are demonstrated on several challenging problems in image and audio processing with several orders of magnitude speed-up compared to the exact optimization algorithms.
Pablo Sprechmann, Alexander M. Bronstein, Guillermo Sapiro
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Questionnaire simplification for fast risk analysis of children's mental health
abstract
Early detection and treatment of psychiatric disorders on children has shown significant impact in their subsequent development and quality of life. The assessment of psychopathology in childhood is commonly carried out by performing long comprehensive interviews such as the widely used Preschool Age Psychiatric Assessment (PAPA). Unfortunately, the time required to complete a full interview is too long to apply it at the scale of the actual population at risk, and most of the population goes undiagnosed or is diagnosed significantly later than desired. In this work, we aim to learn from unique and very rich previously collected PAPA examples the inter-correlations between different questions in order to provide a reliable risk analysis in the form of a much shorter interview. This helps to put such important risk analysis at the hands of regular practitioners, including teachers and family doctors. We use for this purpose the alternating decision trees algorithm, which combines decision trees with boosting to produce small and interpretable decision rules. Rather than a binary prediction, the algorithm provides a measure of confidence in the classification outcome. This is highly desirable from a clinical perspective, where it is preferable to abstain a decision on the low-confidence cases and recommend further screening. In order to prevent over-fitting, we propose to use network inference analysis to predefine a set of candidate question with consistent high correlation with the diagnosis. We report encouraging results with high levels of prediction using two independently collected datasets. The length and accuracy of the developed method suggests that it could be a valuable tool for preliminary evaluation in everyday care.
Kimberly L. H. Carpenter, Pablo Sprechmann, Marcelo Fiori, A. Robert Calderbank, Helen Link Egger, Guillermo Sapiro
ICASSP2
2013 Learnable low rank sparse models for speech denoising
abstract
In this paper we present a framework for real time enhancement of speech signals. Our method leverages a new process-centric approach for sparse and parsimonious models, where the representation pursuit is obtained applying a deterministic function or process rather than solving an optimization problem. We first propose a rank-regularized robust version of non-negative matrix factorization (NMF) for modeling time-frequency representations of speech signals in which the spectral frames are decomposed as sparse linear combinations of atoms of a low-rank dictionary. Then, a parametric family of pursuit processes is derived from the iteration of the proximal descent method for solving this model. We present several experiments showing successful results and the potential of the proposed framework. Incorporating discriminative learning makes the proposed method significantly outperform exact NMF algorithms, with fixed latency and at a fraction of it's computational complexity.
Pablo Sprechmann, Alexander M. Bronstein, Michael M. Bronstein, Guillermo Sapiro
ICASSP1
2013 Audio restoration from multiple copies
abstract
A method for removing impulse noise from audio signals by fusing multiple copies of the same recording is introduced in this paper. The proposed algorithm exploits the fact that while in general multiple copies of a given recording are available, all sharing the same master, most degradations in audio signals are record-dependent. Our method first seeks for the optimal non-rigid alignment of the signals that is robust to the presence of sparse outliers with arbitrary magnitude. Unlike previous approaches, we simultaneously find the optimal alignment of the signals and impulsive degradation. This is obtained via continuous dynamic time warping computed solving an Eikonal equation. We propose to use our approach in the derivative domain, reconstructing the signal by solving an inverse problem that resembles the Poisson image editing technique. The proposed framework is here illustrated and tested in the restoration of old gramophone recordings showing promising results; however, it can be used in other applications where different copies of the signal of interest are available and the degradations are copy-dependent.
Pablo Sprechmann, Alexander M. Bronstein, Jean-Michel Morel, Guillermo Sapiro
ICASSP1
2013 Robust Multimodal Graph Matching: Sparse Coding Meets Graph Matching
abstract
Graph matching is a challenging problem with very important applications in a wide range of fields, from image and video analysis to biological and biomedical problems. We propose a robust graph matching algorithm inspired in sparsity-related techniques. We cast the problem, resembling group or collaborative sparsity formulations, as a non-smooth convex optimization problem that can be efficiently solved using augmented Lagrangian techniques. The method can deal with weighted or unweighted graphs, as well as multimodal data, where different graphs represent different types of data. The proposed approach is also naturally integrated with collaborative graph inference techniques, solving general network inference problems where the observed variables, possibly coming from different modalities, are not in correspondence. The algorithm is tested and compared with state-of-the-art graph matching techniques in both synthetic and real graphs. We also present results on multimodal graphs and applications to collaborative inference of brain connectivity from alignment-free functional magnetic resonance imaging (fMRI) data.
Marcelo Fiori, Pablo Sprechmann, Joshua T. Vogelstein, Pablo Musé, Guillermo Sapiro
NIPS2
2013 Supervised Sparse Analysis and Synthesis Operators
abstract
In this paper, we propose a new and computationally efficient framework for learning sparse models. We formulate a unified approach that contains as particular cases models promoting sparse synthesis and analysis type of priors, and mixtures thereof. The supervised training of the proposed model is formulated as a bilevel optimization problem, in which the operators are optimized to achieve the best possible performance on a specific task, e.g., reconstruction or classification. By restricting the operators to be shift invariant, our approach can be thought as a way of learning analysis+synthesis sparsity-promoting convolutional operators. Leveraging recent ideas on fast trainable regressors designed to approximate exact sparse codes, we propose a way of constructing feed-forward neural networks capable of approximating the learned models at a fraction of the computational cost of exact solvers. In the shift-invariant case, this leads to a principled way of constructing task-specific convolutional networks. We illustrate the proposed models on several experiments in music analysis and image processing applications.
Pablo Sprechmann, Roee Litman, Tal Ben Yakar, Alexander M. Bronstein, Guillermo Sapiro
NIPS1
2013 Sparse Modeling of Intrinsic Correspondences
abstract
Abstract We present a novel sparse modeling approach to non‐rigid shape matching using only the ability to detect repeatable regions. As the input to our algorithm, we are given only two sets of regions in two shapes; no descriptors are provided so the correspondence between the regions is not know, nor we know how many regions correspond in the two shapes. We show that even with such scarce information, it is possible to establish very accurate correspondence between the shapes by using methods from the field of sparse modeling, being this, the first non‐trivial use of sparse models in shape correspondence. We formulate the problem ofpermuted sparse coding, in which we solve simultaneously for an unknown permutation ordering the regions on two shapes and for an unknown correspondence in functional representation. We also propose a robust variant capable of handling incomplete matches. Numerically, the problem is solved efficiently by alternating the solution of a linear assignment and a sparse coding problem. The proposed methods are evaluated qualitatively and quantitatively on standard benchmarks containing both synthetic and scanned objects.
Jonathan Pokrass, Alexander M. Bronstein, Michael M. Bronstein, Pablo Sprechmann, Guillermo Sapiro
Comput. Graph. Forum4
2012 Gaussian mixture models for score-informed instrument separation
abstract
A new framework for representing quasi-harmonic signals, and its application to score-informed single channel musical instruments separation, is introduced in this paper. In the proposed approach, the signal's pitch and spectral envelope are modeled separately. The model combines parametric filters enforcing an harmonic structure in the representation, with Gaussian modeling for representing the spectral envelope. The estimation of the signal's model is cast as an inverse problem efficiently solved via a maximum a posteriori expectation-maximization algorithm. The relation of the proposed framework with common non-negative factorization methods is also discussed. The algorithm is evaluated with both real and synthetic instruments mixtures, and comparisons with recently proposed techniques are presented.
Pablo Sprechmann, Pablo Cancela, Guillermo Sapiro
ICASSP1
2012 Learning Efficient Structured Sparse Models
Alexander M. Bronstein, Pablo Sprechmann, Guillermo Sapiro
ICML2
2011 Collaborative sources identification in mixed signals via hierarchical sparse modeling
abstract
A collaborative framework for detecting the different sources in mixed signals is presented in this paper. The approach is based on C-HiLasso, a convex collaborative hierarchical sparse model, and proceeds as follows. First, we build a structured dictionary for mixed signals by concatenating a set of sub-dictionaries, each one of them learned to sparsely model one of a set of possible classes. Then, the coding of the mixed signal is performed by efficiently solving a convex optimization problem that combines standard sparsity with group and collaborative sparsity. The present sources are identified by looking at the sub-dictionaries automatically selected in the coding. The collaborative filtering in C-HiLasso takes advantage of the temporal/spatial redundancy in the mixed signals, letting collections of samples collaborate in identifying the classes, while allowing in dividualsamples to have different internal sparse representations. This collaboration is critical to further stabilize the sparse representation of signals, in particular the class/sub-dictionary selection. The internal sparsity inside the sub-dictionaries, as naturally incorporated by the hierarchical aspects of C-HiLasso, is critical to make the model consistent with the essence of the sub-dictionaries that have been trained for sparse representation of each individual class. We present applications from speaker and instrument identification and texture separation. In the case of audio signals, we use sparse modeling to describe the short-term power spectrum envelopes of harmonic sounds. The proposed pitch independent method automatically detects the number of sources on a recording.
Pablo Sprechmann, Ignacio Ramírez, Pablo Cancela, Guillermo Sapiro
ICASSP1
2010 Classification and clustering via dictionary learning with structured incoherence and shared features
abstract
A clustering framework within the sparse modeling and dictionary learning setting is introduced in this work. Instead of searching for the set of centroid that best fit the data, as in k-means type of approaches that model the data as distributions around discrete points, we optimize for a set of dictionaries, one for each cluster, for which the signals are best reconstructed in a sparse coding manner. Thereby, we are modeling the data as a union of learned low dimensional subspaces, and data points associated to subspaces spanned by just a few atoms of the same learned dictionary are clustered together. An incoherence promoting term encourages dictionaries associated to different classes to be as independent as possible, while still allowing for different classes to share features. This term directly acts on the dictionaries, thereby being applicable both in the supervised and unsupervised settings. Using learned dictionaries for classification and clustering makes this method robust and well suited to handle large datasets. The proposed framework uses a novel measurement for the quality of the sparse representation, inspired by the robustness of the ℓ1regularization term in sparse coding. In the case of unsupervised classification and/or clustering, a new initialization based on combining sparse coding with spectral clustering is proposed. This initialization clusters the dictionary atoms, and therefore is based on solving a low dimensional eigen-decomposition problem, being applicable to large datasets. We first illustrate the proposed framework with examples on standard image and speech datasets in the supervised classification setting, obtaining results comparable to the state-of-the-art with this simple approach. We then present experiments for fully unsupervised clustering on extended standard datasets and texture images, obtaining excellent performance.
Ignacio Ramírez, Pablo Sprechmann, Guillermo Sapiro
CVPR2
2010 Dictionary learning and sparse coding for unsupervised clustering
abstract
A clustering framework within the sparse modeling and dictionary learning setting is introduced in this work. Instead of searching for the set of centroid that best fit the data, as in k-means type of approaches that model the data as distributions around discrete points, we optimize for a set of dictionaries, one for each cluster, for which the signals are best reconstructed in a sparse coding manner. Thereby, we are modeling the data as the of union of learned low dimensional subspaces, and data points associated to subspaces spanned by just a few atoms of the same learned dictionary are clustered together. Using learned dictionaries makes this method robust and well suited to handle large datasets. The proposed clustering algorithm uses a novel measurement for the quality of the sparse representation, inspired by the robustness of the ℓ1regularization term in sparse coding. We first illustrate this measurement with examples on standard image and speech datasets in the supervised classification setting, showing with a simple approach its discriminative power and obtaining results comparable to the state-of-the-art. We then conclude with experiments for fully unsupervised clustering on extended standard datasets and texture images, obtaining excellent performance.
Pablo Sprechmann, Guillermo Sapiro
ICASSP1
2007 Ultrasound Image Segmentation With Shape Priors: Application to Automatic Cattle Rib-Eye Area Estimation
abstract
Automatic ultrasound (US) image segmentation is a difficult task due to the quantity of noise present in the images and the lack of information in several zones produced by the acquisition conditions. In this paper, we propose a method that combines shape priors and image information to achieve this task. In particular, we introduce knowledge about the rib-eye shape using a set of images manually segmented by experts. A method is proposed for the automatic segmentation of new samples in which a closed curve is fitted taking into account both the US image information and the geodesic distance between the evolving curve and the estimated mean rib-eye shape in a shape space. This method can be used to solve similar problems that arise when dealing with US images in other fields. The method was successfully tested over a database composed of 610 US images, for which we have the manual segmentations of two experts.
Pablo Arias 0001, Alejandro Pini, Gonzalo Sanguinetti, Pablo Sprechmann, Pablo Cancela, Alicia Fernández, Alvaro Gómez, Gregory Randall
IEEE Trans. Image Process.4