EDBT 2026 Demo / reviewers in the wild / expert
Jeremias Sulam
dblp:156/3028
· DBLP profile ↗
30ranked-venue papers
7as first author
17since 2021 · last 2025
0000-0003-0946-1957ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Disentangling Safe and Unsafe Image Corruptions via Anisotropy and LocalityabstractState-of-the-art machine learning systems are vulnerable to small perturbations to their input, where "small" is defined according to a threat model that assigns a positive threat to each perturbation. Most prior works define a task-agnostic, isotropic, and global threat, like the ℓpnorm, where the magnitude of the perturbation fully determines the degree of the threat and neither the direction of the attack nor its position in space matter. However, common corruptions in computer vision, such as blur, compression, or occlusions, are not well captured by such threat models. This paper proposes a novel threat model called ProjectedDisplacement (PD) to study robustness beyond existing isotropic and global threat models. The proposed threat model measures the threat of a perturbation via its alignment with unsafe directions, defined as directions in the input space along which a perturbation of sufficient magnitude changes the ground truth class label. Unsafe directions are identified locally for each input based on observed training data. In this way, the PD-threat model exhibits anisotropy and locality. Experiments on Imagenet-1k data indicate that, for any input, the set of perturbations with small PD threat includes safe perturbations of large ℓpnorm that preserve the true label, such as noise, blur and compression, while simultaneously excluding unsafe perturbations that alter the true label. Unlike perceptual threat models based on embeddings of large-vision models, the PD-threat model can be readily computed for arbitrary classification tasks without pre-training or finetuning. Further additional task information such as sensitivity to image regions or concept hierarchies can be easily integrated into the assessment of threat and thus the PD threat model presents practitioners with a flexible, task-driven threat specification that alleviates the limitations of ℓp-threat models. Ramchandran Muthukumar, Ambar Pal, Jeremias Sulam, René Vidal |
CVPR | 3 |
| 2025 | Multiaccuracy and Multicalibration via Proxy GroupsabstractAs the use of predictive machine learning algorithms increases in high-stakes decision-making, it is imperative that these algorithms are fair across sensitive groups. However, measuring and enforcing fairness in real-world applications can be challenging due to missing or incomplete sensitive group information. Proxy-sensitive attributes have been proposed as a practical and effective solution in these settings, but only for parity-based fairness notions. Knowing how to evaluate and control for fairness with missing sensitive group data for newer, different, and more flexible frameworks, such as multiaccuracy and multicalibration, remains unexplored. In this work, we address this gap by demonstrating that in the absence of sensitive group data, proxy-sensitive attributes can provably be used to derive actionable upper bounds on the true multiaccuracy and multicalibration violations, providing insights into a predictive model’s potential worst-case fairness violations. Additionally, we show that adjusting models to satisfy multiaccuracy and multicalibration across proxy-sensitive attributes can significantly mitigate these violations for the true, but unknown, sensitive groups. Through several experiments on real-world datasets, we illustrate that approximate multiaccuracy and multicalibration can be achieved even when sensitive group data is incomplete or unavailable. Beepul Bharti, Mary Versa Clemens-Sewall, Paul H. Yi, Jeremias Sulam |
ICML | 4 |
| 2025 | Conformal Risk Control for Semantic Uncertainty Quantification in Computed Tomography
Jacopo Teneggi, J. Webster Stayman, Jeremias Sulam |
MICCAI (14) | 3 |
| 2025 | Beyond Scores: Proximal Diffusion ModelsabstractDiffusion models have quickly become some of the most popular and powerful generative models for high-dimensional data. The key insight that enabled their development was the realization that access to the score---the gradient of the log-density at different noise levels---allows for sampling from data distributions by solving a reverse-time stochastic differential equation (SDE) via forward discretization, and that popular denoisers allow for unbiased estimators of this score. In this paper, we demonstrate that an alternative, backward discretization of these SDEs, using proximal maps in place of the score, leads to theoretical and practical benefits. We leverage recent results in _proximal matching_ to learn proximal operators of the log-density and, with them, develop Proximal Diffusion Models (`ProxDM`). Theoretically, we prove that $\widetilde{\mathcal O}(d/\sqrt{\varepsilon})$ steps suffice for the resulting discretization to generate an $\varepsilon$-accurate distribution w.r.t. the KL divergence.
Empirically, we show that two variants of `ProxDM` achieve significantly faster convergence within just a few sampling steps compared to conventional score-matching methods. Zhenghan Fang, Mateo Díaz, Sam Buchanan, Jeremias Sulam |
NeurIPS | 4 |
| 2025 | Fourier Diffusion Models: A Method to Control MTF and NPS in Score-Based Stochastic Image GenerationabstractScore-based diffusion models are new and powerful tools for image generation. They are based on a forward stochastic process where an image is degraded with additive white noise and optional input scaling. A neural network can be trained to estimate the time-dependent score function, and used to run the reverse-time stochastic process to generate new samples from the training image distribution. However, one issue is that sampling the reverse process requires many passes of the neural network. In this work we present Fourier Diffusion Models which replace the scalar operations of the forward process with linear shift invariant systems and additive spatially-stationary noise. This allows for a model of continuous probability flow from true images to measurements with a specific modulation transfer function (MTF) and noise power spectrum (NPS). We also derive the reverse process for posterior sampling of high-quality images given blurry noisy measurements. We conducted a computational experiment using the Lung Image Database Consortium dataset of chest CT images and simulated CT measurements with correlated noise and system blur. Our results show that Fourier diffusion models can improve image quality for supervised diffusion posterior sampling relative to existing conditional diffusion models. Matthew Tivnan, Jacopo Teneggi, Tzu-Cheng Lee, Ruoqiao Zhang, Kirsten Boedeker, Grace J. Gang, Jeremias Sulam, J. Webster Stayman |
IEEE Trans. Medical Imaging | 8 |
| 2024 | What's in a Prior? Learned Proximal Networks for Inverse ProblemsabstractProximal operators are ubiquitous in inverse problems, commonly appearing as part of algorithmic strategies to regularize problems that are otherwise ill-posed. Modern deep learning models have been brought to bear for these tasks too, as in the framework of plug-and-play or deep unrolling, where they loosely resemble proximal operators. Yet, something essential is lost in employing these purely data-driven approaches: there is no guarantee that a general deep network represents the proximal operator of any function, nor is there any characterization of the function for which the network might provide some approximate proximal. This not only makes guaranteeing convergence of iterative schemes challenging but, more fundamentally, complicates the analysis of what has been learned by these networks about their training data. Herein we provide a framework to develop *learned proximal networks* (LPN), prove that they provide exact proximal operators for a data-driven nonconvex regularizer, and show how a new training strategy, dubbed *proximal matching*, provably promotes the recovery of the log-prior of the true data distribution. Such LPN provide general, unsupervised, expressive proximal operators that can be used for general inverse problems with convergence guarantees. We illustrate our results in a series of cases of increasing complexity, demonstrating that these models not only result in state-of-the-art performance, but provide a window into the resulting priors learned from data. Zhenghan Fang, Sam Buchanan, Jeremias Sulam |
ICLR | 3 |
| 2024 | Testing Semantic Importance via BettingabstractRecent works have extended notions of feature importance to semantic concepts that are inherently interpretable to the users interacting with a black-box predictive model. Yet, precise statistical guarantees such as false positive rate and false discovery rate control are needed to communicate findings transparently, and to avoid unintended consequences in real-world scenarios. In this paper, we formalize the global (i.e., over a population) and local (i.e., for a sample) statistical importance of semantic concepts for the predictions of opaque models by means of conditional independence, which allows for rigorous testing. We use recent ideas of sequential kernelized independence testing to induce a rank of importance across concepts, and we showcase the effectiveness and flexibility of our framework on synthetic datasets as well as on image classification using several vision-language models. Jacopo Teneggi, Jeremias Sulam |
NeurIPS | 2 |
| 2023 | Sparsity-aware generalization theory for deep neural networksabstractDeep artificial neural networks achieve surprising generalization abilities that remain poorly understood. In this paper, we present a new approach to analyzing generalization for deep feed-forward ReLU networks that takes advantage of the degree of sparsity that is achieved in the hidden layer activations. By developing a framework that accounts for this reduced effective model size for each input sample, we are able to show fundamental trade-offs between sparsity and generalization. Importantly, our results make no strong assumptions about the degree of sparsity achieved by the model, and it improves over recent norm-based approaches. We illustrate our results numerically, demonstrating non-vacuous bounds when coupled with data-dependent priors even in over-parametrized settings. Ramchandran Muthukumar, Jeremias Sulam |
COLT | 2 |
| 2023 | How to Trust Your Diffusion Model: A Convex Optimization Approach to Conformal Risk ControlabstractScore-based generative modeling, informally referred to as diffusion models, continue to grow in popularity across several important domains and tasks. While they provide high-quality and diverse samples from empirical distributions, important questions remain on the reliability and trustworthiness of these sampling procedures for their responsible use in critical scenarios. Conformal prediction is a modern tool to construct finite-sample, distribution-free uncertainty guarantees for any black-box predictor. In this work, we focus on image-to-image regression tasks and we present a generalization of the Risk-Controlling Prediction Sets (RCPS) procedure, that we term $K$-RCPS, which allows to $(i)$ provide entrywise calibrated intervals for future samples of any diffusion model, and $(ii)$ control a certain notion of risk with respect to a ground truth image with minimal mean interval length. Differently from existing conformal risk control procedures, ours relies on a novel convex optimization approach that allows for multidimensional risk control while provably minimizing the mean interval length. We illustrate our approach on two real-world image denoising problems: on natural images of faces as well as on computed tomography (CT) scans of the abdomen, demonstrating state of the art performance. Jacopo Teneggi, Matthew Tivnan, J. Webster Stayman, Jeremias Sulam |
ICML | 4 |
| 2023 | Estimating and Controlling for Equalized Odds via Sensitive Attribute PredictorsabstractAs the use of machine learning models in real world high-stakes decision settings continues to grow, it is highly important that we are able to audit and control for any potential fairness violations these models may exhibit towards certain groups. To do so, one naturally requires access to sensitive attributes, such as demographics, biological sex, or other potentially sensitive features that determine group membership. Unfortunately, in many settings, this information is often unavailable. In this work we study the well known equalized odds (EOD) definition of fairness. In a setting without sensitive attributes, we first provide tight and computable upper bounds for the EOD violation of a predictor. These bounds precisely reflect the worst possible EOD violation. Second, we demonstrate how one can provably control the worst-case EOD by a new post-processing correction method. Our results characterize when directly controlling for EOD with respect to the predicted sensitive attributes is -- and when is not -- optimal when it comes to controlling worst-case EOD. Our results hold under assumptions that are milder than previous works, and we illustrate these results with experiments on synthetic and real datasets. Beepul Bharti, Paul H. Yi, Jeremias Sulam |
NeurIPS | 3 |
| 2023 | Adversarial Examples Might be Avoidable: The Role of Data Concentration in Adversarial RobustnessabstractThe susceptibility of modern machine learning classifiers to adversarial examples has motivated theoretical results suggesting that these might be unavoidable. However, these results can be too general to be applicable to natural data distributions. Indeed, humans are quite robust for tasks involving vision. This apparent conflict motivates a deeper dive into the question: Are adversarial examples truly unavoidable?
In this work, we theoretically demonstrate that a key property of the data distribution -- concentration on small-volume subsets of the input space -- determines whether a robust classifier exists. We further demonstrate that, for a data distribution concentrated on a union of low-dimensional linear subspaces, utilizing structure in data naturally leads to classifiers that enjoy data-dependent polyhedral robustness guarantees, improving upon methods for provable certification in certain regimes. Ambar Pal, Jeremias Sulam, René Vidal |
NeurIPS | 2 |
| 2023 | DeepSTI: Towards tensor reconstruction using fewer orientations in susceptibility tensor imagingabstractSusceptibility tensor imaging (STI) is an emerging magnetic resonance imaging technique that characterizes the anisotropic tissue magnetic susceptibility with a second-order tensor model. STI has the potential to provide information for both the reconstruction of white matter fiber pathways and detection of myelin changes in the brain at mm resolution or less, which would be of great value for understanding brain structure and function in healthy and diseased brain. However, the application of STI in vivo has been hindered by its cumbersome and time-consuming acquisition requirement of measuring susceptibility induced MR phase changes at multiple head orientations. Usually, sampling at more than six orientations is required to obtain sufficient information for the ill-posed STI dipole inversion. This complexity is enhanced by the limitation in head rotation angles due to physical constraints of the head coil. As a result, STI has not yet been widely applied in human studies in vivo. In this work, we tackle these issues by proposing an image reconstruction algorithm for STI that leverages data-driven priors. Our method, called DeepSTI, learns the data prior implicitly via a deep neural network that approximates the proximal operator of a regularizer function for STI. The dipole inversion problem is then solved iteratively using the learned proximal network. Experimental results using both simulation and in vivo human data demonstrate great improvement over state-of-the-art algorithms in terms of the reconstructed tensor image, principal eigenvector maps and tractography results, while allowing for tensor reconstruction with MR phase measured at much less than six different orientations. Notably, promising reconstruction results are achieved by our method from only one orientation in human in vivo, and we demonstrate a potential application of this technique for estimating lesion susceptibility anisotropy in patients with multiple sclerosis. Zhenghan Fang, Kuo-Wei Lai, Peter C. M. van Zijl, Xu Li 0003, Jeremias Sulam |
Medical Image Anal. | 5 |
| 2023 | Fast Hierarchical Games for Image ExplanationsabstractAs modern complex neural networks keep breaking records and solving harder problems, their predictions also become less and less intelligible. The current lack of interpretability often undermines the deployment of accurate machine learning tools in sensitive settings. In this work, we present a model-agnostic explanation method for image classification based on a hierarchical extension of Shapley coefficients-Hierarchical Shap (h-Shap)-that resolves some of the limitations of current approaches. Unlike other Shapley-based explanation methods, h-Shap is scalable and can be computed without the need of approximation. Under certain distributional assumptions, such as those common in multiple instance learning, h-Shap retrieves the exact Shapley coefficients with an exponential improvement in computational complexity. We compare our hierarchical approach with popular Shapley-based and non-Shapley-based methods on a synthetic dataset, a medical imaging scenario, and a general computer vision problem, showing that h-Shap outperforms the state-of-the-art in both accuracy and runtime. Code and experiments are made publicly available. Jacopo Teneggi, Alexandre Luster, Jeremias Sulam |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Entrywise Recovery Guarantees for Sparse PCA via Sparsistent AlgorithmsabstractSparse Principal Component Analysis (PCA) is a prevalent tool across a plethora of subfield of applied statistics. While several results have characterized the recovery error of the principal eigenvectors, these are typically in spectral or Frobenius norms. In this paper, we provide entrywise $\ell_{2,\infty}$ bounds for Sparse PCA under a general high-dimensional subgaussian design. In particular, our bounds hold for any algorithm that selects the correct support with high probability, those that are sparsistent. Our bound improves upon known results by providing a finer characterization of the estimation error, and our proof uses techniques recently developed for entrywise subspace perturbation theory. Joshua Agterberg, Jeremias Sulam |
AISTATS | 2 |
| 2022 | Recovery and Generalization in Over-Realized Dictionary LearningabstractIn over two decades of research, the field of dictionary learning has gathered a large collection of successful applications, and theoretical guarantees for model recovery are known only whenever optimization is carried out in the same model class as that of the underlying dictionary. This work characterizes the surprising phenomenon that dictionary recovery can be facilitated by searching over the space of larger over-realized models. This observation is general and independent of the specific dictionary learning algorithm used. We thoroughly demonstrate this observation in practice and provide an analysis of this phenomenon by tying recovery measures to generalization bounds. In particular, we show that model recovery can be upper-bounded by the empirical risk, a model-dependent quantity and the generalization gap, reflecting our empirical findings. We further show that an efficient and provably correct distillation approach can be employed to recover the correct atoms from the over-realized model. As a result, our meta-algorithm provides dictionary estimates with consistently better recovery of the ground-truth model. Jeremias Sulam, Chong You, Zhihui Zhu |
J. Mach. Learn. Res. | 1 |
| 2022 | Label Cleaning Multiple Instance Learning: Refining Coarse Annotations on Single Whole-Slide ImagesabstractAnnotating cancerous regions in whole-slide images (WSIs) of pathology samples plays a critical role in clinical diagnosis, biomedical research, and machine learning algorithms development. However, generating exhaustive and accurate annotations is labor-intensive, challenging, and costly. Drawing only coarse and approximate annotations is a much easier task, less costly, and it alleviates pathologists' workload. In this paper, we study the problem of refining these approximate annotations in digital pathology to obtain more accurate ones. Some previous works have explored obtaining machine learning models from these inaccurate annotations, but few of them tackle the refinement problem where the mislabeled regions should be explicitly identified and corrected, and all of them require a - often very large - number of training samples. We present a method, named Label Cleaning Multiple Instance Learning (LC-MIL), to refine coarse annotations on a single WSI without the need for external training data. Patches cropped from a WSI with inaccurate labels are processed jointly within a multiple instance learning framework, mitigating their impact on the predictive model and refining the segmentation. Our experiments on a heterogeneous WSI set with breast cancer lymph node metastasis, liver cancer, and colorectal cancer samples show that LC-MIL significantly refines the coarse annotations, outperforming state-of-the-art alternatives, even while learning from a single slide. Moreover, we demonstrate how real annotations drawn by pathologists can be efficiently refined and improved by the proposed approach. All these results demonstrate that LC-MIL is a promising, lightweight tool to provide fine-grained annotations from coarsely annotated pathology sets. Carla Saoud, Sintawat Wangsiricharoen, Aaron W. James, Aleksander S. Popel, Jeremias Sulam |
IEEE Trans. Medical Imaging | 6 |
| 2021 | A Geometric Analysis of Neural Collapse with Unconstrained FeaturesabstractWe provide the first global optimization landscape analysis of Neural Collapse -- an intriguing empirical phenomenon that arises in the last-layer classifiers and features of neural networks during the terminal phase of training. As recently reported by Papyan et al., this phenomenon implies that (i) the class means and the last-layer classifiers all collapse to the vertices of a Simplex Equiangular Tight Frame (ETF) up to scaling, and (ii) cross-example within-class variability of last-layer activations collapses to zero. We study the problem based on a simplified unconstrained feature model, which isolates the topmost layers from the classifier of the neural network. In this context, we show that the classical cross-entropy loss with weight decay has a benign global landscape, in the sense that the only global minimizers are the Simplex ETFs while all other critical points are strict saddles whose Hessian exhibit negative curvature directions. Our analysis of the simplified model not only explains what kind of features are learned in the last layer, but also shows why they can be efficiently optimized, matching the empirical observations in practical deep network architectures. These findings provide important practical implications. As an example, our experiments demonstrate that one may set the feature dimension equal to the number of classes and fix the last-layer classifier to be a Simplex ETF for network training, which reduces memory cost by over 20% on ResNet18 without sacrificing the generalization performance. The source code is available at https://github.com/tding1/Neural-Collapse. Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li 0026, Chong You, Jeremias Sulam, Qing Qu 0001 |
NeurIPS | 6 |
| 2020 | Learned Proximal Networks for Quantitative Susceptibility Mapping
Kuo-Wei Lai, Manisha Aggarwal, Peter C. M. van Zijl, Xu Li 0003, Jeremias Sulam |
MICCAI (2) | 5 |
| 2020 | Learning to solve TV regularised problems with unrolled algorithmsabstractTotal Variation (TV) is a popular regularization strategy that promotes piece-wise constant signals by constraining the ℓ1-norm of the first order derivative of the estimated signal. The resulting optimization problem is usually solved using iterative algorithms such as proximal gradient descent, primal-dual algorithms or ADMM. However, such methods can require a very large number of iterations to converge to a suitable solution. In this paper, we accelerate such iterative algorithms by unfolding proximal gradient descent solvers in order to learn their parameters for 1D TV regularized problems. While this could be done using the synthesis formulation, we demonstrate that this leads to slower performances. The main difficulty in applying such methods in the analysis formulation lies in proposing a way to compute the derivatives through the proximal operator. As our main contribution, we develop and characterize two approaches to do so, describe their benefits and limitations, and discuss the regime where they can actually improve over iterative procedures. We validate those findings with experiments on synthetic and real data. Hamza Cherkaoui, Jeremias Sulam, Thomas Moreau 0001 |
NeurIPS | 2 |
| 2020 | Conformal Symplectic and Relativistic OptimizationabstractArguably, the two most popular accelerated or momentum-based optimization methods are Nesterov's accelerated gradient and Polyaks's heavy ball, both corresponding to different discretizations of a particular second order differential equation with a friction term. Such connections with continuous-time dynamical systems have been instrumental in demystifying acceleration phenomena in optimization. Here we study structure-preserving discretizations for a certain class of dissipative (conformal) Hamiltonian systems, allowing us to analyze the symplectic structure of both Nesterov and heavy ball, besides providing several new insights into these methods. Moreover, we propose a new algorithm based on a dissipative relativistic system that normalizes the momentum and may result in more stable/faster optimization. Importantly, such a method generalizes both Nesterov and heavy ball, each being recovered as distinct limiting cases, and has potential advantages at no additional cost. Guilherme França, Jeremias Sulam, Daniel P. Robinson, René Vidal |
NeurIPS | 2 |
| 2020 | Adversarial Robustness of Supervised Sparse CodingabstractSeveral recent results provide theoretical insights into the phenomena of adversarial examples. Existing results, however, are often limited due to a gap between the simplicity of the models studied and the complexity of those deployed in practice. In this work, we strike a better balance by considering a model that involves learning a representation while at the same time giving a precise generalization bound and a robustness certificate. We focus on the hypothesis class obtained by combining a sparsity-promoting encoder coupled with a linear classifier, and show an interesting interplay between the expressivity and stability of the (supervised) representation map and a notion of margin in the feature space. We bound the robust risk (to $\ell_2$-bounded perturbations) of hypotheses parameterized by dictionaries that achieve a mild encoder gap on training data. Furthermore, we provide a robustness certificate for end-to-end classification. We demonstrate the applicability of our analysis by computing certified accuracy on real data, and compare with other alternatives for certified robustness. Jeremias Sulam, Ramchandran Muthukumar, Raman Arora |
NeurIPS | 1 |
| 2020 | Geometric potentials from deep learning improve prediction of CDR H3 loop structuresabstractMOTIVATION: Antibody structure is largely conserved, except for a complementarity-determining region featuring six variable loops. Five of these loops adopt canonical folds which can typically be predicted with existing methods, while the remaining loop (CDR H3) remains a challenge due to its highly diverse set of observed conformations. In recent years, deep neural networks have proven to be effective at capturing the complex patterns of protein structure. This work proposes DeepH3, a deep residual neural network that learns to predict inter-residue distances and orientations from antibody heavy and light chain sequence. The output of DeepH3 is a set of probability distributions over distances and orientation angles between pairs of residues. These distributions are converted to geometric potentials and used to discriminate between decoy structures produced by RosettaAntibody and predict new CDR H3 loop structures de novo. RESULTS: When evaluated on the Rosetta antibody benchmark dataset of 49 targets, DeepH3-predicted potentials identified better, same and worse structures [measured by root-mean-squared distance (RMSD) from the experimental CDR H3 loop structure] than the standard Rosetta energy function for 33, 6 and 10 targets, respectively, and improved the average RMSD of predictions by 32.1% (1.4 Å). Analysis of individual geometric potentials revealed that inter-residue orientations were more effective than inter-residue distances for discriminating near-native CDR H3 loops. When applied to de novo prediction of CDR H3 loop structures, DeepH3 achieves an average RMSD of 2.2 ± 1.1 Å on the Rosetta antibody benchmark. AVAILABILITY AND IMPLEMENTATION: DeepH3 source code and pre-trained model parameters are freely available at https://github.com/Graylab/deepH3-distances-orientations. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jeffrey A. Ruffolo, Sai Pooja Mahajan, Jeremias Sulam, Jeffrey J. Gray |
Bioinform. | 4 |
| 2020 | On Multi-Layer Basis Pursuit, Efficient Algorithms and Convolutional Neural NetworksabstractParsimonious representations are ubiquitous in modeling and processing information. Motivated by the recent Multi-Layer Convolutional Sparse Coding (ML-CSC) model, we herein generalize the traditional Basis Pursuit problem to a multi-layer setting, introducing similar sparse enforcing penalties at different representation layers in a symbiotic relation between synthesis and analysis sparse priors. We explore different iterative methods to solve this new problem in practice, and we propose a new Multi-Layer Iterative Soft Thresholding Algorithm (ML-ISTA), as well as a fast version (ML-FISTA). We show that these nested first order algorithms converge, in the sense that the function value of near-fixed points can get arbitrarily close to the solution of the original problem. We further show how these algorithms effectively implement particular recurrent convolutional neural networks (CNNs) that generalize feed-forward ones without introducing any parameters. We present and analyze different architectures resulting from unfolding the iterations of the proposed pursuit algorithms, including a new Learned ML-ISTA, providing a principled way to construct deep recurrent CNNs. Unlike other similar constructions, these architectures unfold a global pursuit holistically for the entire network. We demonstrate the emerging constructions in a supervised learning setting, consistently improving the performance of classical CNNs while maintaining the number of parameters constant. Jeremias Sulam, Aviad Aberdam, Amir Beck, Michael Elad |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | A Local Block Coordinate Descent Algorithm for the CSC ModelabstractThe Convolutional Sparse Coding (CSC) model has recently gained considerable traction in the signal and image processing communities. By providing a global, yet tractable, model that operates on the whole image, the CSC was shown to overcome several limitations of the patch-based sparse model while achieving superior performance in various applications. Contemporary methods for pursuit and learning the CSC dictionary often rely on the Alternating Direction Method of Multipliers (ADMM) in the Fourier domain for the computational convenience of convolutions, while ignoring the local characterizations of the image. In this work we propose a new and simple approach that adopts a localized strategy, based on the Block Coordinate Descent algorithm. The proposed method, termed Local Block Coordinate Descent (LoBCoD), operates locally on image patches. Furthermore, we introduce a novel stochastic gradient descent version of LoBCoD for training the convolutional filters. This Stochastic-LoBCoD leverages the benefits of online learning, while being applicable even to a single training image. We demonstrate the advantages of the proposed algorithms for image inpainting and multi-focus image fusion, achieving state-of-the-art results. Ev Zisselman, Jeremias Sulam, Michael Elad |
CVPR | 2 |
| 2018 | Projecting on to the Multi-Layer Convolutional Sparse Coding ModelabstractThe recently proposed Multi-Layer Convolutional Sparse Coding (ML-CSC) model, consisting of a cascade of convolutional sparse layers, provides a new interpretation of Convolutional Neural Networks (CNNs). Under this framework, the forward pass in a CNN is equivalent to an algorithm that estimates nested sparse representation vectors from a given input signal. Despite having served as a pivotal connection between CNNs and sparse modeling, it is still unclear how to develop pursuit algorithms that serve this model exactly. In this work, we propose a new pursuit formulation by adopting a projection approach. We provide new and improved bounds on the stability of the resulting convolutional sparse representations, and we propose a multi-layer projection algorithm to retrieve them. We demonstrate this algorithm numerically, showing that it is superior to the Layered Basis Pursuit alternative in retrieving the representations of signals belonging to the ML-CSC model. Jeremias Sulam, Vardan Papyan, Yaniv Romano, Michael Elad |
ICASSP | 1 |
| 2017 | Convolutional Dictionary Learning via Local ProcessingabstractConvolutional sparse coding is an increasingly popular model in the signal and image processing communities, tackling some of the limitations of traditional patch-based sparse representations. Although several works have addressed the dictionary learning problem under this model, these relied on an ADMM formulation in the Fourier domain, losing the sense of locality and the relation to the traditional patch-based sparse pursuit. A recent work suggested a novel theoretical analysis of this global model, providing guarantees that rely on a localized sparsity measure. Herein, we extend this local-global relation by showing how one can efficiently solve the convolutional sparse pursuit problem and train the filters involved, while operating locally on image patches. Our approach provides an intuitive algorithm that can leverage standard techniques from the sparse representations field. The proposed method is fast to train, simple to implement, and flexible enough that it can be easily deployed in a variety of applications. We demonstrate the proposed training scheme for image inpainting and image separation, achieving state-of-the-art results. Vardan Papyan, Yaniv Romano, Michael Elad, Jeremias Sulam |
ICCV | 4 |
| 2017 | Dynamical system classification with diffusion embedding for ECG-based person identification
Jeremias Sulam, Yaniv Romano, Ronen Talmon |
Signal Process. | 1 |
| 2016 | Large Inpainting of Face Images With TrainletsabstractImage inpainting is concerned with the completion of missing data in an image. When the area to inpaint is relatively large, this problem becomes challenging. In these cases, traditional methods based on patch models and image propagation are limited, since they fail to consider a global perspective of the problem. In this letter, we employ a recently proposed dictionary learning framework, coined Trainlets, to design large adaptable atoms from a corpus of various datasets of face images by leveraging the online sparse dictionary learning algorithm. We, therefore, formulate the inpainting task as an inverse problem with a sparse-promoting prior based on the learned global model. Our results show the effectiveness of our scheme, obtaining much more plausible results than competitive methods. Jeremias Sulam, Michael Elad |
IEEE Signal Process. Lett. | 1 |
| 2015 | Fusion of ultrasound harmonic imaging with clutter removal using sparse signal separationabstractIn ultrasound, second harmonic imaging is usually preferred due to the higher clutter artifacts and speckle noise common in the first harmonic image. Typical ultrasound use either one or the other image, applying corresponding filters for each case. In this work we propose a method based on a joint sparsity model that fuses the first and second harmonic images while performing clutter mitigation and noise reduction. Our approach, Fused Morphological Component Analysis (FMCA), uses two adaptive dictionaries for characterizing the clutter components in each image, and a common dictionary for the tissue representation. Our results indicate that the obtained images contain less clutter artifacts, less speckle noise and as such enjoy of the benefits of both harmonic input images. Javier Turek, Jeremias Sulam, Michael Elad, Irad Yavneh |
ICASSP | 2 |
| 2014 | Image denoising through multi-scale learnt dictionariesabstractOver the last decade, a number of algorithms have shown promising results in removing additive white Gaussian noise from natural images, and though different, they all share in common a patch based strategy by locally denoising overlapping patches. While this lowers the complexity of the problem, it also causes noticeable artifacts when dealing with large smooth areas. In this paper we present a patch-based denoising algorithm relying on a sparsity-inspired model (K-SVD), which uses a multi-scale analysis framework. This allows us to overcome some of the disadvantages of the popular algorithms. We look for a sparse representation under an already sparsifying wavelet transform by adaptively training a dictionary on the different decomposition bands of the noisy image itself, leading to a multi-scale version of the K-SVD algorithm. We then combine the single scale and multi-scale approaches by merging both outputs by weighted joint sparse coding of the images. Our experiments on natural images indicate that our method is competitive with state of the art algorithms in terms of PSNR while giving superior results with respect to visual quality. Jeremias Sulam, Boaz Ophir, Michael Elad |
ICIP | 1 |