VLDB 2026 Research / reviewers in the wild / expert
Ruud van Sloun
dblp:162/9715 · also Ruud J. G. van Sloun
· DBLP profile ↗
60ranked-venue papers
6as first author
43since 2021 · last 2026
0000-0003-2845-0495ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 1 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SVD-NO: Learning PDE Solution Operators with SVD Integral KernelsabstractNeural operators have emerged as a promising paradigm for learning solution operators of partial differential equations (PDEs) directly from data. Existing methods, such as those based on Fourier or graph techniques, make strong assumptions about the structure of the kernel integral operator, assumptions which may limit expressivity. We present SVD-NO, a neural operator that explicitly parameterizes the kernel by its singular-value decomposition (SVD) and then carries out the integral directly in the low-rank basis. Two lightweight networks learn the left and right singular functions, a diagonal parameter matrix learns the singular values, and a Gram-matrix regularizer enforces orthonormality. As SVD-NO approximates the full kernel, it obtains a high degree of expressivity. Furthermore, due to its low-rank structure the computational complexity of applying the operator remains reasonable, leading to a practical system. In extensive evaluations on five diverse benchmark equations, SVD-NO achieves a new state of the art. In particular, SVD-NO provides greater performance gains on PDEs whose solutions are highly spatially variable. Noam Koren, Ralf J. J. Mackenbach, Ruud van Sloun, Kira Radinsky, Daniel Freedman |
AAAI | 3 |
| 2026 | Patient-Adaptive Echocardiography Using Cognitive UltrasoundabstractFocused transmits are the most commonly used transmit strategy for echocardiograms, but suffer from relatively low frame rates, and in 3D, even lower volume rates. Fast imaging based on unfocused transmits has disadvantages such as motion decorrelation and limited harmonic imaging capabilities. This work introduces a patient-adaptive focused transmit and receive scheme that has the ability to drastically reduce the number of transmits needed to produce a high-quality ultrasound image. The method relies on posterior sampling with a temporal diffusion model to perceive and reconstruct the anatomy based on partial observations, while subsequently acquiring the most informative transmits. This cognitive ultrasound modality outperforms random and equispaced subsampling in terms of distortion and perceptual metrics on the 2D EchoNet-Dynamic dataset and a 3D Philips dataset, where we actively select focused elevation planes. Furthermore, our method improves generalized contrast-to-noise ratio from 0.83 to 0.89 compared to the same number of diverging wave transmits on six in-house echocardiograms. Additionally, we can segment the left ventricle, with on average 0.91 Dice-Sørensen coefficient, through simulating using 2 out of 112 lines. Finally, our method can be run in real-time on GPU accelerators from 2023, increasing the maximum achievable frame-rate from 46 Hz to 58 Hz. The code is publicly available at https://tue-bmd.github.io/casl/. Wessel L. van Nierop, Oisín Nolan, Tristan S. W. Stevens, Ruud van Sloun |
IEEE Trans. Medical Imaging | 4 |
| 2026 | High Volume Rate 3-D Ultrasound Reconstruction With Diffusion ModelsabstractThree-dimensional ultrasound enables real-time volumetric visualization of anatomical structures. Unlike traditional 2D ultrasound, 3D imaging reduces reliance on precise probe orientation, potentially making ultrasound more accessible to clinicians with varying levels of experience and improving automated measurements and post-exam analysis. However, achieving both high volume rates and high image quality remains a significant challenge. While 3D diverging waves can provide high volume rates, they suffer from limited tissue harmonic generation and increased multipath effects, which degrade image quality. One compromise is to retain focus in elevation while leveraging unfocused diverging waves in the lateral direction to reduce the number of transmissions per elevation plane. Reaching the volume rates achieved by full 3D diverging waves, however, requires dramatically undersampling the number of elevation planes. Subsequently, to render the full volume, simple interpolation techniques are applied. This paper introduces a novel approach to 3D ultrasound reconstruction from a reduced set of elevation planes by employing diffusion models (DMs) to achieve increased spatial and temporal resolution. We compare both traditional and supervised deep learning-based interpolation methods on a 3D cardiac ultrasound dataset. Our results show that DM-based reconstruction consistently outperforms the baselines in image quality and downstream task performance. Additionally, we accelerate inference by leveraging the temporal consistency inherent to ultrasound sequences. Finally, we explore the robustness of the proposed method by exploiting the probabilistic nature of diffusion posterior sampling to quantify reconstruction uncertainty and demonstrate improved recall on out-of-distribution data with synthetic anomalies under strong subsampling. Code is available at 3d-ultrasound-diffusion.github.io. Tristan S. W. Stevens, Oisín Nolan, Oudom Somphone, Jean-Luc Robert, Ruud van Sloun |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Deep Learning for Modulo Sampling of FRI SignalsabstractThe finite rate of innovation (FRI) of certain signal classes allows for sub-Nyquist sampling, while modulo folding enables sampling signals with dynamic ranges beyond that of an analog-to-digital converter (ADC). Combining these techniques enables sampling at sub-Nyquist rates using an ADC with a lower dynamic range than the signal. Current signal recovery techniques use algorithmic approaches with high oversampling factors (OFs) and constraints on the signal parameters. In this paper, we introduce a deep learning method to modulo unfolding and combine it with an annihilating filter approach for signal parameter recovery. Our method significantly improves unfolding accuracy and reduces error compared to the state-of-the-art, and performs well at OFs as low as 2 times the rate of innovation, with no constraints on the signal parameters. This approach offers a practical tool for signal recovery from modulo sampled FRI signals, potentially reducing the hardware demands of measurement devices. Sem Koenen, Vincent van de Schaft, Ruud van Sloun |
ICASSP | 3 |
| 2025 | Deep Variational Sequential Monte Carlo for High-Dimensional ObservationsabstractSequential Monte Carlo (SMC), or particle filtering, is widely used in nonlinear state-space systems, but its performance often suffers from poorly approximated proposal and state-transition distributions. This work introduces a differentiable particle filter that leverages the unsupervised variational SMC objective to parameterize the proposal and transition distributions with a neural network, designed to learn from high-dimensional observations. Experimental results demonstrate that our approach outperforms established baselines in tracking the challenging Lorenz attractor from high-dimensional and partial observations. Furthermore, an evidence lower bound based evaluation indicates that our method offers a more accurate representation of the posterior distribution. Wessel L. van Nierop, Nir Shlezinger, Ruud van Sloun |
ICASSP | 3 |
| 2025 | Deep Sylvester Posterior Inference for Adaptive Compressed Sensing in Ultrasound ImagingabstractUltrasound images are commonly formed by sequential acquisition of beam-steered scan-lines. Minimizing the number of required scan-lines can significantly enhance frame rate, field of view, energy efficiency, and data transfer speeds. Existing approaches typically use static subsampling schemes in combination with sparsity-based or, more recently, deep-learning-based recovery. In this work, we introduce an adaptive subsampling method that maximizes intrinsic information gain in-situ, employing a Sylvester Normalizing Flow encoder to infer an approximate Bayesian posterior under partial observation in real-time. Using the Bayesian posterior and a deep generative model for future observations, we determine the subsampling scheme that maximizes the mutual information between the subsampled observations, and the next frame of the video. We evaluate our approach using the EchoNet cardiac ultrasound video dataset and demonstrate that our active sampling method outperforms competitive baselines, including uniform and variable-density random sampling, as well as equidistantly spaced scanlines, improving mean absolute reconstruction error by 15%. Moreover, posterior inference and the sampling scheme generation are performed in just 0.015 seconds (66Hz), making it fast enough for real-time 2D ultrasound imaging applications. Simon W. Penninga, Hans Van Gorp, Ruud van Sloun |
ICASSP | 3 |
| 2025 | Sequential Posterior Sampling with Diffusion ModelsabstractDiffusion models have quickly risen in popularity for their ability to model complex distributions and perform effective posterior sampling. Unfortunately, the iterative nature of these generative models makes them computationally expensive and unsuitable for real-time sequential inverse problems such as ultrasound imaging. Considering the strong temporal structure across sequences of frames, we propose a novel approach that models the transition dynamics to improve the efficiency of sequential diffusion posterior sampling in conditional image synthesis. Through modeling sequence data using a video vision transformer (ViViT) transition model based on previous diffusion outputs, we can initialize the reverse diffusion trajectory at a lower noise scale, greatly reducing the number of iterations required for convergence. We demonstrate the effectiveness of our approach on a real-world dataset of high frame rate cardiac ultrasound images and show that it achieves the same performance as a full diffusion trajectory while accelerating inference 25×, enabling real-time posterior sampling. Furthermore, we show that the addition of a transition model improves the PSNR up to 8% in cases with severe motion. Our method opens up new possibilities for real-time applications of diffusion models in imaging and other domains requiring real-time inference. Tristan S. W. Stevens, Oisín Nolan, Jean-Luc Robert, Ruud van Sloun |
ICASSP | 4 |
| 2025 | Learning Structured Compressed Sensing with Automatic Resource AllocationabstractMultidimensional data acquisition often requires extensive time and poses significant challenges for hardware and software regarding data storage and processing. Rather than designing a single compression matrix as in conventional compressed sensing, structured compressed sensing yields dimension-specific compression matrices, reducing the number of optimizable parameters. Recent advances in machine learning (ML) have enabled task-based supervised learning of subsampling matrices, albeit at the expense of complex downstream models. Additionally, the sampling resource allocation across dimensions is often determined in advance through heuristics. To address these challenges, we introduce Structured COmpressed Sensing with Automatic Resource Allocation (SCOSARA) with an information theory-based unsupervised learning strategy. SCOSARA adaptively distributes samples across sampling dimensions while maximizing Fisher information content. Using ultrasound localization as a case study, we compare SCOSARA to state-of-the-art ML-based and greedy search algorithms. Simulation results demonstrate that SCOSARA can produce high-quality subsampling matrices that achieve lower Cramér Rao Bound values than the baselines. In addition, SCOSARA outperforms other ML-based algorithms in terms of the number of trainable parameters, computational complexity, and memory requirements while automatically choosing the number of samples per axis. Han Wang 0052, Iris A. M. Huijben, Hans Van Gorp, Ruud van Sloun, Florian Roemer |
ICASSP | 5 |
| 2025 | Deep Unfolding Using Score-based Generative Networks for Automotive Radar Interference MitigationabstractAutomotive frequency-modulated continuous wave (FMCW) radars, essential in Advanced Driver Assistance Systems, encounter mutual interference issues that degrade their detection capabilities. Model-based algorithms, though widely used, rely heavily on predetermined assumptions about the statistical properties. General-purpose black-box deep learning approaches, while effective in their training distribution, often lack flexibility and generalizability in dynamic environments. We introduce a novel hybrid method that combines model-based techniques with deep learning, treating interference mitigation as a source separation problem. Specifically, our method employs score-based deep generative networks to accurately capture the structure of FMCW interference. Additionally, we employ deep unfolding to accelerate inference, critical for automotive radar applications. Empirical results from simulated data demonstrate that the proposed algorithm outperforms the baseline models by 3.26 dB in signal-to-interference-plus-noise ratio in the presence of aggressive interference, and also shows good generalizability with measured data. Xinyi Wei, Jihwan Youn, Jeroen Overdevest, Jun Li 0091, Satish Ravindran, Ruud van Sloun |
ICASSP | 6 |
| 2025 | A Deep Generative Model for Five-Class Sleep Staging With Arbitrary Sensor InputabstractGold-standard sleep scoring is based on epoch-based assignment of sleep stages based on a combination of EEG, EOG and EMG signals. However, a polysomnographic recording consists of many other signals that could be used for sleep staging, including cardio-respiratory modalities. Leveraging this signal variety would offer important advantages, for example increasing reliability, resilience to signal loss, and application to long-term non-obtrusive recordings. We developed a deep generative model for automatic sleep staging from a plurality of sensors and any -arbitrary- combination thereof. We trained a score-based diffusion model using a dataset of 1947 expert-labelled overnight recordings with 36 different signals, and achieved zero-shot inference on any sensor set by leveraging a novel Bayesian factorization of the score function across the sensors. On single-channel EEG, the model reaches the performance limit in terms of polysomnography inter-rater agreement (5- class accuracy 85.6%, Cohen's kappa 0.791). Moreover, the method offers full flexibility to use any sensor set, for example finger photoplethysmography, nasal flow and thoracic respiratory movements, (5-class accuracy 79.0%, Cohen's kappa of 0.697), or even derivations very unconventional for sleep staging, such as tibialis and sternocleidomastoid EMG (5-class accuracy 71.0%, kappa 0.575). Additionally, we propose a novel interpretability metric in terms of information gain per sensor and show this is linearly correlated with classification performance. Finally, our model allows for post- hoc addition of entirely new sensor modalities by merely training a score estimator on the novel input instead of having to retrain from scratch on all inputs. Hans Van Gorp, Merel van Gilst, Pedro Fonseca 0002, Fokke B. van Meulen, Johannes P. van Dijk, Sebastiaan Overeem, Ruud van Sloun |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Conditional Contrastive Predictive Coding for Assessment of Fetal Health From the CardiotocogramabstractFetal well-being during labor is currently assessed by medical professionals through visual interpretation of the cardiotocogram, a simultaneous recording of Fetal Heart Rate and Uterine Activity. This method is disputed due to high inter- and intra-observer variability and a resulting high number of unnecessary interventions. Recently, an unsupervised deep learning model for automated anomaly detection in the cardiotocogram was presented. Anomalies were defined as out-of-distribution behaviour or deviations from subject-specific behaviour and the model was based on the WaveNet architecture, but required a two-step training. The current work improves this previous work by leveraging Contrastive Predictive Coding (CPC), which uses a contrastive loss to make latent predictions without requiring a decoder network. In this work, CPC was extended with a stochastic, recurrent, and conditioned (upon Uterine Activity) future predictor. We, moreover, introduce a new training objective that was found better suitable for the task of anomaly detection. Evaluated on annotations made by experienced gynecologists, all proposed extensions were shown to be beneficial, and the proposed method is shown to rival or outperform the WaveNet-based method on different annotation categories. Ivar R. de Vries, Raoul Melaet, Iris A. M. Huijben, Judith O. E. H. Van Laar, René D. Kok, S. Guid Oei, Ruud van Sloun, Rik Vullings |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Investigating and Improving Latent Density Segmentation Models for Aleatoric Uncertainty Quantification in Medical ImagingabstractData uncertainties, such as sensor noise, occlusions or limitations in the acquisition method can introduce irreducible ambiguities in images, which result in varying, yet plausible, semantic hypotheses. In Machine Learning, this ambiguity is commonly referred to as aleatoric uncertainty. In image segmentation, latent density models can be utilized to address this problem. The most popular approach is the Probabilistic U-Net (PU-Net), which uses latent Normal densities to optimize the conditional data log-likelihood Evidence Lower Bound. In this work, we demonstrate that the PU-Net latent space is severely sparse and heavily under-utilized. To address this, we introduce mutual information maximization and entropy-regularized Sinkhorn Divergence in the latent space to promote homogeneity across all latent dimensions, effectively improving gradient-descent updates and latent space informativeness. Our results show that by applying this on public datasets of various clinical segmentation problems, our proposed methodology receives up to 11% performance gains compared against preceding latent variable models for probabilistic segmentation on the Hungarian-Matched Intersection over Union. The results indicate that encouraging a homogeneous latent space significantly improves latent density modeling for medical image segmentation. M. M. Amaan Valiuddin, Christiaan G. A. Viviers, Ruud van Sloun, Peter H. N. de With, Fons van der Sommen |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Retaining Informative Latent Variables in Probabilistic SegmentationabstractConditional latent-variable models can successfully quantify annotation variability in segmentation. Training such models involves tuning the dimensionality of the latent space to optimally capture the inherent data ambiguity. Nevertheless, we discover after careful tuning, that the latent space does not always reflect this. In fact, some latent dimensions are completely neglected. For such segmentation models the latent dimensionality is often poorly motivated or based on computational constraints. In this paper, we offer an information-theoretic approach to optimally leverage all latent dimensions. We adapt and improve the Probabilistic U-Net to maximize the mutual information between the latent and output variables, leading to improved latent space properties and higher segmentation performance. M. M. Amaan Valiuddin, Christiaan G. A. Viviers, Ruud van Sloun, Peter H. N. de With, Fons van der Sommen |
ICASSP | 3 |
| 2024 | Residual Quantization with Implicit Neural CodebooksabstractVector quantization is a fundamental operation for data compression and vector search. To obtain high accuracy, multi-codebook methods represent each vector using codewords across several codebooks. Residual quantization (RQ) is one such method, which iteratively quantizes the error of the previous step. While the error distribution is dependent on previously-selected codewords, this dependency is not accounted for in conventional RQ as it uses a fixed codebook per quantization step. In this paper, we propose QINCo, a neural RQ variant that constructs specialized codebooks per step that depend on the approximation of the vector from previous steps. Experiments show that QINCo outperforms state-of-the-art methods by a large margin on several datasets and code sizes. For example, QINCo achieves better nearest-neighbor search accuracy using 12-byte codes than the state-of-the-art UNQ using 16 bytes on the BigANN1M and Deep1M datasets. Iris A. M. Huijben, Matthijs Douze, Matthew J. Muckley, Ruud van Sloun, Jakob Verbeek |
ICML | 4 |
| 2024 | Exploring the trade-off between deep-learning and explainable models for brain-machine interfacesabstractPeople with brain or spinal cord-related paralysis often need to rely on others for basic tasks, limiting their independence. A potential solution is brain-machine interfaces (BMIs), which could allow them to voluntarily control external devices (e.g., robotic arm) by decoding brain activity to movement commands. In the past decade, deep-learning decoders have achieved state-of-the-art results in most BMI applications, ranging from speech production to finger control. However, the 'black-box' nature of deep-learning decoders could lead to unexpected behaviors, resulting in major safety concerns in real-world physical control scenarios. In these applications, explainable but lower-performing decoders, such as the Kalman filter (KF), remain the norm. In this study, we designed a BMI decoder based on KalmanNet, an extension of the KF that augments its operation with recurrent neural networks to compute the Kalman gain. This results in a varying “trust” that shifts between inputs and dynamics. We used this algorithm to predict finger movements from the brain activity of two monkeys. We compared KalmanNet results offline (pre-recorded data, $n=13$ days) and online (real-time predictions, $n=5$ days) with a simple KF and two recent deep-learning algorithms: tcFNN (non-ReFIT version) and LSTM. KalmanNet achieved comparable or better results than other deep learning models in offline and online modes, relying on the dynamical model for stopping while depending more on neural inputs for initiating movements. We further validated this mechanism by implementing a heteroscedastic KF that used the same strategy, and it also approached state-of-the-art performance while remaining in the explainable domain of standard KFs. However, we also see two downsides to KalmanNet. KalmanNet shares the limited generalization ability of existing deep-learning decoders, and its usage of the KF as an inductive bias limits its performance in the presence of unseen noise distributions. Despite this trade-off, our analysis successfully integrates traditional controls and modern deep-learning approaches to motivate high-performing yet still explainable BMI designs. Luis Cubillos, Guy Revach, Matthew Mender, Joseph T. Costello, Hisham Temmar, Aren Hite, Diksha Anoop Kumar Zutshi, Dylan Wallace, Xiaoyong Ni, Madison Kelberman, Matt S. Willsey, Ruud van Sloun, Nir Shlezinger, Parag G. Patil, Anne Draelos, Cynthia A. Chestek |
NeurIPS | 12 |
| 2024 | ULTRA-SR Challenge: Assessment of Ultrasound Localization and TRacking Algorithms for Super-Resolution ImagingabstractWith the widespread interest and uptake of super-resolution ultrasound (SRUS) through localization and tracking of microbubbles, also known as ultrasound localization microscopy (ULM), many localization and tracking algorithms have been developed. ULM can image many centimeters into tissue in-vivo and track microvascular flow non-invasively with sub-diffraction resolution. In a significant community effort, we organized a challenge, Ultrasound Localization and TRacking Algorithms for Super-Resolution (ULTRA-SR). The aims of this paper are threefold: to describe the challenge organization, data generation, and winning algorithms; to present the metrics and methods for evaluating challenge entrants; and to report results and findings of the evaluation. Realistic ultrasound datasets containing microvascular flow for different clinical ultrasound frequencies were simulated, using vascular flow physics, acoustic field simulation and nonlinear bubble dynamics simulation. Based on these datasets, 38 submissions from 24 research groups were evaluated against ground truth using an evaluation framework with six metrics, three for localization and three for tracking. In-vivo mouse brain and human lymph node data were also provided, and performance assessed by an expert panel. Winning algorithms are described and discussed. The publicly available data with ground truth and the defined metrics for both localization and tracking present a valuable resource for researchers to benchmark algorithms and software, identify optimized methods/software for their data, and provide insight into the current limits of the field. In conclusion, Ultra-SR challenge has provided benchmarking data and tools as well as direct comparison and insights for a number of the state-of-the art localization and tracking algorithms. Marcelo Lerendegui, Kai Riemer, Georgios K. Papageorgiou, Bingxue Wang, Lachlan Arthur, Arthur Chavignon, Olivier Couture, Pingtong Huang, Md Ashikuzzaman, Stefanie Dencks, Christopher Dunsby, Brandon Helfield, Jørgen Arendt Jensen, Thomas Lisson, Matthew R. Lowerison, Hassan Rivaz, Anthony E. Samir, Georg Schmitz, Scott J. Schoen, Ruud van Sloun, Tristan S. W. Stevens, Jipeng Yan 0001, Vassilis Sboros, Meng-Xing Tang |
IEEE Trans. Medical Imaging | 21 |
| 2024 | Dehazing Ultrasound Using Diffusion ModelsabstractEchocardiography has been a prominent tool for the diagnosis of cardiac disease. However, these diagnoses can be heavily impeded by poor image quality. Acoustic clutter emerges due to multipath reflections imposed by layers of skin, subcutaneous fat, and intercostal muscle between the transducer and heart. As a result, haze and other noise artifacts pose a real challenge to cardiac ultrasound imaging. In many cases, especially with difficult-to-image patients such as patients with obesity, a diagnosis from B-Mode ultrasound imaging is effectively rendered unusable, forcing sonographers to resort to contrast-enhanced ultrasound examinations or refer patients to other imaging modalities. Tissue harmonic imaging has been a popular approach to combat haze, but in severe cases is still heavily impacted by haze. Alternatively, denoising algorithms are typically unable to remove highly structured and correlated noise, such as haze. It remains a challenge to accurately describe the statistical properties of structured haze, and develop an inference method to subsequently remove it. Diffusion models have emerged as powerful generative models and have shown their effectiveness in a variety of inverse problems. In this work, we present a joint posterior sampling framework that combines two separate diffusion models to model the distribution of both clean ultrasound and haze in an unsupervised manner. Furthermore, we demonstrate techniques for effectively training diffusion models on radio-frequency ultrasound data and highlight the advantages over image data. Experiments on both in-vitro and in-vivo cardiac datasets show that the proposed dehazing method effectively removes haze while preserving signals from weakly reflected tissue. Tristan S. W. Stevens, Faik C. Meral, Jason Yu, Iason Zacharias Apostolakis, Jean-Luc Robert, Ruud van Sloun |
IEEE Trans. Medical Imaging | 6 |
| 2024 | Dynamic Probabilistic Pruning: A General Framework for Hardware-Constrained Pruning at Different GranularitiesabstractUnstructured neural network pruning algorithms have achieved impressive compression ratios. However, the resulting-typically irregular-sparse matrices hamper efficient hardware implementations, leading to additional memory usage and complex control logic that diminishes the benefits of unstructured pruning. This has spurred structured coarse-grained pruning solutions that prune entire feature maps or even layers, enabling efficient implementation at the expense of reduced flexibility. Here, we propose a flexible new pruning mechanism that facilitates pruning at different granularities (weights, kernels, and feature maps) while retaining efficient memory organization (e.g., pruning exactly k -out-of- n weights for every output neuron or pruning exactly k -out-of- n kernels for every feature map). We refer to this algorithm as dynamic probabilistic pruning (DPP). DPP leverages the Gumbel-softmax relaxation for differentiable k -out-of- n sampling, facilitating end-to-end optimization. We show that DPP achieves competitive compression ratios and classification accuracy when pruning common deep learning models trained on different benchmark datasets for image classification. Relevantly, the dynamic masking of DPP facilitates for joint optimization of pruning and weight quantization in order to even further compress the network, which we show as well. Finally, we propose novel information-theoretic metrics that show the confidence and pruning diversity of pruning masks within a layer. Lizeth Gonzalez-Carabarin, Iris A. M. Huijben, Bastiaan S. Veeling, Alexandre Schmid, Ruud van Sloun |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | The Transform-and-Perform Framework: Explainable Deep Learning Beyond ClassificationabstractIn recent years, visual analytics (VA) has shown promise in alleviating the challenges of interpreting black-box deep learning (DL) models. While the focus of VA for explainable DL has been mainly on classification problems, DL is gaining popularity in high-dimensional-to-high-dimensional (H-H) problems such as image-to-image translation. In contrast to classification, H-H problems have no explicit instance groups or classes to study. Each output is continuous, high-dimensional, and changes in an unknown non-linear manner with changes in the input. These unknown relations between the input, model and output necessitate the user to analyze them in conjunction, leveraging symmetries between them. Since classification tasks do not exhibit some of these challenges, most existing VA systems and frameworks allow limited control of the components required to analyze models beyond classification. Hence, we identify the need for and present a unified conceptual framework, the Transform-and-Perform framework (T&P), to facilitate the design of VA systems for DL model analysis focusing on H-H problems. T&P provides a checklist to structure and identify workflows and analysis strategies to design new VA systems, and understand existing ones to uncover potential gaps for improvements. The goal is to aid the creation of effective VA systems that support the structuring of model understanding and identifying actionable insights for model improvements. We highlight the growing need for new frameworks like T&P with a real-world image-to-image translation application. We illustrate how T&P effectively supports the understanding and identification of potential gaps in existing VA systems. Vidya Prasad, Ruud van Sloun, Stef van den Elzen, Anna Vilanova, Nicola Pezzotti |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | ProactiV: Studying Deep Learning Model Behavior Under Input TransformationsabstractDeep learning (DL) models have shown performance benefits across many applications, from classification to image-to-image translation. However, low interpretability often leads to unexpected model behavior once deployed in the real world. Usually, this unexpected behavior is because the training data domain does not reflect the deployment data domain. Identifying a model's breaking points under input conditions and domain shifts, i.e., input transformations, is essential to improve models. Although visual analytics (VA) has shown promise in studying the behavior of model outputs under continually varying inputs, existing methods mainly focus on per-class or instance-level analysis. We aim to generalize beyond classification where classes do not exist and provide a global view of model behavior under co-occurring input transformations. We present a DL model-agnostic VA method (ProactiV) to help model developers proactively study output behavior under input transformations to identify and verify breaking points. ProactiV relies on a proposed input optimization method to determine the changes to a given transformed input to achieve the desired output. The data from this optimization process allows the study of global and local model behavior under input transformations at scale. Additionally, the optimization method provides insights into the input characteristics that result in desired outputs and helps recognize model biases. We highlight how ProactiV effectively supports studying model behavior with example classification and image-to-image translation tasks. Vidya Prasad, Ruud van Sloun, Anna Vilanova, Nicola Pezzotti |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Learned Kalman Filtering in Latent Space with High-Dimensional DataabstractThe Kalman filter (KF) is a widely-used algorithm for tracking dynamical systems that can be faithfully captured by state space (SS) models. The need to fully describe an SS model limits its applicability under complex settings, e.g., when tracking based on visual or graphical data. This challenge can be treated by mapping the measurements into latent features obeying some postulated closed-form SS model, and applying the KF in the latent space. However, the validity of this approximated SS model may constitute a limiting factor. In this work we tackle the challenges associated with tracking from high-dimensional measurements by jointly learning the KF along with the latent space mapping. Our proposed approach combines a learned encoder while tracking in the latent space using the recently proposed data-driven Kalman-Net, and having both modules jointly tuned from data. Our empirical results demonstrate that the proposed approach achieves improved performance over both model-based and data-driven techniques, by learning a surrogate latent representation that most facilitates tracking. Itay Buchnik, Damiano Steger, Guy Revach, Ruud van Sloun, Tirza Routtenberg, Nir Shlezinger |
ICASSP | 4 |
| 2023 | Active Subsampling Using Deep Generative Models by Maximizing Expected Information GainabstractWe introduce an adaptive, fully probabilistic pipeline for optimized signal subsampling in sampling-budget constrained systems. Our pipeline equips an agent with a deep generative model of its measurement-generating environment with which it infers posterior distributions over high-dimensional signals. This posterior distribution is subsequently used by the agent to adaptively select next samples that maximize the expected information gain. Experiments on the MNIST and fastMRI data sets show strong adaptability of selected sampling sequences to the signal modality, resulting in high-quality reconstructions for high acceleration factors. Performance is upper-bounded by the representation error of the used generative model, which is mainly evident at low acceleration factors. Koen C. E. van de Camp, Hamdi Joudeh, Duarte Antunes, Ruud van Sloun |
ICASSP | 4 |
| 2023 | Aleatoric Uncertainty Estimation of Overnight Sleep Statistics Through Posterior Sampling Using Conditional Normalizing FlowsabstractIn sleep staging, a polysomnography is visually scored by a human expert, who creates a hypnogram that classifies the measurement into a sequence of sleep stages, from which overnight sleep statistics, such as total sleep time, are derived. Because inter-scorer agreement between humans is limited, deep learning methods trained to automate sleep staging have aleatoric uncertainty about both hypnogram and overnight statistics. We would like to estimate this aleatoric uncertainty, which can be achieved by means of posterior sampling. Current approaches model the hypnogram through a time-based factorization of categorical distributions over sleep stages. This discards time-dependent information, invalidating posterior sampling of the overnight statistics. Instead of factorizing, we propose to jointly model the sequence of sleep stages, by introducing U-Flow, a conditional normalizing flow network. We compare U-Flow to factorized baselines, leveraging 921 recordings, and show that it achieves similar performance in terms of accuracy and Cohen’s kappa on the majority voted hypnograms, while outperforming in terms of uncertainty estimation of the overnight sleep statistics. Hans Van Gorp, Merel van Gilst, Pedro Fonseca 0002, Sebastiaan Overeem, Ruud van Sloun |
ICASSP | 5 |
| 2023 | Hierarchical Filtering With Online Learned Priors for ECG DenoisingabstractElectrocardiographic signals (ECG) are used in many healthcare applications, including at-home monitoring of vital signs. These applications often rely on wearable technology and provide low quality ECG signals. Although many methods have been proposed for denoising the ECG to boost its quality and enable clinical interpretation, these methods typically fall short for ECG data obtained with wearable technology, because of either their limited tolerance to noise or their limited flexibility to capture ECG dynamics. This paper presents HKF, a hierarchical Kalman filtering method, that leverages a patient-specific learned structured prior of the ECG signal, and integrates it into a state space model to yield filters that capture both intra- and inter-heartbeat dynamics. HKF is demonstrated to outperform previously proposed methods such as the model-based Kalman filter and data-driven autoencoders, in ECG denoising task in terms of mean-squared error, making it a suitable candidate for application in extramural healthcare settings. Timur Locher, Guy Revach, Nir Shlezinger, Ruud van Sloun, Rik Vullings |
ICASSP | 4 |
| 2023 | Neural Maximum-a-Posteriori Beamforming for Ultrasound ImagingabstractUltrasound imaging is an attractive imaging modality due to its low-cost and real-time feedback, although it often falls short in image quality compared to MRI and CT imaging. Conventional ultrasound image reconstruction, such as Delay-and-Sum beamforming, is derived from maximum-likelihood estimation. As such, no prior information is exploited in the image formation process, which limits potential image quality. Maximum-a-posteriori (MAP) beamforming aims to overcome this issue, but often relies on rough approximations of the underlying signal statistics. Deep learning based reconstruction methods have demonstrated impressive results over the past years, but often lack interpretability and require vast amounts of data.In this work we present a neural MAP beamforming technique, which efficiently combines deep learning in the MAP beamforming framework. We show that this model-based deep learning approach can achieve high-quality imaging, improving over the state-of-the-art, without compromising the real-time abilities of ultrasound imaging. Ben Luijten, Boudewine W. Ossenkoppele, Nico de Jong, Martin D. Verweij, Yonina C. Eldar, Massimo Mischi, Ruud van Sloun |
ICASSP | 7 |
| 2023 | Signal Reconstruction for FMCW Radar Interference Mitigation Using Deep UnfoldingabstractRemoval of frequency-modulated continuous wave (FMCW) interference by zeroing corrupted samples causes significant distortions and peak power losses in the range-Doppler map. Existing methods aim to diminish these distortions by utilizing data from one dimension to reconstruct the corrupted samples, which do not perform well when a large number of samples are interfered and have difficulty recovering weak target signals.In this paper, model-based deep learning interference mitigation algorithms, called ALISTA and ALFISTA, are presented that reduce these artifacts by leveraging the full integration gain using all uncorrupted fast-time and slow-time samples. Simulations with 50% corrupted samples show that target peak power loss and velocity peak-to-sidelobe ratio (VPSR) with a 20-layer ALFISTA improves with 5.5 and 9.6 dB compared to zeroing. Furthermore, significant improvements in precision and recall are observed, even when large amounts (50-80%) of samples are missing. Jeroen Overdevest, A. G. C. Koppelaar, M. J. G. Bekooij, J. Youn, Ruud van Sloun |
ICASSP | 5 |
| 2023 | Deep Root Music Algorithm for Data-Driven Doa EstimationabstractDirection of arrival (DoA) estimation is a fundamental task in array processing. A popular family of DoA estimation algorithms are subspace methods, which operate by dividing the measurements into distinct signal and noise subspaces. Subspace methods, such as Root-MUSIC, require the sources to be non-coherent, and are considerably degraded when this does not hold. In this work we propose Deep Root-MUSIC (DR-MUSIC); a data-driven DoA estimator which augments Root-MUSIC with a deep neural network applied to the empirical autocorrelation of the input. DR-MUSIC learns how to divide the observations into distinguishable subspaces, thus leveraging data to cope with coherent sources, low SNR and limited snapshots, while preserving the interpretability and the suitability of the model-based algorithm. Dor Haim Shmuel, Julian P. Merkofer, Guy Revach, Ruud van Sloun, Nir Shlezinger |
ICASSP | 4 |
| 2023 | SOM-CPC: Unsupervised Contrastive Learning with Self-Organizing Maps for Structured Representations of High-Rate Time SeriesabstractContinuous monitoring with an ever-increasing number of sensors has become ubiquitous across many application domains. However, acquired time series are typically high-dimensional and difficult to interpret. Expressive deep learning (DL) models have gained popularity for dimensionality reduction, but the resulting latent space often remains difficult to interpret. In this work we propose SOM-CPC, a model that visualizes data in an organized 2D manifold, while preserving higher-dimensional information. We address a largely unexplored and challenging set of scenarios comprising high-rate time series, and show on both synthetic and real-life data (physiological data and audio recordings) that SOM-CPC outperforms strong baselines like DL-based feature extraction, followed by conventional dimensionality reduction techniques, and models that jointly optimize a DL model and a Self-Organizing Map (SOM). SOM-CPC has great potential to acquire a better understanding of latent patterns in high-rate data streams. Iris A. M. Huijben, Arthur A. Nijdam, Sebastiaan Overeem, Merel van Gilst, Ruud van Sloun |
ICML | 5 |
| 2023 | A Review of the Gumbel-max Trick and its Extensions for Discrete Stochasticity in Machine LearningabstractThe Gumbel-max trick is a method to draw a sample from a categorical distribution, given by its unnormalized (log-)probabilities. Over the past years, the machine learning community has proposed several extensions of this trick to facilitate, e.g., drawing multiple samples, sampling from structured domains, or gradient estimation for error backpropagation in neural network optimization. The goal of this survey article is to present background about the Gumbel-max trick, and to provide a structured overview of its extensions to ease algorithm selection. Moreover, it presents a comprehensive outline of (machine learning) literature in which Gumbel-based algorithms have been leveraged, reviews commonly-made design choices, and sketches a future perspective. Iris A. M. Huijben, Wouter Kool 0001, Max B. Paulus, Ruud van Sloun |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Modeling the Impact of Inter-Rater Disagreement on Sleep Statistics Using Deep Generative LearningabstractSleep staging is the process by which an overnight polysomnographic measurement is segmented into epochs of 30 seconds, each of which is annotated as belonging to one of five discrete sleep stages. The resulting scoring is graphically depicted as a hypnogram, and several overnight sleep statistics are derived, such as total sleep time and sleep onset latency. Gold standard sleep staging as performed by human technicians is time-consuming, costly, and comes with imperfect inter-scorer agreement, which also results in inter-scorer disagreement about the overnight statistics. Deep learning algorithms have shown promise in automating sleep scoring, but struggle to model inter-scorer disagreement in sleep statistics. To that end, we introduce a novel technique using conditional generative models based on Normalizing Flows that permits the modeling of the inter-rater disagreement of overnight sleep statistics, termed U-Flow. We compare U-Flow to other automatic scoring methods on a hold-out test set of 70 subjects, each scored by six independent scorers. The proposed method achieves similar sleep staging performance in terms of accuracy and Cohen's kappa on the majority-voted hypnograms. At the same time, U-Flow outperforms the other methods in terms of modeling the inter-rater disagreement of overnight sleep statistics. The consequences of inter-rater disagreement about overnight sleep statistics may be great, and the disagreement potentially carries diagnostic and scientifically relevant information about sleep structure. U-Flow is able to model this disagreement efficiently and can support further investigations into the impact inter-rater disagreement has on sleep medicine and basic sleep research. Hans Van Gorp, Merel van Gilst, Pedro Fonseca 0002, Sebastiaan Overeem, Ruud van Sloun |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Deep Proximal Unfolding For Image Recovery from Under-Sampled Channel Data in Intravascular UltrasoundabstractIntravascular UltraSound (IVUS) is a key tool in guiding the treatment and diagnosis of various coronary heart diseases. However, due to its nature IVUS is a very challenging modality to interpret, and suffers from a severely restricted data transfer rate. This forces a trade-off between temporal and spatial resolution. Here, we propose a model-based deep learning solution that aims to reconstruct images from data that has been beamformed by under-sampling the number of channels by a factor of 4. By exploiting the physics based measurement model, we achieve better performance and consistency in our predictions when compared to benchmark models. This lowers the computational load on existing hardware and enables in exploring our ability to run multiple visualisation modalities simultaneously, without a loss of temporal resolution. Nishith Chennakeshava, Tristan S. W. Stevens, Frederik J. de Bruijn, Andrew Hancock, Martin Pekar, Yonina C. Eldar, Massimo Mischi, Ruud van Sloun |
ICASSP | 8 |
| 2022 | Unfolding Model-Based Beamforming for High Quality Ultrasound ImagingabstractAperture Domain Model Image REconstruction (ADMIRE) is an advanced ultrasound beamforming method that uses a model-based approach to suppress sources of acoustic clutter and improve ultrasound image quality. However, it requires solving an ill-posed inverse problem for which regularization is utilized. As a result, the iterative nature of solving this problem is computationally expensive, and the choice of regularization bounds the fidelity of the obtained solution. Therefore, in this work, we pose ADMIRE as a sparse coding problem and unfold the iterations of the iterative shrinkage and thresholding algorithm (ISTA), training it end to end to yield learned ISTA (LISTA). This enables effective tailoring of the solver to the specific data distribution and task at hand, while enjoying higher data efficiency and robustness than generic deep learning methods. Evaluation of our proposed method on both simulated cyst data and in vivo liver data demonstrates its potential to outperform conventional ADMIRE. Christopher Khan, Ruud van Sloun, Brett C. Byram |
ICASSP | 2 |
| 2022 | Uncertainty in Data-Driven Kalman Filtering for Partially Known State-Space ModelsabstractProviding a metric of uncertainty alongside a state estimate is often crucial when tracking a dynamical system. Classic state estimators, such as the Kalman filter (KF), provide a time-dependent uncertainty measure from knowledge of the underlying statistics; however, deep learning based tracking systems struggle to reliably characterize uncertainty. In this paper, we investigate the ability of KalmanNet, a recently proposed; hybrid; model-based; deep state tracking algorithm, to estimate an uncertainty measure. By exploiting the interpretable nature of KalmanNet, we show that the error covariance matrix can be computed based on its internal features, as an uncertainty measure. We demonstrate that when the system dynamics are known, KalmanNet—which learns its mapping from data without access to the statistics—provides uncertainty similar to that provided by the KF; and while in the presence of evolution model-mismatch, KalmanNet provides a more accurate error estimation. Itzik Klein, Guy Revach, Nir Shlezinger, Jonas E. Mehr, Ruud van Sloun, Yonina C. Eldar |
ICASSP | 5 |
| 2022 | Deep Augmented Music Algorithm for Data-Driven Doa EstimationabstractDirection of arrival (DoA) estimation is a crucial task in sensor array signal processing, giving rise to various successful model-based (MB) algorithms as well as recently developed data-driven (DD) methods. This paper introduces a new hybrid MB/DD DoA estimation architecture, based on the classical multiple signal classification (MUSIC) algorithm. Our approach augments crucial aspects of the original MUSIC structure with specifically designed neural architectures, allowing it to overcome certain limitations of the purely MB method, such as its inability to successfully localize coherent sources. The deep augmented MUSIC algorithm is shown to outperform its unaltered version with a superior resolution. Julian P. Merkofer, Guy Revach, Nir Shlezinger, Ruud van Sloun |
ICASSP | 4 |
| 2022 | RTSNet: Deep Learning Aided Kalman SmoothingabstractThe smoothing task is the core of many signal processing applications. It deals with the recovery of a sequence of hidden state variables from a sequence of noisy observations in a one-shot manner. In this work we propose RTSNet, a highly efficient model-based and data-driven smoothing algorithm. RTSNet integrates dedicated trainable models into the flow of the classical Rauch-Tung-Striebel (RTS) smoother, and is able to outperform it when operating under model mismatch and non-linearities while retaining its efficiency and interpretability. Our numerical study demonstrates that although RTSNet is based on more compact neural networks, which leads to faster training and inference times, it outperforms the state-of-the-art, data-driven smoother in a non-linear use case. Xiaoyong Ni, Guy Revach, Nir Shlezinger, Ruud van Sloun, Yonina C. Eldar |
ICASSP | 4 |
| 2022 | Accelerated Intravascular Ultrasound Imaging using Deep Reinforcement LearningabstractIntravascular ultrasound (IVUS) offers a unique perspective in the treatment of vascular diseases by creating a sequence of ultrasound-slices acquired from within the vessel. However, unlike conventional hand-held ultrasound, the thin catheter only provides room for a small number of physical channels for signal transfer from a transducer-array at the tip. For continued improvement of image quality and frame rate, we present the use of deep reinforcement learning to deal with the current physical information bottleneck. Valuable inspiration has come from the field of magnetic resonance imaging (MRI), where learned acquisition schemes have brought significant acceleration in image acquisition at competing image quality. To efficiently accelerate IVUS imaging, we propose a framework that utilizes deep reinforcement learning for an optimal adaptive acquisition policy on a per-frame basis enabled by actor-critic methods and Gumbel top-K sampling. Tristan S. W. Stevens, Nishith Chennakeshava, Frederik J. de Bruijn, Martin Pekar, Ruud van Sloun |
ICASSP | 5 |
| 2022 | Contrastive Predictive Coding for Anomaly Detection of Fetal Health from the CardiotocogramabstractFetal well-being during labor is currently assessed by medical professionals through visual interpretation of the cardiotocogram (CTG), a simultaneous recording of Fetal Heart Rate (FHR) and Uterine Contractions (UC). This method is disputed due to high inter- and intra-observer variability and a resulting increase in the number of unnecessary interventions. A method for computerized interpretation of the CTG, based on Contrastive Predictive Coding (CPC) is presented here. We hypothesize that the CPC framework, when trained on healthy fetuses only, can predict the FHR response of healthy fetuses to UC, but will provide significant prediction error in case of fetuses with compromised condition. To that end, we have extended the original CPC model by making stochastic, recurrent, and conditioned (upon Uterine Contractions) predictions. We, moreover, introduce a new training objective that was found more suitable for the task of anomaly detection. Based on the detection of out-of-distribution behaviour and deviations from subject-specific behaviour, the proposed model is capable of achieving promising results for identification of suspicious and anomalous FHR events in the CTG, with an average correlation coefficient of 0.80±0.13 with respect to expert annotations. Bert de Vries, Iris A. M. Huijben, René D. Kok, Ruud van Sloun, Rik Vullings |
ICASSP | 4 |
| 2022 | Image Denoising with Deep Unfolding And Normalizing FlowsabstractMany application domains, spanning from low-level computer vision to medical imaging, require high-fidelity images from noisy measurements. State-of-the-art methods for solving denoising problems combine deep learning with iterative model-based solvers, a concept known as deep algorithm unfolding or unrolling. By combining a-priori knowledge of the forward measurement model with learned proximal image-to-image mappings based on deep networks, these methods yield solutions that are both physically feasible (data-consistent) and perceptually plausible (consistent with prior belief). However, current proximal mappings based on (predominantly convolutional) neural networks only implicitly learn such image priors. In this paper, we propose to make these image priors fully explicit by embedding deep generative models in the form of normalizing flows within the unfolded proximal gradient algorithm, and training the entire algorithm in an end-to-end fashion. We demonstrate that the proposed method outperforms competitive baselines on image denoising. Xinyi Wei, Hans Van Gorp, Lizeth Gonzalez-Carabarin, Daniel Freedman, Yonina C. Eldar, Ruud van Sloun |
ICASSP | 6 |
| 2022 | Structured and tiled-based pruning of Deep Learning models targeting FPGA implementationsabstractModel compression techniques have lead to a reduction of size and number of computations of Deep Learning models. However, techniques such as pruning mostly lack of a real co-optimization with hardware platforms. For instance, implementing unstructured pruning in dedicated hardware is not a straightforward task, which increases memory and reduces the effective bandwidth usage. Moreover, such pruning algorithms should be adapted to certain hardware requirements, such as the use of tiling. Therefore, in this work, we leverage the use of the Gumbel-Softmax relaxation sampling to structurally prune tiles, which benefits further hardware implementations, and additionally allows to jointly optimize with quantization. Additionally, we show that the combination of different pruning scenarios leads to a larger sparsity. Finally, we demonstrate the benefit of using structured pruning on fine-grained elements (weights) in an FPGA design. Lizeth Gonzalez-Carabarin, Alexandre Schmid, Ruud van Sloun |
ISCAS | 3 |
| 2021 | Kalmannet: Data-Driven Kalman FilteringabstractThe Kalman filter (KF) is a celebrated signal processing algorithm, implementing optimal state estimation of dynamical systems that are well represented by a linear Gaussian state-space model. The KF is model-based, and therefore relies on full and accurate knowledge of the underlying model. We present KalmanNet, a hybrid data-driven/model-based filter that does not require full knowledge of the underlying model parameters. KalmanNet is inspired by the classical KF flow and implemented by integrating a dedicated and compact neural network for the Kalman gain computation. We present an offline training method, and numerically illustrate that KalmanNet can achieve optimal performance without full knowledge of the model parameters. We demonstrate that when facing inaccurate parameters KalmanNet learns to achieve notably improved performance compared to KF. Guy Revach, Nir Shlezinger, Ruud van Sloun, Yonina C. Eldar |
ICASSP | 3 |
| 2021 | Active Deep Probabilistic SubsamplingabstractSubsampling a signal of interest can reduce costly data transfer, battery drain, radiation exposure and acquisition time in a wide range of problems. The recently proposed Deep Probabilistic Subsampling (DPS) method effectively integrates subsampling in an end-to-end deep learning model, but learns a static pattern for all datapoints. We generalize DPS to a sequential method that actively picks the next sample based on the information acquired so far; dubbed Active-DPS (A-DPS). We validate that A-DPS improves over DPS for MNIST classification at high subsampling rates. Moreover, we demonstrate strong performance in active acquisition Magnetic Resonance Image (MRI) reconstruction, outperforming DPS and other deep learning methods. Hans Van Gorp, Iris A. M. Huijben, Bastiaan S. Veeling, Nicola Pezzotti, Ruud van Sloun |
ICML | 5 |
| 2021 | Joint Performance Optimization of Monostatic and Bistatic SAR ConfigurationsabstractThis study focuses on the optimization of phased array antenna excitation coefficients for a spaceborne synthetic aperture radar (SAR) system, comprising a joint mono static and bistatic configuration. An amplitude-only binary-coded Genetic algorithm (BGA) is proposed for the joint non-linear optimization problem to improve monostatic and bistatic SAR performance metrics. Based on the developed optimizer, significant performance improvement is achieved for both configurations. Nehir Berk Onat, Ozan Dogan, Mario Azcueta, Ruud van Sloun |
IGARSS | 4 |
| 2021 | Super-Resolution Ultrasound Localization Microscopy Through Deep LearningabstractUltrasound localization microscopy has enabled super-resolution vascular imaging through precise localization of individual ultrasound contrast agents (microbubbles) across numerous imaging frames. However, analysis of high-density regions with significant overlaps among the microbubble point spread responses yields high localization errors, constraining the technique to low-concentration conditions. As such, long acquisition times are required to sufficiently cover the vascular bed. In this work, we present a fast and precise method for obtaining super-resolution vascular images from high-density contrast-enhanced ultrasound imaging data. This method, which we term Deep Ultrasound Localization Microscopy (Deep-ULM), exploits modern deep learning strategies and employs a convolutional neural network to perform localization microscopy in dense scenarios, learning the nonlinear image-domain implications of overlapping RF signals originating from such sets of closely spaced microbubbles. Deep-ULM is trained effectively using realistic on-line synthesized data, enabling robust inference in-vivo under a wide variety of imaging conditions. We show that deep learning attains super-resolution with challenging contrast-agent densities, both in-silico as well as in-vivo. Deep-ULM is suitable for real-time applications, resolving about 70 high-resolution patches ( 128×128 pixels) per second on a standard PC. Exploiting GPU computation, this number increases to 1250 patches per second. Ruud van Sloun, Oren Solomon, Matthew Bruce, Zin Z. Khaing, Hessel Wijkstra, Yonina C. Eldar, Massimo Mischi |
IEEE Trans. Medical Imaging | 1 |
| 2020 | Learning Sampling and Model-Based Signal Recovery for Compressed Sensing MRIabstractCompressed sensing (CS) MRI relies on adequate under-sampling of the k-space to accelerate the acquisition without compromising image quality. Consequently, the design of optimal sampling patterns for these k-space coefficients has received significant attention, with many CS MRI methods exploiting variable-density probability distributions. Realizing that an optimal sampling pattern may depend on the downstream task (e.g. image reconstruction, segmentation, or classification), we here propose joint learning of both task-adaptive k-space sampling and a subsequent model-based proximal-gradient recovery network. The former is enabled through a probabilistic generative model that leverages the Gumbel-softmax relaxation to sample across trainable beliefs while maintaining differentiability. The proposed combination of a highly flexible sampling model and a model-based (sampling-adaptive) image reconstruction network facilitates exploration and efficient training, yielding improved MR image quality compared to other sampling baselines. Iris A. M. Huijben, Bastiaan S. Veeling, Ruud van Sloun |
ICASSP | 3 |
| 2020 | Learning Task-Based Analog-to-Digital Conversion for MIMO ReceiversabstractAnalog-to-digital conversion allows physical signals to be processed using digital hardware. This conversion consists of two stages: Sampling, which maps a continuous-time signal into discrete-time, and quantization, i.e., representing the continuous-amplitude quantities using a finite number of bits. This conversion is typically carried out using generic uniform mappings that are ignorant of the task for which the signal is acquired, and can be costly when operating in high rates and fine resolutions. In this work we design task-oriented analog-to-digital converters (ADCs) which operate in a data-driven manner, namely they learn how to map an analog signal into a sampled digital representation such that the system task can be efficiently carried out. We propose a model for sampling and quantization which both faithfully represents these operations while allowing the system to learn non-uniform mappings from training data. We focus on the task of symbol detection in multiple-input multiple-output (MIMO) digital receivers, where multiple analog signals are simultaneously acquired in order to recover a set of discrete information symbols. Our numerical results demonstrate that the proposed approach achieves performance which is comparable to operating without quantization constraints, while achieving more accurate digital representation compared to utilizing conventional uniform ADCs. Nir Shlezinger, Ruud van Sloun, Iris A. M. Huijben, Georgee Tsintsadze, Yonina C. Eldar |
ICASSP | 2 |
| 2020 | ProxSGD: Training Structured Neural Networks under Regularization and Constraints
Yang Yang 0033, Yaxiong Yuan, Avraam Chatzimichailidis, Ruud van Sloun, Lei Lei 0001, Symeon Chatzinotas |
ICLR | 4 |
| 2020 | Deep probabilistic subsampling for task-adaptive compressed sensing
Iris A. M. Huijben, Bastiaan S. Veeling, Ruud van Sloun |
ICLR | 3 |
| 2020 | Deep Learning in Ultrasound ImagingabstractIn this article, we consider deep learning strategies in ultrasound systems, from the front end to advanced applications. Our goal is to provide the reader with a broad understanding of the possible impact of deep learning methodologies on many aspects of ultrasound imaging. In particular, we discuss methods that lie at the interface of signal acquisition and machine learning, exploiting both data structure (e.g., sparsity in some domain) and data dimensionality (big data) already at the raw radio-frequency channel stage. As some examples, we outline efficient and effective deep learning solutions for adaptive beamforming and adaptive spectral Doppler through artificial agents, learn compressive encodings for the color Doppler, and provide a framework for structured signal recovery by learning fast approximations of iterative minimization problems, with applications to clutter suppression and super-resolution ultrasound. These emerging technologies may have a considerable impact on ultrasound imaging, showing promise across key components in the receive processing chain. Ruud van Sloun, Regev Cohen, Yonina C. Eldar |
Proc. IEEE | 1 |
| 2020 | Localizing B-Lines in Lung Ultrasonography by Weakly Supervised Deep Learning, In-Vivo ResultsabstractLung ultrasound (LUS) is nowadays gaining growing attention from both the clinical and technical world. Of particular interest are several imaging-artifacts, e.g., A- and B- line artifacts. While A-lines are a visual pattern which essentially represent a healthy lung surface, B-line artifacts correlate with a wide range of pathological conditions affecting the lung parenchyma. In fact, the appearance of B-lines correlates to an increase in extravascular lung water, interstitial lung diseases, cardiogenic and non-cardiogenic lung edema, interstitial pneumonia and lung contusion. Detection and localization of B-lines in a LUS video are therefore tasks of great clinical interest, with accurate, objective and timely evaluation being critical. This is particularly true in environments such as the emergency units, where timely decision may be crucial. In this work, we present and describe a method aimed at supporting clinicians by automatically detecting and localizing B-lines in an ultrasound scan. To this end, we employ modern deep learning strategies and train a fully convolutional neural network to perform this task on B-mode images of dedicated ultrasound phantoms in-vitro, and on patients in-vivo. An accuracy, sensitivity, specificity, negative and positive predictive value equal to 0.917, 0.915, 0.918, 0.950 and 0.864 were achieved in-vitro, respectively. Using a clinical system in-vivo, these statistics were 0.892, 0.871, 0.930, 0.798 and 0.958, respectively. We moreover calculate neural attention maps that visualize which components in the image triggered the network, thereby offering simultaneous weakly-supervised localization. These promising results confirm the capability of the proposed method to identify and localize the presence of B-lines in clinical lung ultrasonography. Ruud van Sloun, Libertario Demi |
IEEE J. Biomed. Health Informatics | 1 |
| 2020 | Learning Sub-Sampling and Signal Recovery With Applications in Ultrasound ImagingabstractLimitations on bandwidth and power consumption impose strict bounds on data rates of diagnostic imaging systems. Consequently, the design of suitable (i.e. task- and data-aware) compression and reconstruction techniques has attracted considerable attention in recent years. Compressed sensing emerged as a popular framework for sparse signal reconstruction from a small set of compressed measurements. However, typical compressed sensing designs measure a (non)linearly weighted combination of all input signal elements, which poses practical challenges. These designs are also not necessarily task-optimal. In addition, real-time recovery is hampered by the iterative and time-consuming nature of sparse recovery algorithms. Recently, deep learning methods have shown promise for fast recovery from compressed measurements, but the design of adequate and practical sensing strategies remains a challenge. Here, we propose a deep learning solution termed Deep Probabilistic Sub-sampling (DPS), that enables joint optimization of a task-adaptive sub-sampling pattern and a subsequent neural task model in an end-to-end fashion. Once learned, the task-based sub-sampling patterns are fixed and straightforwardly implementable, e.g. by non-uniform analog-to-digital conversion, sparse array design, or slow-time ultrasound pulsing schemes. The effectiveness of our framework is demonstrated in-silico for sparse signal recovery from partial Fourier measurements, and in-vivo for both anatomical image and tissue-motion (Doppler) reconstruction from sub-sampled medical ultrasound imaging data. Iris A. M. Huijben, Bastiaan S. Veeling, Kees Janse, Massimo Mischi, Ruud van Sloun |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Adaptive Ultrasound Beamforming Using Deep LearningabstractBiomedical imaging is unequivocally dependent on the ability to reconstruct interpretable and high-quality images from acquired sensor data. This reconstruction process is pivotal across many applications, spanning from magnetic resonance imaging to ultrasound imaging. While advanced data-adaptive reconstruction methods can recover much higher image quality than traditional approaches, their implementation often poses a high computational burden. In ultrasound imaging, this burden is significant, especially when striving for low-cost systems, and has motivated the development of high-resolution and high-contrast adaptive beamforming methods. Here we show that deep neural networks, that adopt the algorithmic structure and constraints of adaptive signal processing techniques, can efficiently learn to perform fast high-quality ultrasound beamforming using very little training data. We apply our technique to two distinct ultrasound acquisition strategies (plane wave, and synthetic aperture), and demonstrate that high image quality can be maintained when measuring at low data-rates, using undersampled array designs. Beyond biomedical imaging, we expect that the proposed deep learning based adaptive processing framework can benefit a variety of array and signal processing applications, in particular when data-efficiency and robustness are of importance. Ben Luijten, Regev Cohen, Frederik J. de Bruijn, Harold A. W. Schmeitz, Massimo Mischi, Yonina C. Eldar, Ruud van Sloun |
IEEE Trans. Medical Imaging | 7 |
| 2020 | Deep Learning for Classification and Localization of COVID-19 Markers in Point-of-Care Lung UltrasoundabstractDeep learning (DL) has proved successful in medical imaging and, in the wake of the recent COVID-19 pandemic, some works have started to investigate DL-based solutions for the assisted diagnosis of lung diseases. While existing works focus on CT scans, this paper studies the application of DL techniques for the analysis of lung ultrasonography (LUS) images. Specifically, we present a novel fully-annotated dataset of LUS images collected from several Italian hospitals, with labels indicating the degree of disease severity at a frame-level, video-level, and pixel-level (segmentation masks). Leveraging these data, we introduce several deep models that address relevant tasks for the automatic analysis of LUS images. In particular, we present a novel deep network, derived from Spatial Transformer Networks, which simultaneously predicts the disease severity score associated to a input frame and provides localization of pathological artefacts in a weakly-supervised way. Furthermore, we introduce a new method based on uninorms for effective frame score aggregation at a video-level. Finally, we benchmark state of the art deep models for estimating pixel-level segmentations of COVID-19 imaging biomarkers. Experiments on the proposed dataset demonstrate satisfactory results on all the considered tasks, paving the way to future research on DL for the assisted diagnosis of COVID-19 from LUS data. Subhankar Roy, Willi Menapace, Sebastiaan Oei, Ben Luijten, Enrico Fini, Cristiano Saltori, Iris A. M. Huijben, Nishith Chennakeshava, Federico Mento, Alessandro Sentelli, Emanuele Peschiera, Riccardo Trevisan, Giovanni Maschietto, Elena Torri, Riccardo Inchingolo, Andrea Smargiassi, Gino Soldati, Paolo Rota, Andrea Passerini, Ruud van Sloun, Elisa Ricci 0001, Libertario Demi |
IEEE Trans. Medical Imaging | 20 |
| 2020 | Deep Unfolded Robust PCA With Application to Clutter Suppression in UltrasoundabstractContrast enhanced ultrasound is a radiation-free imaging modality which uses encapsulated gas microbubbles for improved visualization of the vascular bed deep within the tissue. It has recently been used to enable imaging with unprecedented subwavelength spatial resolution by relying on super-resolution techniques. A typical preprocessing step in super-resolution ultrasound is to separate the microbubble signal from the cluttering tissue signal. This step has a crucial impact on the final image quality. Here, we propose a new approach to clutter removal based on robust principle component analysis (PCA) and deep learning. We begin by modeling the acquired contrast enhanced ultrasound signal as a combination of low rank and sparse components. This model is used in robust PCA and was previously suggested in the context of ultrasound Doppler processing and dynamic magnetic resonance imaging. We then illustrate that an iterative algorithm based on this model exhibits improved separation of microbubble signal from the tissue signal over commonly practiced methods. Next, we apply the concept of deep unfolding to suggest a deep network architecture tailored to our clutter filtering problem which exhibits improved convergence speed and accuracy with respect to its iterative counterpart. We compare the performance of the suggested deep network on both simulations and in-vivo rat brain scans, with a commonly practiced deep-network architecture and with the fast iterative shrinkage algorithm. We show that our architecture exhibits better image quality and contrast. Oren Solomon, Regev Cohen, Yi Zhang 0117, Yi Yang 0045, Qiong He, Jianwen Luo 0001, Ruud van Sloun, Yonina C. Eldar |
IEEE Trans. Medical Imaging | 7 |
| 2019 | Deep Convolutional Robust PCA with Application to Ultrasound ImagingabstractSparse and low-rank decomposition, also known as robust principle component analysis, has been applied successfully in numerous applications. Typically, this approach leads to a minimization problem which is solved using iterative algorithms. Drawing inspiration from recurrent networks, in recent years deep-learning strategies have been extended to mimic the behavior of iterative algorithms, with reduced complexity. In this work, we propose an extension of these deep architectures to robust principle component analysis in which fully-connected layers are replaced with convolutional ones. This strategy offers spatial invariance and significant reduction in the number of learned parameters. We then apply the proposed method to contrast-enhanced ultrasound, in which low-rank tissue signal needs to be removed in order to visualize blood vessels. We demonstrate the effectiveness of our approach on simulations and in-vivo rat brain scans. The resulting images exhibit improved visual quality and contrast compared with images obtained by commonly practiced methods. Regev Cohen, Yi Zhang 0117, Oren Solomon, Daniel Toberman, Liran Taieb, Ruud van Sloun, Yonina C. Eldar |
ICASSP | 6 |
| 2019 | Deep Learning for Fast Adaptive BeamformingabstractThe real-time nature that makes diagnostic ultrasonography so appealing to clinicians imposes strong constraints on the computational complexity of image reconstruction algorithms. As such, these typically rely on traditional delay-and-sum beamforming, a low-complexity approach that unfortunately comes at the cost of reduced image quality as compared to more advanced and content-adaptive beamformers. Here, we propose a model-aware deep learning strategy to ultrasound image reconstruction, which leverages knowledge of minimum variance beamforming while exploiting the efficiency of deep neural networks. Our approach yields high quality images with strong contrast at real-time reconstruction rates. The neural network is trained using in vivo and simulated radio frequency channel data of a single plane wave transmit, and corresponding high-quality minimum-variance beamformed reconstructions. Performance is benchmarked using simulated acquisitions from the PICMUS [1] dataset, demonstrating the convincing generalizability and image quality of the proposed beamformer. Ben Luijten, Regev Cohen, Frederik J. de Bruijn, Harold A. W. Schmeitz, Massimo Mischi, Yonina C. Eldar, Ruud van Sloun |
ICASSP | 7 |
| 2019 | Deep Learning for Super-resolution Vascular Ultrasound ImagingabstractBased on the intravascular infusion of gas microbubbles, which act as ultrasound contrast agents, ultrasound localization microscopy has enabled super resolution vascular imaging through precise detection of individual microbubbles across numerous imaging frames. However, analysis of high-density regions with significant overlaps among the microbubble point spread functions typically yields high localization errors, constraining the technique to low-concentration conditions. As such, long acquisition times are required for sufficient coverage of the vascular bed. Algorithms based on sparse recovery have been developed specifically to cope with the overlapping point-spread-functions of multiple microbubbles. While successful localization of densely-spaced emitters has been demonstrated, even highly optimized fast sparse recovery techniques involve a time-consuming iterative procedure. In this work, we used deep learning to improve upon standard ultrasound localization microscopy (Deep-ULM), and obtain super-resolution vascular images from high-density contrast-enhanced ultrasound data. Deep-ULM is suitable for real-time applications, resolving about 1250 high-resolution patches (128×128 pixels) per second using GPU acceleration. Ruud van Sloun, Oren Solomon, Matthew Bruce, Zin Z. Khaing, Yonina C. Eldar, Massimo Mischi |
ICASSP | 1 |
| 2019 | Super-resolution Using Flow Estimation in Contrast Enhanced Ultrasound ImagingabstractUltrasound localization microscopy offers new radiation-free diagnostic tools for vascular imaging deep within the tissue. Despite its high spatial resolution, low microbubble concentrations dictate the acquisition of tens of thousands of images, over the course of several seconds to tens of seconds, to produce a single super-resolved image. To address this limitation, sparsity-based approaches have recently been proposed to significantly reduce the total acquisition time, by resolving the vasculature in settings with considerable microbubble overlap. Here, we report on initial results of improving the spatial resolution and visual vascular reconstruction quality of sparsity-based super-resolution ultrasound imaging from low frame-rate acquisitions, by exploiting the inherent kinematics of microbubbles' flow. Our method relies on simultaneous tracking and sparsity-based detection of individual microbubbles. Oren Solomon, Ruud van Sloun, Massimo Mischi, Yonina C. Eldar |
ICASSP | 2 |
| 2018 | Convective-Dispersion Modeling in 3D Contrast-Ultrasound Imaging for the Localization of Prostate CancerabstractDespite being the solid tumor with the highest incidence in western men, prostate cancer (PCa) still lacks reliable imaging solutions that can overcome the need for systematic biopsies. Dynamic contrast-enhanced ultrasound imaging (DCE-US) allows us to quantitatively characterize the vascular bed in the prostate, due to its ability to visualize an intravenously administered bolus of contrast agents. Previous research has demonstrated that DCE-US parameters related to the vascular architecture are useful markers for the localization of PCa lesions. In this paper, we propose a novel method to assess the convective dispersion (D) and velocity (v) of the contrast bolus spreading through the prostate from three-dimensional (3D) DCE-US recordings. By assuming that D and v are locally constant, we solve the convective-dispersion equation by minimizing the corresponding regularized least-squares problem. 3D multiparametric maps of D and v were compared with 3D histopathology retrieved from the radical prostatectomy specimens of six patients. With a pixel-wise area under the receiver operating characteristic curve of 0.72 and 0.80, respectively, the method shows diagnostic value for the localization of PCa. Rogier R. Wildeboer, Ruud van Sloun, Stefan G. Schalk, Christophe K. Mannaerts, J. C. Van Der Linden, Pintong Huang, Hessel Wijkstra, Massimo Mischi |
IEEE Trans. Medical Imaging | 2 |
| 2017 | Ultrasound-contrast-agent dispersion and velocity imaging for prostate cancer localization
Ruud van Sloun, Libertario Demi, Arnoud Postema, Jean J. M. C. H. de la Rosette, Hessel Wijkstra, Massimo Mischi |
Medical Image Anal. | 1 |
| 2017 | Entropy of Ultrasound-Contrast-Agent Velocity Fields for Angiogenesis Imaging in Prostate CancerabstractProstate cancer care can benefit from accurate and cost-efficient imaging modalities that are able to reveal prognostic indicators for cancer. Angiogenesis is known to play a central role in the growth of tumors towards a metastatic or a lethal phenotype. With the aim of localizing angiogenic activity in a non-invasive manner, Dynamic Contrast Enhanced Ultrasound (DCE-US) has been widely used. Usually, the passage of ultrasound contrast agents thought the organ of interest is analyzed for the assessment of tissue perfusion. However, the heterogeneous nature of blood flow in angiogenic vasculature hampers the diagnostic effectiveness of perfusion parameters. In this regard, quantification of the heterogeneity of flow may provide a relevant additional feature for localizing angiogenesis. Statistics based on flow magnitude as well as its orientation can be exploited for this purpose. In this paper, we estimate the microbubble velocity fields from a standard bolus injection and provide a first statistical characterization by performing a spatial entropy analysis. By testing the method on 24 patients with biopsy-proven prostate cancer, we show that the proposed method can be applied effectively to clinically acquired DCE-US data. The method permits estimation of the in-plane flow vector fields and their local intricacy, and yields promising results (receiver-operating-characteristic curve area of 0.85) for the detection of prostate cancer. Ruud van Sloun, Libertario Demi, Arnoud Postema, Jean J. M. C. H. de la Rosette, Hessel Wijkstra, Massimo Mischi |
IEEE Trans. Medical Imaging | 1 |