EDBT 2026 Demo / reviewers in the wild / expert
Carola-Bibiane Schönlieb
dblp:07/8184 · also Carola B. Schönlieb, Carola Schönlieb, Carola-Bibiane Schoenlieb
· DBLP profile ↗
105ranked-venue papers
0as first author
74since 2021 · last 2026
0000-0003-0099-6306ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 35 since 2021Artificial intelligence and machine learning · 42 · 35 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 19 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Blessing of Dimensionality for Approximating Sobolev Classes on ManifoldsabstractThe manifold hypothesis says that natural high-dimensional data lie on or around a low-dimensional manifold. The recent success of statistical and learning-based methods in very high dimensions empirically supports this hypothesis, suggesting that typical worst-case analysis does not provide practical guarantees. A natural step for analysis is thus to assume the manifold hypothesis and derive bounds that are independent of any ambient dimensions that the data may be embedded in. Theoretical implications in this direction have recently been explored in terms of generalization of ReLU networks and convergence of Langevin methods. In this work, we consider optimal uniform approximations with functions of finite statistical complexity. While upper bounds on uniform approximation exist in the literature using ReLU neural networks, we consider the opposite: lower bounds to quantify the fundamental difficulty of approximation on manifolds. In particular, we demonstrate that the statistical complexity required to approximate a class of bounded Sobolev functions on a compact manifold is bounded from below, and moreover that this bound is dependent only on the intrinsic properties of the manifold, such as curvature, volume, and injectivity radius. Hong Ye Tan, Subhadip Mukherjee, Junqi Tang, Carola-Bibiane Schönlieb |
AAAI | 4 |
| 2026 | Brain foundation models with hypergraph dynamic adapter for brain disease analysisabstractBrain diseases, such as Alzheimer’s disease and brain tumors, present profound challenges due to their complexity and societal impact. Recent advancements in brain foundation models have shown significant promise in addressing a range of brain-related tasks. However, current brain foundation models are limited by task and data homogeneity, restricted generalization beyond segmentation or classification, and inefficient adaptation to diverse clinical tasks. In this work, we propose SAM-Brain3D, a brain-specific foundation model trained on over 66,000 brain image-label pairs across 14 MRI sub-modalities, and Hypergraph Dynamic Adapter (HyDA), a lightweight adapter for efficient and effective downstream adaptation. SAM-Brain3D captures detailed brain-specific anatomical and modality priors for segmenting diverse brain targets and broader downstream tasks. HyDA leverages hypergraphs to fuse complementary multi-modal data and dynamically generate patient-specific convolutional kernels for multi-scale feature fusion and personalized patient-wise adaptation. Together, our framework excels across a broad spectrum of brain disease segmentation and classification tasks. Extensive experiments demonstrate that our method consistently outperforms existing state-of-the-art approaches, offering a new paradigm for brain disease analysis through multi-modal, multi-scale, and dynamic foundation modeling. Zhongying Deng, Ziyan Huang, Lipei Zhang, Angelica I. Avilés-Rivero, Chaoyu Liu, Junjun He, Zoe Kourtzi, Carola-Bibiane Schönlieb |
Pattern Recognit. | 9 |
| 2026 | HIBMatch: Hypergraph Information Bottleneck for Semi-Supervised Alzheimer's ProgressionabstractAlzheimer's disease progression prediction is critical for patients with early Mild Cognitive Impairment (MCI) to enable timely intervention and improve their quality of life. While existing progression prediction techniques demonstrate potential with multimodal data, they are highly limited by their reliance on labelled data and fail to account for a key element of future progression prediction: not all features extracted at the current moment may be relevant for predicting progression several years later. To address these limitations in the literature, we design a novel semi-supervised multimodal learning hypergraph architecture, termed HIBMatch, by harnessing hypergraph knowledge based on information bottleneck and consistency regularisation strategies. Firstly, our framework utilises hypergraphs to represent multimodal data, encompassing both imaging and non-imaging modalities. Secondly, to harmonise relevant information from the currently captured data for future MCI conversion prediction, we propose a Hypergraph Information Bottleneck (HIB) that discriminates against irrelevant information, thereby focusing exclusively on harmonising relevant information for future MCI conversion prediction. Thirdly, our method enforces consistency regularisation between the HIB and a discriminative classifier to enhance the robustness and generalisation capabilities of HIBMatch under both topological and feature perturbations. Finally, to fully exploit the unlabeled data, HIBMatch incorporates a cross-modal contrastive loss for data efficiency. Extensive experiments on the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset demonstrate that our proposed HIBMatch framework surpasses existing state-of-the-art methods in Alzheimer's disease prognosis. Zhongying Deng, Angelica I. Avilés-Rivero, Zoe Kourtzi, Carola-Bibiane Schönlieb |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | Training-Free Dual Hyperbolic Adapters for Better Cross-Modal ReasoningabstractRecent research in Vision-Language Models (VLMs) has significantly advanced our capabilities in cross-modal reasoning. However, existing methods suffer from performance degradation with domain changes or require substantial computational resources for fine-tuning in new domains. To address this issue, we develop a new adaptation method for large vision-language models, calledTraining-free Dual Hyperbolic Adapters(T-DHA). We characterize vision-language relationship between semantic concepts, which typically has a hierarchical tree structure, in the hyperbolic space instead of the traditional Euclidean space. Hyperbolic spaces exhibit exponential volume growth with radius, unlike the polynomial growth in Euclidean space. We find that this unique property is particularly effective for embedding hierarchical data structures using the Poincaré ball model, achieving significantly improved representation and discrimination power. Coupled with negative learning, it provides more accurate and robust classifications with fewer feature dimensions. Our extensive experimental results on various datasets demonstrate that the T-DHA method significantly outperforms existing state-of-the-art methods in few-shot image recognition and domain generalization tasks. Yi Zhang 0109, Chun-Wun Cheng, Ke Yu 0004, Yushun Tang, Carola-Bibiane Schönlieb, Zhihai He, Angelica I. Avilés-Rivero |
IEEE Trans. Multim. | 6 |
| 2025 | Learning Regularization for Graph Inverse ProblemsabstractIn recent years, Graph Neural Networks (GNNs) have been utilized for various applications ranging from drug discovery to network design and social networks. In many applications, it is impossible to observe some properties of the graph directly; instead, noisy and indirect measurements of these properties are available. These scenarios are coined as Graph Inverse Problems (GRIPs). In this work, we introduce a framework leveraging GNNs to solve GRIPs. The framework is based on a combination of likelihood and prior terms, which are used to find a solution that fits the data while adhering to learned prior information. Specifically, we propose to combine recent deep learning techniques that were developed for inverse problems, together with GNN architectures, to formulate and solve GRIPs. We study our approach on a number of representative problems that demonstrate the effectiveness of the framework. Moshe Eliasof, Md Shahriar Rahim Siddiqui, Carola-Bibiane Schönlieb, Eldad Haber |
AAAI | 3 |
| 2025 | On Oversquashing in Graph Neural Networks Through the Lens of Dynamical SystemsabstractA common problem in Message-Passing Neural Networks is oversquashing -- the limited ability to facilitate effective information flow between distant nodes. Oversquashing is attributed to the exponential decay in information transmission as node distances increase. This paper introduces a novel perspective to address oversquashing, leveraging dynamical systems properties of global and local non-dissipativity, that enable the maintenance of a constant information flow rate. We present SWAN, a uniquely parameterized GNN model with antisymmetry both in space and weight domains, as a means to obtain non-dissipativity. Our theoretical analysis asserts that by implementing these properties, SWAN offers an enhanced ability to transmit information over extended distances. Empirical evaluations on synthetic and real-world benchmarks that emphasize long-range interactions validate the theoretical understanding of SWAN, and its ability to mitigate oversquashing. Alessio Gravina, Moshe Eliasof, Claudio Gallicchio, Davide Bacciu, Carola-Bibiane Schönlieb |
AAAI | 5 |
| 2025 | Cross-Modal Few-Shot Learning with Second-Order Neural Ordinary Differential EquationsabstractWe introduce SONO, a novel method leveraging Second-Order Neural Ordinary Differential Equations (Second-Order NODEs) to enhance cross-modal few-shot learning. By employing a simple yet effective architecture consisting of a Second-Order NODEs model paired with a cross-modal classifier, SONO addresses the significant challenge of overfitting, which is common in few-shot scenarios due to limited training examples. Our second-order approach can approximate a broader class of functions, enhancing the model's expressive power and feature generalization capabilities. We initialize our cross-modal classifier with text embeddings derived from class-relevant prompts, streamlining training efficiency by avoiding the need for frequent text encoder processing. Additionally, we utilize text-based image augmentation, exploiting CLIP’s robust image-text correlation to enrich training data significantly. Extensive experiments across multiple datasets demonstrate that SONO outperforms existing state-of-the-art methods in few-shot learning performance. Yi Zhang 0109, Chun-Wun Cheng, Zhihai He, Carola-Bibiane Schönlieb, Yuyan Chen, Angelica I. Avilés-Rivero |
AAAI | 5 |
| 2025 | DiTASK: Multi-Task Fine-Tuning with Diffeomorphic TransformationsabstractPre-trained Vision Transformers now serve as powerful tools for computer vision. Yet, efficiently adapting them for multiple tasks remains a challenge that arises from the need to modify the rich hidden representations encoded by the learned weight matrices, without inducing interference between tasks. Current parameter-efficient methods like LoRA, which apply low-rank updates, force tasks to compete within constrained subspaces, ultimately degrading performance. We introduce DiTASK a novel Diffeomorphic Multi-Task Fine-Tuning approach that maintains pre-trained representations by preserving weight matrix singular vectors, while enabling task-specific adaptations through neural diffeomorphic transformations of the singular values. By following this approach, DiTASK enables both shared and task-specific feature modulations with minimal added parameters. Our theoretical analysis shows that DiTASK achieves full-rank updates during optimization, preserving the geometric structure of pretrained features, and establishing a new paradigm for efficient multi-task learning (MTL). Our experiments on PASCAL MTL and NYUD show that DiTASK achieves state-of-the-art performance across four dense prediction tasks, using 75% fewer parameters than existing methods. Our code is available here. Krishna Sri Ipsit Mantri, Carola-Bibiane Schönlieb, Bruno Ribeiro 0001, Chaim Baskin, Moshe Eliasof |
CVPR | 2 |
| 2025 | Iterative Operator Sketching Framework for Large-Scale Imaging Inverse ProblemsabstractDespite impressive empirical performance in various imaging applications, iterative data-driven reconstruction (IDR) schemes such as plug-and-play algorithms and deep unrolling networks can have significant computational limitations, especially for large-scale imaging inverse problems. This is mostly because they need to involve the high-dimensional forward/adjoint operators that are expensive to compute in each iteration. In this work, we propose a new operator sketching framework tailored for designing efficient IDR schemes, which are currently state-of-the-art solutions for imaging inverse problems. Our framework performs dimensionality reduction in both image and measurement data domains, leading to efficient computations. Using this framework, we derive several accelerated IDR schemes, such as the plug-and-play multi-stage sketched gradient (PnP-MS2G) and sketching-based primal-dual (LSPD and Sk-LSPD) deep unrolling networks. Our experiments on X-ray CT image reconstruction demonstrate the remarkable effectiveness of the proposed sketched IDR methods. Junqi Tang, Subhadip Mukherjee, Carola-Bibiane Schönlieb |
ICASSP | 3 |
| 2025 | Estimation of single-cell and tissue perturbation effect in spatial transcriptomics via Spatial Causal DisentanglementabstractModels of Virtual Cells and Virtual Tissues at single-cell resolution would allow us to test perturbations in silico and accelerate progress in tissue and cell engineering.
However, most such models are not rooted in causal inference and as a result, could mistake correlation for causation.
We introduce Celcomen, a novel generative graph neural network grounded in mathematical causality to disentangle intra- and inter-cellular gene regulation in spatial transcriptomics and single-cell data.
Celcomen can also be prompted by perturbations to generate spatial counterfactuals, thus offering insights into experimentally inaccessible states, with potential applications in human health.
We validate the model's disentanglement and identifiability through simulations, and demonstrate its counterfactual predictions in clinically relevant settings, including human glioblastoma and fetal spleen, recovering inflammation-related gene programs post immune system perturbation.
Moreover, it supports mechanistic interpretability, as its parameters can be reverse-engineered from observed behavior, making it an accessible model for understanding both neural networks and complex biological systems. Stathis Megas, Daniel G. Chen, Krzysztof Polanski, Moshe Eliasof, Carola-Bibiane Schönlieb, Sarah A. Teichmann |
ICLR | 5 |
| 2025 | Lie Algebra Canonicalization: Equivariant Neural Operators under arbitrary Lie GroupsabstractThe quest for robust and generalizable machine learning models has driven recent interest in exploiting symmetries through equivariant neural networks. In the context of PDE solvers, recent works have shown that Lie point symmetries can be a useful inductive bias for Physics-Informed Neural Networks (PINNs) through data and loss augmentation. Despite this, directly enforcing equivariance within the model architecture for these problems remains elusive. This is because many PDEs admit non-compact symmetry groups, oftentimes not studied beyond their infinitesimal generators, making them incompatible with most existing equivariant architectures. In this work, we propose Lie aLgebrA Canonicalization (LieLAC), a novel approach that exploits only the action of infinitesimal generators of the symmetry group, circumventing the need for knowledge of the full group structure. To achieve this, we address existing theoretical issues in the canonicalization literature, establishing connections with frame averaging in the case of continuous non-compact groups. Operating within the framework of canonicalization, LieLAC can easily be integrated with unconstrained pre-trained models, transforming inputs to a canonical form before feeding them into the existing model, effectively aligning the input for model inference according to allowed symmetries. LieLAC utilizes standard Lie group descent schemes, achieving equivariance in pre-trained models. Finally, we showcase LieLAC's efficacy on tasks of invariant image classification and Lie point symmetry equivariant neural PDE solvers using pre-trained models. Zakhar Shumaylov, Peter Zaika, James Rowbottom, Ferdia Sherry, Melanie Weber 0001, Carola-Bibiane Schönlieb |
ICLR | 6 |
| 2025 | Score-based Pullback Riemannian Geometry: Extracting the Data Manifold Geometry using Anisotropic FlowsabstractData-driven Riemannian geometry has emerged as a powerful tool for interpretable representation learning, offering improved efficiency in downstream tasks. Moving forward, it is crucial to balance cheap manifold mappings with efficient training algorithms. In this work, we integrate concepts from pullback Riemannian geometry and generative models to propose a framework for data-driven Riemannian geometry that is scalable in both geometry and learning: score-based pullback Riemannian geometry. Focusing on unimodal distributions as a first step, we propose a score-based Riemannian structure with closed-form geodesics that pass through the data probability density. With this structure, we construct a Riemannian autoencoder (RAE) with error bounds for discovering the correct data manifold dimension. This framework can naturally be used with anisotropic normalizing flows by adopting isometry regularization during training. Through numerical experiments on diverse datasets, including image data, we demonstrate that the proposed framework produces high-quality geodesics passing through the data support, reliably estimates the intrinsic dimension of the data manifold, and provides a global chart of the manifold. To the best of our knowledge, this is the first scalable framework for extracting the complete geometry of the data manifold. Willem Diepeveen, Georgios Batzolis, Zakhar Shumaylov, Carola-Bibiane Schönlieb |
ICML | 4 |
| 2025 | Graph Adaptive Autoregressive Moving Average ModelsabstractGraph State Space Models (SSMs) have recently been introduced to enhance Graph Neural Networks (GNNs) in modeling long-range interactions. Despite their success, existing methods either compromise on permutation equivariance or limit their focus to pairwise interactions rather than sequences. Building on the connection between Autoregressive Moving Average (ARMA) and SSM, in this paper, we introduce GRAMA, a Graph Adaptive method based on a learnable ARMA framework that addresses these limitations.
By transforming from static to sequential graph data, GRAMA leverages the strengths of the ARMA framework, while preserving permutation equivariance. Moreover, GRAMA incorporates a selective attention mechanism for dynamic learning of ARMA coefficients, enabling efficient and flexible long-range information propagation. We also establish theoretical connections between GRAMA and Selective SSMs, providing insights into its ability to capture long-range dependencies. Experiments on 26 synthetic and real-world datasets demonstrate that GRAMA consistently outperforms backbone models and performs competitively with state-of-the-art methods. Moshe Eliasof, Alessio Gravina, Andrea Ceni, Claudio Gallicchio, Davide Bacciu, Carola-Bibiane Schönlieb |
ICML | 6 |
| 2025 | G-Adaptivity: optimised graph-based mesh relocation for finite element methodsabstractWe present a novel, and effective, approach to achieve optimal mesh relocation in finite element methods (FEMs). The cost and accuracy of FEMs is critically dependent on the choice of mesh points. Mesh relocation (r-adaptivity) seeks to optimise the mesh geometry to obtain the best solution accuracy at given computational budget. Classical r-adaptivity relies on the solution of a separate nonlinear “meshing” PDE to determine mesh point locations. This incurs significant cost at remeshing, and relies on estimates that relate interpolation- and FEM-error. Recent machine learning approaches have focused on the construction of fast surrogates for such classical methods. Instead, our new approach trains a graph neural network (GNN) to determine mesh point locations by directly minimising the FE solution error from the PDE system Firedrake to achieve higher solution accuracy. Our GNN architecture closely aligns the mesh solution space to that of classical meshing methodologies, thus replacing classical estimates for optimality with a learnable strategy. This allows for rapid and robust training and results in an extremely efficient and effective GNN approach to online r-adaptivity. Our method outperforms both classical, and prior ML, approaches to r-adaptive meshing. In particular, it achieves lower FE solution error, whilst retaining the significant speed-up over classical methods observed in prior ML work. James Rowbottom, Georg Maierhofer, Teo Deveney, Eike Hermann Müller, Alberto Paganini, Katharina Schratz, Pietro Liò, Carola-Bibiane Schönlieb, Chris J. Budd |
ICML | 8 |
| 2025 | Implicit U-KAN2.0: Dynamic, Efficient and Interpretable Medical Image Segmentation
Chun-Wun Cheng, Yanqi Cheng, Javier A. Montoya-Zegarra, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero |
MICCAI (11) | 5 |
| 2025 | Return of ChebNet: Understanding and Improving an Overlooked GNN on Long Range TasksabstractChebNet, one of the earliest spectral GNNs, has largely been overshadowed by Message Passing Neural Networks (MPNNs), which gained popularity for their simplicity and effectiveness in capturing local graph structure. Despite their success, MPNNs are limited in their ability to capture long-range dependencies between nodes. This has led researchers to adapt MPNNs through *rewiring* or make use of *Graph Transformers*, which compromise the computational efficiency that characterized early spatial message passing architectures, and typically disregard the graph structure. Almost a decade after its original introduction, we revisit ChebNet to shed light on its ability to model distant node interactions. We find that out-of-box, ChebNet already shows competitive advantages relative to classical MPNNs and GTs on long-range benchmarks, while maintaining good scalability properties for high-order polynomials. However, we uncover that this polynomial expansion leads ChebNet to an unstable regime during training. To address this limitation, we cast ChebNet as a stable and non-dissipative dynamical system, which we coin Stable-ChebNet. Our Stable-ChebNet model allows for stable information propagation, and has controllable dynamics which do not require the use of eigendecompositions, positional encodings, or graph rewiring. Across several benchmarks, Stable-ChebNet achieves near state-of-the-art performance. Ali Hariri, Alvaro Arroyo, Alessio Gravina, Moshe Eliasof, Carola-Bibiane Schönlieb, Davide Bacciu, Xiaowen Dong 0001, Kamyar Azizzadenesheli, Pierre Vandergheynst |
NeurIPS | 5 |
| 2025 | Approximation theory for 1-Lipschitz ResNetsabstract$1$-Lipschitz neural networks are fundamental for generative modelling, inverse problems, and robust classifiers. In this paper, we focus on $1$-Lipschitz residual networks (ResNets) based on explicit Euler steps of negative gradient flows and study their approximation capabilities. Leveraging the Restricted Stone–Weierstrass Theorem, we first show that these $1$-Lipschitz ResNets are dense in the set of scalar $1$-Lipschitz functions on any compact domain when width and depth are allowed to grow. We also show that these networks can exactly represent scalar piecewise affine $1$-Lipschitz functions. We then prove a stronger statement: by inserting norm-constrained linear maps between the residual blocks, the same density holds when the hidden width is fixed. Because every layer obeys simple norm constraints, the resulting models can be trained with off-the-shelf optimisers. This paper provides the first universal approximation guarantees for $1$-Lipschitz ResNets, laying a rigorous foundation for their practical use. Davide Murari, Takashi Furuya, Carola-Bibiane Schönlieb |
NeurIPS | 3 |
| 2025 | D2SA: Dual-Stage Distribution and Slice Adaptation for Efficient Test-Time Adaptation in MRI ReconstructionabstractVariations in Magnetic resonance imaging (MRI) scanners and acquisition protocols cause distribution shifts that degrade reconstruction performance on unseen data. Test-time adaptation (TTA) offers a promising solution to address this discrepancies. However, previous single-shot TTA approaches are inefficient due to repeated training and suboptimal distributional models. Self-supervised learning methods may risk over-smoothing in scarce data scenarios. To address these challenges, we propose a novel Dual-Stage Distribution and Slice Adaptation (D2SA) via MRI implicit neural representation (MR-INR) to improve MRI reconstruction performance and efficiency, which features two stages. In the first stage, an MR-INR branch performs patient-wise distribution adaptation by learning shared representations across slices and modelling patient-specific shifts with mean and variance adjustments. In the second stage, single-slice adaptation refines the output from frozen convolutional layers with a learnable anisotropic diffusion module, preventing over-smoothing and reducing computation. Experiments across five MRI distribution shifts demonstrate that our method can integrate well with various self-supervised learning (SSL) framework, improving performance and accelerating convergence under diverse conditions. Lipei Zhang, Zhongying Deng, Yanqi Cheng, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero |
NeurIPS | 5 |
| 2025 | Artificial immunofluorescence in a flash: Rapid synthetic imaging from brightfield through residual diffusionabstractImmunofluorescent (IF) imaging is crucial for visualising biomarker expressions, cell morphology and assessing the effects of drug treatments on sub-cellular components. IF imaging needs extra staining process and often requiring cell fixation, therefore it may also introduce artefacts and alter endogenous cell morphology. Some IF stains are expensive or not readily available hence hindering experiments. Recent diffusion models, which synthesise high-fidelity IF images from easy-to-acquire brightfield (BF) images, offer a promising solution but are hindered by training instability and slow inference times due to the noise diffusion process. This paper presents a novel method for the conditional synthesis of IF images directly from BF images along with cell segmentation masks. Our approach employs a Residual Diffusion process that enhances stability and significantly reduces inference time. We performed a critical evaluation against other image-to-image synthesis models, including UNets, GANs, and advanced diffusion models. Our model demonstrates significant improvements in image quality ( p < 0 . 05 in MSE, PSNR, and SSIM), inference speed (26 times faster than competing diffusion models), and accurate segmentation results for both nuclei and cell bodies (0.77 and 0.63 mean IOU for nuclei and cell true positives, respectively). This paper is a substantial advancement in the field, providing robust and efficient tools for cell image analysis. • We introduce a novel diffusion model to synthesise fluorescence images from brightfield images. • CellResDM improves quality, speed, and segmentation accuracy, surpassing existing models. • CellResDM model can simultaneously generates IF images and cell/nuclei segmentation. Xiaodan Xing, Chunling Tang, Siofra Murdoch, Giorgos Papanastasiou, Yunzhe Guo, Xianglu Xiao, Jan Oscar Cross-Zamirski, Carola-Bibiane Schönlieb, Kristina Xiao Liang, Zhangming Niu, Evandro Fei Fang, Yinhai Wang, Guang Yang 0006 |
Neurocomputing | 8 |
| 2025 | Enhancing global sensitivity and uncertainty quantification in medical image reconstruction with Monte Carlo arbitrary-masked mambaabstractDeep learning has been extensively applied in medical image reconstruction, where Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) represent the predominant paradigms, each possessing distinct advantages and inherent limitations: CNNs exhibit linear complexity with local sensitivity, whereas ViTs demonstrate quadratic complexity with global sensitivity. The emerging Mamba has shown superiority in learning visual representation, which combines the advantages of linear scalability and global sensitivity. In this study, we introduce MambaMIR, an Arbitrary-Masked Mamba-based model with wavelet decomposition for joint medical image reconstruction and uncertainty estimation. A novel Arbitrary Scan Masking (ASM) mechanism "masks out" redundant information to introduce randomness for further uncertainty estimation. Compared to the commonly used Monte Carlo (MC) dropout, our proposed MC-ASM provides an uncertainty map without the need for hyperparameter tuning and mitigates the performance drop typically observed when applying dropout to low-level tasks. For further texture preservation and better perceptual quality, we employ the wavelet transformation into MambaMIR and explore its variant based on the Generative Adversarial Network, namely MambaMIR-GAN. Comprehensive experiments have been conducted for multiple representative medical image reconstruction tasks, demonstrating that the proposed MambaMIR and MambaMIR-GAN outperform other baseline and state-of-the-art methods in different reconstruction tasks, where MambaMIR achieves the best reconstruction fidelity and MambaMIR-GAN has the best perceptual quality. In addition, our MC-ASM provides uncertainty maps as an additional tool for clinicians, while mitigating the typical performance drop caused by the commonly used dropout. Liutao Yang, Fanwen Wang, Yinzhe Wu 0001, Yang Nan 0002, Weiwen Wu, Chengyan Wang, Kuangyu Shi, Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb, Daoqiang Zhang, Guang Yang 0006 |
Medical Image Anal. | 10 |
| 2025 | Learning homeomorphic image registration via conformal-invariant hyperelastic regularisation
Noémie Debroux, Harry Qin, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero |
Medical Image Anal. | 5 |
| 2025 | Nested Bregman Iterations for Decomposition ProblemsabstractAbstract. We consider the task of image reconstruction while simultaneously decomposing the reconstructed image into components with different features. A commonly used tool for this is a variational approach with an infimal convolution of appropriate functions as a regularizer. Especially for noise corrupted observations, incorporating these functionals into the classical method of Bregman iterations provides a robust method for obtaining an overall good approximation of the true image by stopping the iteration early according to a discrepancy principle. However, crucially, the quality of the separate components depends further on the proper choice of the regularization weights associated with the infimally convoluted functionals. Here, we propose the method of Nested Bregman iterations to improve a decomposition in a structured way. This allows for the transformation of the task of choosing the weights into the problem of stopping the iteration according to a meaningful criterion based on normalized cross-correlation. We discuss the well-definedness and the convergence behavior of the proposed method and illustrate its strength numerically with various image decomposition tasks employing infimal convolution functionals. Tobias Wolf, Derek Driggs, Kostas Papafitsoros, Elena Resmerita, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 5 |
| 2025 | ADFound: A Foundation Model for Diagnosis and Prognosis of Alzheimer's DiseaseabstractAlzheimer's disease (AD) is an incurable neurodegenerative disorder characterized by progressive cognitive and functional decline. Consequently, early diagnosis and accurate prediction of disease progression are of paramount importance and inherently complex, necessitating the integration of multi-modal data. While existing methods are typically task-specific and lack generalization, we present ADFound, the first multi-modal foundation model for AD capable of simultaneously addressing diagnosis and prognosis tasks through a unified framework. ADFound leverages a substantial amount of unlabeled 3D multi-modal neuroimaging, including paired and unpaired data, to achieve its objectives. Specifically, ADFound is developed upon the Multi-modal Vim encoder by Vision Mamba block to capture long-range dependencies inherent in 3D multi-modal medical images. To efficiently pre-train ADFound on unlabeled paired and upaired multi-modal neuroimaging data, we proposed a novel self-supervised learning framework that integrates multi-modal masked autoencoder (MAE) and contrastive learning. The multi-modal MAE aims to learn local relations among modalities by reconstructing images with unmasked image patches. Additionally, we introduce a Dual Contrastive Learning for Multi-modal Data to enhance the discriminative capabilities of multi-modal representations from intra-modal and inter-modal perspectives. Our experiments demonstrate that ADFound outperforms state-of-the-art methods across a wide range of downstream tasks relevant to the diagnosis and prognosis of AD. Furthermore, the results indicate that our foundation model can be extended to more modalities, such as non-image data, showing its versatility. Guangqian Yang, Kangrui Du, Zhihan Yang 0001, Ye Du 0002, Eva Yi Wah Cheung, Zoe Kourtzi, Carola-Bibiane Schönlieb |
IEEE J. Biomed. Health Informatics | 9 |
| 2025 | TrafficCAM: A Versatile Dataset for Traffic Flow SegmentationabstractTraffic flow analysis is revolutionising traffic management. By leveraging traffic flow data, traffic control bureaus could provide drivers with real-time alerts, advising the fastest routes and therefore optimising transportation logistics and reducing congestion. The existing traffic flow datasets have two major limitations. They feature a limited number of classes, usually limited to one type of vehicle, and the scarcity of unlabelled data. In this paper, we introduce a new benchmark traffic flow image dataset called TrafficCAM. Our dataset distinguishes itself by two major highlights. Firstly, TrafficCAM provides both pixel-level and instance-level semantic labelling along with a large range of types of vehicles and pedestrians. It is composed of a large and diverse set of video sequences recorded in streets from eight Indian cities with stationary cameras. Secondly, TrafficCAM aims to establish a new benchmark for developing fully-supervised tasks, and importantly, semi-supervised learning techniques. It is the first dataset that provides a vast amount of unlabelled data, helping to better capture traffic flow qualification under a low-cost annotation requirement. More precisely, our dataset has 4,364 image frames with semantic and instance annotations along with 58,689 unlabelled image frames. We validate our new dataset through a large and comprehensive range of experiments on several state-of-the-art approaches under four different settings: fully-supervised semantic and instance segmentation, and semi-supervised semantic and instance segmentation tasks. Our benchmark dataset and official toolkit are released athttps://math-ml-x.github.io/TrafficCAM/. Zhongying Deng, Yanqi Cheng, Rihuan Ke, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Contrastive Registration for Unsupervised Medical Image SegmentationabstractMedical image segmentation is an important task in medical imaging, as it serves as the first step for clinical diagnosis and treatment planning. While major success has been reported using deep learning supervised techniques, they assume a large and well-representative labeled set. This is a strong assumption in the medical domain where annotations are expensive, time-consuming, and inherent to human bias. To address this problem, unsupervised segmentation techniques have been proposed in the literature. Yet, none of the existing unsupervised segmentation techniques reach accuracies that come even near to the state-of-the-art of supervised segmentation methods. In this work, we present a novel optimization model framed in a new convolutional neural network (CNN)-based contrastive registration architecture for unsupervised medical image segmentation called CLMorph. The core idea of our approach is to exploit image-level registration and feature-level contrastive learning, to perform registration-based segmentation. First, we propose an architecture to capture the image-to-image transformation mapping via registration for unsupervised medical image segmentation. Second, we embed a contrastive learning mechanism in the registration architecture to enhance the discriminative capacity of the network at the feature level. We show that our proposed CLMorph technique mitigates the major drawbacks of existing unsupervised techniques. We demonstrate, through numerical and visual experiments, that our technique substantially outperforms the current state-of-the-art unsupervised segmentation methods on two major medical image datasets. Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | On The Temporal Domain of Differential Equation Inspired Graph Neural NetworksabstractGraph Neural Networks (GNNs) have demonstrated remarkable success in modeling complex relationships in graph-structured data. A recent innovation in this field is the family of Differential Equation-Inspired Graph Neural Networks (DE-GNNs), which leverage principles from continuous dynamical systems to model information flow on graphs with built-in properties such as feature smoothing or preservation. However, existing DE-GNNs rely on first or second-order temporal dependencies. In this paper, we propose a neural extension to those pre-defined temporal dependencies. We show that our model, called TDE-GNN, can capture a wide range of temporal dynamics that go beyond typical first or second-order methods, and provide use cases where existing temporal models are challenged. We demonstrate the benefit of learning the temporal dependencies using our method rather than using pre-defined temporal dynamics on several graph benchmarks. Moshe Eliasof, Eldad Haber, Eran Treister, Carola-Bibiane Schönlieb |
AISTATS | 4 |
| 2024 | Resilient Graph Neural Networks: A Coupled Dynamical Systems ApproachabstractGraph Neural Networks (GNNs) have established themselves as a key component in addressing diverse graph-based tasks. Despite their notable successes, GNNs remain susceptible to input perturbations in the form of adversarial attacks. This paper introduces an innovative approach to fortify GNNs against adversarial perturbations through the lens of coupled dynamical systems. Our method introduces graph neural layers based on differential equations with contractive properties, which, as we show, improve the robustness of GNNs. A distinctive feature of the proposed approach is the simultaneous learned evolution of both the node features and the adjacency matrix, yielding an intrinsic enhancement of model robustness to perturbations in the input features and the connectivity of the graph. We mathematically derive the underpinnings of our novel architecture and provide theoretical insights to reason about its expected behavior. We demonstrate the efficacy of our method through numerous real-world benchmarks, reading on par or improved performance compared to existing methods. Moshe Eliasof, Davide Murari, Ferdia Sherry, Carola-Bibiane Schönlieb |
ECAI | 4 |
| 2024 | Data-Driven Convex Regularizers for Inverse ProblemsabstractWe propose to learn a data-adaptive convex regularizer, which is parameterized using an input-convex neural network (ICNN), for variational image reconstruction. The regularizer parameters are learned adversarially by telling apart clean images from the artifact-ridden ones in a training dataset. Convexity of the regularizer is theoretically and practically important since (i) one can establish well-posedness guarantees for the corresponding variational reconstruction problem and (ii) devise provably convergent optimization algorithms for reconstruction. In particular, the resulting method is shown to be convergent in the sense of regularization and can be solved provably using a gradient-based solver. To demonstrate the performance of our approach for solving inverse problems, we consider deblurring natural images and reconstruction in X-ray computed tomography (CT) and show that the proposed convex regularizer is on par with and sometimes superior to state-of-the-art classical and data-driven techniques for inverse problems, especially with severely ill-posed forward operators (such as in limited-angle tomography). Subhadip Mukherjee, Sören Dittmer, Zakhar Shumaylov, Sebastian Lunz, Ozan Öktem, Carola-Bibiane Schönlieb |
ICASSP | 6 |
| 2024 | HAMLET: Graph Transformer Neural Operator for Partial Differential EquationsabstractWe present a novel graph transformer framework, HAMLET, designed to address the challenges in solving partial differential equations (PDEs) using neural networks. The framework uses graph transformers with modular input encoders to directly incorporate differential equation information into the solution process. This modularity enhances parameter correspondence control, making HAMLET adaptable to PDEs of arbitrary geometries and varied input formats. Notably, HAMLET scales effectively with increasing data complexity and noise, showcasing its robustness. HAMLET is not just tailored to a single type of physical simulation, but can be applied across various domains. Moreover, it boosts model resilience and performance, especially in scenarios with limited data. We demonstrate, through extensive experiments, that our framework is capable of outperforming current techniques for PDEs. Andrey Bryutkin, Zhongying Deng, Guang Yang 0006, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero |
ICML | 5 |
| 2024 | Weakly Convex Regularisers for Inverse Problems: Convergence of Critical Points and Primal-Dual OptimisationabstractVariational regularisation is the primary method for solving inverse problems, and recently there has been considerable work leveraging deeply learned regularisation for enhanced performance. However, few results exist addressing the convergence of such regularisation, particularly within the context of critical points as opposed to global minimisers. In this paper, we present a generalised formulation of convergent regularisation in terms of critical points, and show that this is achieved by a class of weakly convex regularisers. We prove convergence of the primal-dual hybrid gradient method for the associated variational problem, and, given a Kurdyka-Łojasiewicz condition, an $\mathcal{O}(\log{k}/k)$ ergodic convergence rate. Finally, applying this theory to learned regularisation, we prove universal approximation for input weakly convex neural networks (IWCNN), and show empirically that IWCNNs can lead to improved performance of learned adversarial regularisers for computed tomography (CT) reconstruction. Zakhar Shumaylov, Jeremy Budd, Subhadip Mukherjee, Carola-Bibiane Schönlieb |
ICML | 4 |
| 2024 | Diffusion Models Encode the Intrinsic Dimension of Data ManifoldsabstractIn this work, we provide a mathematical proof that diffusion models encode data manifolds by approximating their normal bundles. Based on this observation we propose a novel method for extracting the intrinsic dimension of the data manifold from a trained diffusion model. Our insights are based on the fact that a diffusion model approximates the score function i.e. the gradient of the log density of a noise-corrupted version of the target distribution for varying levels of corruption. We prove that as the level of corruption decreases, the score function points towards the manifold, as this direction becomes the direction of maximal likelihood increase. Therefore, at low noise levels, the diffusion model provides us with an approximation of the manifold’s normal bundle, allowing for an estimation of the manifold’s intrinsic dimension. To the best of our knowledge our method is the first estimator of intrinsic dimension based on diffusion models and it outperforms well established estimators in controlled experiments on both Euclidean and image data. Jan Stanczuk, Georgios Batzolis, Teo Deveney, Carola-Bibiane Schönlieb |
ICML | 4 |
| 2024 | Spatiotemporal Graph Neural Network Modelling Perfusion MRI
Ruodan Yan, Carola-Bibiane Schönlieb, Chao Li 0031 |
MICCAI (2) | 2 |
| 2024 | Biophysics Informed Pathological Regularisation for Brain Tumour Segmentation
Lipei Zhang, Yanqi Cheng, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero |
MICCAI (12) | 4 |
| 2024 | TrafficMOT: A Challenging Dataset for Multi-Object Tracking in Complex Traffic ScenariosabstractACM Multimedia 2024, Melbourne, Australia, Oct 28 - Nov 1, 2024 Yanqi Cheng, Zhongying Deng, Dongdong Chen 0001, Xiaowei Hu 0001, Pietro Liò, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero |
ACM Multimedia | 8 |
| 2024 | GRANOLA: Adaptive Normalization for Graph Neural NetworksabstractDespite the widespread adoption of Graph Neural Networks (GNNs), these models often incorporate off-the-shelf normalization layers like BatchNorm or InstanceNorm, which were not originally designed for GNNs. Consequently, these normalization layers may not effectively capture the unique characteristics of graph-structured data, potentially even weakening the expressive power of the overall architecture.
While existing graph-specific normalization layers have been proposed, they often struggle to offer substantial and consistent benefits. In this paper, we propose GRANOLA, a novel graph-adaptive normalization layer. Unlike existing normalization layers, GRANOLA normalizes node features by adapting to the specific characteristics of the graph, particularly by generating expressive representations of its nodes, obtained by leveraging the propagation of Random Node Features (RNF) in the graph. We provide theoretical results that support our design choices as well as an extensive empirical evaluation demonstrating the superior performance of GRANOLA over existing normalization techniques. Furthermore, GRANOLA emerges as the top-performing method among all baselines in the same time complexity class of Message Passing Neural Networks (MPNNs). Moshe Eliasof, Beatrice Bevilacqua, Carola-Bibiane Schönlieb, Haggai Maron |
NeurIPS | 3 |
| 2024 | DiGRAF: Diffeomorphic Graph-Adaptive Activation FunctionabstractIn this paper, we propose a novel activation function tailored specifically for graph data in Graph Neural Networks (GNNs). Motivated by the need for graph-adaptive and flexible activation functions, we introduce DiGRAF, leveraging Continuous Piecewise-Affine Based (CPAB) transformations, which we augment with an additional GNN to learn a graph-adaptive diffeomorphic activation function in an end-to-end manner. In addition to its graph-adaptivity and flexibility, DiGRAF also possesses properties that are widely recognized as desirable for activation functions, such as differentiability, boundness within the domain, and computational efficiency.
We conduct an extensive set of experiments across diverse datasets and tasks, demonstrating a consistent and superior performance of DiGRAF compared to traditional and graph-specific activation functions, highlighting its effectiveness as an activation function for GNNs. Our code is available at https://github.com/ipsitmantri/DiGRAF. Krishna Sri Ipsit Mantri, Carola-Bibiane Schönlieb, Bruno Ribeiro 0001, Beatrice Bevilacqua, Moshe Eliasof |
NeurIPS | 3 |
| 2024 | A linear transportation Lp distance for pattern recognitionabstractThe transportation Lp distance, denoted TLp, has been proposed as a generalisation of Wasserstein Wp distances motivated by the property that it can be applied directly to colour or multi-channelled images, as well as multivariate time-series without normalisation or mass constraints. These distances, as with Wp, are powerful tools in modelling data with spatial or temporal perturbations. However, their computational cost can make them infeasible to apply to even moderate pattern recognition tasks. We propose linear versions of these distances and show that the linear TLp distance significantly improves over the linear Wp distance on signal processing tasks, whilst being several orders of magnitude faster to compute than the TLp distance. Oliver M. Crook, Mihai Cucuringu, Tim Hurst, Carola-Bibiane Schönlieb, Matthew Thorpe, Konstantinos C. Zygalakis |
Pattern Recognit. | 4 |
| 2024 | NF-ULA: Normalizing Flow-Based Unadjusted Langevin Algorithm for Imaging Inverse ProblemsabstractAbstract. Bayesian methods for solving inverse problems are a powerful alternative to classical methods since the Bayesian approach offers the ability to quantify the uncertainty in the solution. In recent years, data-driven techniques for solving inverse problems have also been remarkably successful, due to their superior representation ability. In this work, we incorporate data-based models into a class of Langevin-based sampling algorithms for Bayesian inference in imaging inverse problems. In particular, we introduce NF-ULA (normalizing flow-based unadjusted Langevin algorithm), which involves learning a normalizing flow (NF) as the image prior. We use NF to learn the prior because a tractable closed-form expression for the log prior enables the differentiation of it using autograd libraries. Our algorithm only requires a normalizing flow-based generative network, which can be pretrained independently of the considered inverse problem and the forward operator. We perform theoretical analysis by investigating the well-posedness and nonasymptotic convergence of the resulting NF-ULA algorithm. The efficacy of the proposed NF-ULA algorithm is demonstrated in various image restoration problems such as image deblurring, image inpainting, and limited-angle X-ray computed tomography reconstruction. NF-ULA is found to perform better than competing methods for severely ill-posed inverse problems. Ziruo Cai, Junqi Tang, Subhadip Mukherjee, Jinglai Li, Carola-Bibiane Schönlieb, Xiaoqun Zhang |
SIAM J. Imaging Sci. | 5 |
| 2024 | Practical Acceleration of the Condat-Vũ AlgorithmabstractAbstract. The Condat–Vũ algorithm is a widely used primal-dual method for optimizing composite objectives of three functions. Several algorithms for optimizing composite objectives of two functions are special cases of Condat–Vũ, including proximal gradient descent (PGD). It is well known that PGD exhibits suboptimal performance, and a simple adjustment to PGD can accelerate its convergence rate from [Formula: see text] to [Formula: see text] on convex objectives, and this accelerated rate is optimal. In this work, we show that a simple adjustment to the Condat–Vũ algorithm allows it to recover accelerated PGD (APGD) as a special case, instead of PGD. We prove that this accelerated Condat–Vũ algorithm achieves optimal convergence rates and significantly outperforms the traditional Condat–Vũ algorithm in regimes where the Condat–Vũ algorithm approximates the dynamics of PGD. We demonstrate the effectiveness of our approach in various applications in machine learning and computational imaging. Derek Driggs, Matthias J. Ehrhardt, Carola-Bibiane Schönlieb, Junqi Tang |
SIAM J. Imaging Sci. | 3 |
| 2024 | Proximal Langevin Sampling with Inexact Proximal MappingabstractAbstract. In order to solve tasks like uncertainty quantification or hypothesis tests in Bayesian imaging inverse problems, we often have to draw samples from the arising posterior distribution. For the usually log-concave but high-dimensional posteriors, Markov chain Monte Carlo methods based on time discretizations of Langevin diffusion are a popular tool. If the potential defining the distribution is nonsmooth, these discretizations are usually of an implicit form leading to Langevin sampling algorithms that require the evaluation of proximal operators. For some of the potentials relevant in imaging problems this is only possible approximately using an iterative scheme. We investigate the behavior of a proximal Langevin algorithm under the presence of errors in the evaluation of proximal mappings. We generalize existing nonasymptotic and asymptotic convergence results of the exact algorithm to our inexact setting and quantify the bias between the target and the algorithm’s stationary distribution due to the errors. We show that the additional bias stays bounded for bounded errors and converges to zero for decaying errors in a strongly convex setting. We apply the inexact algorithm to sample numerically from the posterior of typical imaging inverse problems in which we can only approximate the proximal operator by an iterative scheme and validate our theoretical convergence results. Matthias J. Ehrhardt, Lorenz Kuger, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 3 |
| 2024 | Provably Convergent Plug-and-Play Quasi-Newton MethodsabstractAbstract. Plug-and-Play (PnP) methods are a class of efficient iterative methods that aim to combine data fidelity terms and deep denoisers using classical optimization algorithms, such as ISTA or ADMM, with applications in inverse problems and imaging. Provable PnP methods are a subclass of PnP methods with convergence guarantees, such as fixed point convergence or convergence to critical points of some energy function. Many existing provable PnP methods impose heavy restrictions on the denoiser or fidelity function, such as nonexpansiveness or strict convexity, respectively. In this work, we propose a novel algorithmic approach incorporating quasi-Newton steps into a provable PnP framework based on proximal denoisers, resulting in greatly accelerated convergence while retaining light assumptions on the denoiser. By characterizing the denoiser as the proximal operator of a weakly convex function, we show that the fixed points of the proposed quasi-Newton PnP algorithm are critical points of a weakly convex function. Numerical experiments on image deblurring and super-resolution demonstrate 2–8x faster convergence as compared to other provable PnP methods with similar reconstruction quality. Hong Ye Tan, Subhadip Mukherjee, Junqi Tang, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 4 |
| 2024 | LaplaceNet: A Hybrid Graph-Energy Neural Network for Deep Semisupervised ClassificationabstractSemisupervised learning (SSL) has received a lot of recent attention as it alleviates the need for large amounts of labeled data which can often be expensive, requires expert knowledge, and be time consuming to collect. Recent developments in deep semisupervised classification have reached unprecedented performance and the gap between supervised and SSL is ever-decreasing. This improvement in performance has been based on the inclusion of numerous technical tricks, strong augmentation techniques, and costly optimization schemes with multiterm loss functions. We propose a new framework, LaplaceNet, for deep semisupervised classification that has a greatly reduced model complexity. We utilize a hybrid approach where pseudolabels are produced by minimizing the Laplacian energy on a graph. These pseudolabels are then used to iteratively train a neural-network backbone. Our model outperforms state-of-the-art methods for deep semisupervised classification, over several benchmark datasets. Furthermore, we consider the application of strong augmentations to neural networks theoretically and justify the use of a multisampling approach for SSL. We demonstrate, through rigorous experimentation, that a multisampling augmentation approach improves generalization and reduces the sensitivity of the network to augmentation. Philip Sellars, Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | SCOTCH and SODA: A Transformer Video Shadow Detection FrameworkabstractShadows in videos are difficult to detect because of the large shadow deformation between frames. In this work, we argue that accounting for shadow deformation is essential when designing a video shadow detection method. To this end, we introduce the shadow deformation attention trajectory (SODA), a new type of video self-attention module, specially designed to handle the large shadow deformations in videos. Moreover, we present a new shadow contrastive learning mechanism (SCOTCH) which aims at guiding the network to learn a unified shadow representation from massive positive shadow pairs across different videos. We demonstrate empirically the effectiveness of our two contributions in an ablation study. Furthermore, we show that SCOTCH and SODA significantly outperforms existing techniques for video shadow detection. Code is available at the project page: https://lihaoliu-cambridge.github.io/scotch_and_soda/ Jean Prost, Lei Zhu 0003, Nicolas Papadakis, Pietro Liò, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero |
CVPR | 6 |
| 2023 | Robust Data-Driven Accelerated Mirror DescentabstractLearning-to-optimize is an emerging framework that leverages training data to speed up the solution of certain optimization problems. One such approach is based on the classical mirror descent algorithm, where the mirror map is modelled using input-convex neural networks. In this work, we extend this functional parameterization approach by introducing momentum into the iterations, based on the classical accelerated mirror descent. Our approach combines short-time accelerated convergence with stable long-time behavior. We empirically demonstrate additional robustness with respect to multiple parameters on denoising and deconvolution experiments. Hong Ye Tan, Subhadip Mukherjee, Junqi Tang, Andreas Hauptmann, Carola-Bibiane Schönlieb |
ICASSP | 5 |
| 2023 | CDiffMR: Can We Replace the Gaussian Noise with K-Space Undersampling for Fast MRI?
Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb, Guang Yang 0006 |
MICCAI (10) | 3 |
| 2023 | DiffMIC: Dual-Guidance Diffusion Network for Medical Image Classification
Huazhu Fu, Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb, Lei Zhu 0003 |
MICCAI (6) | 4 |
| 2023 | A Continuous-time Stochastic Gradient Descent Method for Continuous DataabstractOptimization problems with continuous data appear in, e.g., robust machine learning, functional data analysis, and variational inference. Here, the target function is given as an integral over a family of (continuously) indexed target functions---integrated with respect to a probability measure. Such problems can often be solved by stochastic optimization methods: performing optimization steps with respect to the indexed target function with randomly switched indices. In this work, we study a continuous-time variant of the stochastic gradient descent algorithm for optimization problems with continuous data. This so-called stochastic gradient process consists in a gradient flow minimizing an indexed target function that is coupled with a continuous-time index process determining the index. Index processes are, e.g., reflected diffusions, pure jump processes, or other Lévy processes on compact spaces. Thus, we study multiple sampling patterns for the continuous data space and allow for data simulated or streamed at runtime of the algorithm. We analyze the approximation properties of the stochastic gradient process and study its longtime behavior and ergodicity under constant and decreasing learning rates. We end with illustrating the applicability of the stochastic gradient process in a polynomial regression problem with noisy functional data, as well as in a physics-informed neural network. Kexin Jin, Jonas Latz, Carola-Bibiane Schönlieb |
J. Mach. Learn. Res. | 4 |
| 2023 | Joint Reconstruction-Segmentation on GraphsabstractAbstract. Practical image segmentation tasks concern images which must be reconstructed from noisy, distorted, and/or incomplete observations. A recent approach for solving such tasks is to perform this reconstruction jointly with the segmentation, using each to guide the other. However, this work has so far employed relatively simple segmentation methods, such as the Chan–Vese algorithm. In this paper, we present a method for joint reconstruction-segmentation using graph-based segmentation methods, which have been seeing increasing recent interest. Complications arise due to the large size of the matrices involved, and we show how these complications can be managed. We then analyze the convergence properties of our scheme. Finally, we apply this scheme to distorted versions of “two cows” images familiar from previous graph-based segmentation literature, first to a highly noised version and second to a blurred version, achieving highly accurate segmentations in both cases. We compare these results to those obtained by sequential reconstruction-segmentation approaches, finding that our method competes with, or even outperforms, those approaches in terms of reconstruction and segmentation accuracy. Jeremy Budd, Yves van Gennip, Jonas Latz, Simone Parisotto, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 5 |
| 2023 | Regularizing Orientation Estimation in Cryogenic Electron Microscopy Three-Dimensional Map Refinement through Measure-Based Lifting over Riemannian ManifoldsabstractAbstract. Motivated by the trade-off between noise robustness and data consistency for joint three-imensional (3D) map reconstruction and rotation estimation in single particle cryogenic-electron microscopy (Cryo-EM), we propose ellipsoidal support lifting (ESL), a measure-based lifting scheme for regularizing and approximating the global minimizer of a smooth function over a Riemannian manifold. Under a uniqueness assumption on the minimizer we show several theoretical results, in particular well-posedness of the method and an error bound due to the induced bias with respect to the global minimizer. Additionally, we use the developed theory to integrate the measure-based lifting scheme into an alternating update method for joint homogeneous 3D map reconstruction and rotation estimation, where typically tens of thousands of manifold-valued minimization problems have to be solved and where regularization is necessary because of the high noise levels in the data. The joint recovery method is used to test both the theoretical predictions and algorithmic performance through numerical experiments with Cryo-EM data. In particular, the induced bias due to the regularizing effect of ESL empirically estimates better rotations, i.e., rotations closer to the ground truth, than global optimization would. Willem Diepeveen, Jan Lellmann, Ozan Öktem, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 4 |
| 2023 | Guest Editorial Special Issue on Geometric Deep Learning in Medical ImagingabstractIn recent years, more and more attention has been devoted to geometric deep learning (GDL) and its applications to various problems in medical imaging. Unlike convolutional neural networks (CNNs) limited to 2-D/3-D grid-structured data, GDL can handle non-Euclidean data (i.e., graphs and manifolds) and is hence well-suited for medical imaging data such as structure-function connectivity networks, imaging genetics and omics, spatio-temporal anatomical representations, physics-informed GDL for optimal imaging sampling and acquisition, GDL in imaging inverse problems, etc. However, despite recent advances in GDL research, questions remain on how best to learn representations of non-Euclidean medical imaging data; how to convolve effectively on graphs; how to perform graph pooling/unpooling; how to handle heterogeneous data; and how to improve the interpretability of GDL. After discussing many other domain experts, we identify the need for a special issue that brings to the attention of the medical imaging community these interesting topics. Huazhu Fu, Yitian Zhao, Pew-Thian Yap, Carola-Bibiane Schönlieb, Alejandro F. Frangi |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Multi-Modal Learning for Predicting the Genotype of GliomaabstractThe isocitrate dehydrogenase (IDH) gene mutation is an essential biomarker for the diagnosis and prognosis of glioma. It is promising to better predict glioma genotype by integrating focal tumor image and geometric features with brain network features derived from MRI. Convolutional neural networks show reasonable performance in predicting IDH mutation, which, however, cannot learn from non-Euclidean data, e.g., geometric and network data. In this study, we propose a multi-modal learning framework using three separate encoders to extract features of focal tumor image, tumor geometrics and global brain networks. To mitigate the limited availability of diffusion MRI, we develop a self-supervised approach to generate brain networks from anatomical multi-sequence MRI. Moreover, to extract tumor-related features from the brain network, we design a hierarchical attention module for the brain network encoder. Further, we design a bi-level multi-modal contrastive loss to align the multi-modal features and tackle the domain gap at the focal tumor and global brain. Finally, we propose a weighted population graph to integrate the multi-modal features for genotype prediction. Experimental results on the testing set show that the proposed model outperforms the baseline deep learning models. The ablation experiments validate the performance of different components of the framework. The visualized interpretation corresponds to clinical knowledge with further validation. In conclusion, the proposed learning framework provides a novel approach for predicting the genotype of glioma. Yiran Wei 0002, Xi Chen 0042, Lei Zhu 0003, Lipei Zhang, Carola-Bibiane Schönlieb, Stephen J. Price, Chao Li 0031 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | S $^3$ Net: Self-Supervised Self-Ensembling Network for Semi-Supervised RGB-D Salient Object DetectionabstractRGB-D salient object detection aims to detect visually distinctive objects or regions from a pair of the RGB image and the depth image. State-of-the-art RGB-D saliency detectors are mainly based on convolutional neural networks but almost suffer from an intrinsic limitation relying on the labeled data, thus degrading detection accuracy in complex cases. In this work, we present a self-supervised self-ensembling network (S$^3$Net) for semi-supervised RGB-D salient object detection by leveraging the unlabeled data and exploring a self-supervised learning mechanism. To be specific, we first build a self-guided convolutional neural network (SG-CNN) as a baseline model by developing a series of three-layer cross-model feature fusion (TCF) modules to leverage complementary information among depth and RGB modalities and formulating an auxiliary task that predicts a self-supervised image rotation angle. After that, to further explore the knowledge from unlabeled data, we assign SG-CNN to a student network and a teacher network, and encourage the saliency predictions and self-supervised rotation predictions from these two networks to be consistent on the unlabeled data. Experimental results on seven widely-used benchmark datasets demonstrate that our network quantitatively and qualitatively outperforms the state-of-the-art methods. Lei Zhu 0003, Xiaoqiang Wang 0007, Ping Li 0016, Xin Yang 0011, Qing Zhang 0006, Weiming Wang 0002, Carola-Bibiane Schönlieb, C. L. Philip Chen |
IEEE Trans. Multim. | 7 |
| 2022 | Mutual Contrastive Low-rank Learning to Disentangle Whole Slide Image Representations for Glioma Grading
Lipei Zhang, Yiran Wei 0002, Ying Fu 0001, Stephen J. Price, Carola-Bibiane Schönlieb, Chao Li 0031 |
BMVC | 5 |
| 2022 | Rethinking Video Rain Streak Removal: A New Synthesis Model and a Deraining Network with Video Rain Prior
Lei Zhu 0003, Huazhu Fu, Harry Qin, Carola-Bibiane Schönlieb, Wei Feng 0005, Song Wang 0002 |
ECCV (19) | 5 |
| 2022 | Stylegan-Induced Data-Driven Regularization for Inverse ProblemsabstractRecent advances in generative adversarial networks (GANs) have opened up the possibility of generating high-resolution photo-realistic images that were impossible to produce previously. The ability of GANs to sample from high-dimensional distributions has naturally motivated researchers to leverage their power for modeling the image prior in inverse problems. We extend this line of research by developing a Bayesian image reconstruction framework that utilizes the full potential of a pre-trained StyleGAN2 generator, which is the currently dominant GAN architecture, for constructing the prior distribution on the underlying image. Our proposed approach, which we refer to as learned Bayesian reconstruction with generative models (L-BRGM), entails joint optimization over the style-code and the input latent code, and enhances the expressive power of a pre-trained StyleGAN2 generator by allowing the style-codes to be different for different generator layers. Considering the inverse problems of image inpainting and super-resolution, we demonstrate that the proposed approach is competitive with, and sometimes superior to, state-of-the-art GAN-based image reconstruction methods. Arthur Conmy, Subhadip Mukherjee, Carola-Bibiane Schönlieb |
ICASSP | 3 |
| 2022 | Multi-modal Hypergraph Diffusion Network with Dual Prior for Alzheimer Classification
Angelica I. Avilés-Rivero, Christina Runkel, Nicolas Papadakis, Zoe Kourtzi, Carola-Bibiane Schönlieb |
MICCAI (3) | 5 |
| 2022 | HERS Superpixels: Deep Affinity Learning for Hierarchical Entropy Rate SegmentationabstractSuperpixels serve as a powerful preprocessing tool in numerous computer vision tasks. By using superpixel representation, the number of image primitives can be largely reduced by orders of magnitudes. With the rise of deep learning in recent years, a few works have attempted to feed deeply learned features / graphs into existing classical superpixel techniques. However, none of them are able to produce superpixels in near real-time, which is crucial to the applicability of superpixels in practice. In this work, we propose a two-stage graph-based framework for superpixel segmentation. In the first stage, we introduce an efficient Deep Affinity Learning (DAL) network that learns pairwise pixel affinities by aggregating multi-scale information. In the second stage, we propose a highly efficient superpixel method called Hierarchical Entropy Rate Segmentation (HERS). Using the learned affinities from the first stage, HERS builds a hierarchical tree structure that can produce any number of highly adaptive superpixels instantaneously. We demonstrate, through visual and numerical experiments, the effectiveness and efficiency of our method compared to various state-of-the-art superpixel methods.1 Hankui Peng, Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb |
WACV | 3 |
| 2022 | On Biased Stochastic Gradient EstimationabstractWe present a uniform analysis of biased stochastic gradient methods for minimizing convex, strongly convex, and non-convex composite objectives, and identify settings where bias is useful in stochastic gradient estimation. The framework we present allows us to extend proximal support to biased algorithms, including SAG and SARAH, for the first time in the convex setting. We also use our framework to develop a new algorithm, Stochastic Average Recursive GradiEnt (SARGE), that achieves the oracle complexity lower-bound for non-convex, finite-sum objectives and requires strictly fewer calls to a stochastic gradient oracle per iteration than SVRG and SARAH. We support our theoretical results with numerical experiments that demonstrate the benefits of certain biased gradient estimators. Derek Driggs, Jingwei Liang, Carola-Bibiane Schönlieb |
J. Mach. Learn. Res. | 3 |
| 2022 | TFPnP: Tuning-free Plug-and-Play Proximal Algorithms with Applications to Inverse Imaging ProblemsabstractPlug-and-Play (PnP) is a non-convex optimization framework that combines proximal algorithms, for example, the alternating direction method of multipliers (ADMM), with advanced denoising priors. Over the past few years, great empirical success has been obtained by PnP algorithms, especially for the ones that integrate deep learning-based denoisers. However, a key problem of PnP approaches is the need for manual parameter tweaking which is essential to obtain high-quality results across the high discrepancy in imaging conditions and varying scene content. In this work, we present a class of tuning-free PnP proximal algorithms that can determine parameters such as denoising strength, termination time, and other optimization-specific parameters automatically. A core part of our approach is a policy network for automated parameter search which can be effectively learned via a mixture of model-free and model-based deep reinforcement learning strategies. We demonstrate, through rigorous numerical and visual experiments, that the learned policy can customize parameters to different settings, and is often more efficient and effective than existing handcrafted criteria. Moreover, we discuss several practical considerations of PnP denoisers, which together with our learned policy yield state-of-the-art results. This advanced performance is prevalent on both linear and nonlinear exemplar inverse imaging problems, and in particular shows promising results on compressed sensing MRI, sparse-view CT, single-photon imaging, and phase retrieval. Kaixuan Wei, Angelica I. Avilés-Rivero, Jingwei Liang, Ying Fu 0001, Hua Huang 0001, Carola-Bibiane Schönlieb |
J. Mach. Learn. Res. | 6 |
| 2022 | Beyond fine-tuning: Classifying high resolution mammograms using function-preserving transformationsabstractThe task of classifying mammograms is very challenging because the lesion is usually small in the high resolution image. The current state-of-the-art approaches for medical image classification rely on using the de-facto method for convolutional neural networks-fine-tuning. However, there are fundamental differences between natural images and medical images, which based on existing evidence from the literature, limits the overall performance gain when designed with algorithmic approaches. In this paper, we propose to go beyond fine-tuning by introducing a novel framework called MorphHR, in which we highlight a new transfer learning scheme. The idea behind the proposed framework is to integrate function-preserving transformations, for any continuous non-linear activation neurons, to internally regularise the network for improving mammograms classification. The proposed solution offers two major advantages over the existing techniques. Firstly and unlike fine-tuning, the proposed approach allows for modifying not only the last few layers but also several of the first ones on a deep ConvNet. By doing this, we can design the network front to be suitable for learning domain specific features. Secondly, the proposed scheme is scalable to hardware. Therefore, one can fit high resolution images on standard GPU memory. We show that by using high resolution images, one prevents losing relevant information. We demonstrate, through numerical and visual experiments, that the proposed approach yields to a significant improvement in the classification performance over state-of-the-art techniques, and is indeed on a par with radiology experts. Moreover and for generalisation purposes, we show the effectiveness of the proposed learning scheme on another large dataset, the ChestX-ray14, surpassing current state-of-the-art techniques. Angelica I. Avilés-Rivero, Shuo Wang 0011, Yuan Huang 0009, Fiona J. Gilbert, Carola-Bibiane Schönlieb, Chang Wen Chen |
Medical Image Anal. | 6 |
| 2022 | Unsupervised Image Restoration Using Partially Linear DenoisersabstractDeep neural network based methods are the state of the art in various image restoration problems. Standard supervised learning frameworks require a set of noisy measurement and clean image pairs for which a distance between the output of the restoration model and the ground truth, clean images is minimized. The ground truth images, however, are often unavailable or very expensive to acquire in real-world applications. We circumvent this problem by proposing a class of structured denoisers that can be decomposed as the sum of a nonlinear image-dependent mapping, a linear noise-dependent term and a small residual term. We show that these denoisers can be trained with only noisy images under the condition that the noise has zero mean and known variance. The exact distribution of the noise, however, is not assumed to be known. We show the superiority of our approach for image denoising, and demonstrate its extension to solving other restoration problems such as image deblurring where the ground truth is not available. Our method outperforms some recent unsupervised and self-supervised deep denoising models that do not require clean images for their training. For deblurring problems, the method, using only one noisy and blurry observation per image, reaches a quality not far away from its fully supervised counterparts on a benchmark dataset. Rihuan Ke, Carola-Bibiane Schönlieb |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | AI-Based Reconstruction for Fast MRI - A Systematic Review and Meta-AnalysisabstractCompressed sensing (CS) has been playing a key role in accelerating the magnetic resonance imaging (MRI) acquisition process. With the resurgence of artificial intelligence, deep neural networks and CS algorithms are being integrated to redefine the state of the art of fast MRI. The past several years have witnessed substantial growth in the complexity, diversity, and performance of deep-learning-based CS techniques that are dedicated to fast MRI. In this meta-analysis, we systematically review the deep-learning-based CS techniques for fast MRI, describe key model designs, highlight breakthroughs, and discuss promising directions. We have also introduced a comprehensive analysis framework and a classification system to assess the pivotal role of deep learning in CS-based acceleration for MRI. Carola-Bibiane Schönlieb, Pietro Liò, Tim Leiner, Pier Luigi Dragotti, Ge Wang 0001, Daniel Rueckert, David N. Firmin, Guang Yang 0006 |
Proc. IEEE | 2 |
| 2022 | GraphXCOVID: Explainable deep graph diffusion pseudo-Labelling for identifying COVID-19 on chest X-rays
Angelica I. Avilés-Rivero, Philip Sellars, Carola-Bibiane Schönlieb, Nicolas Papadakis |
Pattern Recognit. | 3 |
| 2022 | Semi-Supervised Superpixel-Based Multi-Feature Graph Learning for Hyperspectral Image DataabstractGraphs naturally lend themselves to model the complexities of hyperspectral image (HSI) data as well as to serve as semi-supervised classifiers by propagating given labels among nearest neighbors. In this work, we present a novel framework for the classification of HSI data in light of a very limited amount of labeled data, inspired by multi-view graph learning and graph signal processing. Given ana priorisuperpixel-segmented HSI, we seek a robust and efficient graph construction and label propagation method to conduct semi-supervised learning (SSL). Since the graph is paramount to the success of the subsequent classification task, particularly in light of the intrinsic complexity of HSI data, we consider the problem of finding the optimal graph to model such data. Our contribution is two-fold. First, we propose a multi-stage edge-efficient semi-supervised graph learning framework for HSI data, which exploits given label information through pseudo-label features embedded in the graph construction. Second, we examine and enhance the contribution of multiple superpixel features embedded in the graph on the basis of pseudo-label features in an extension of the previous framework, which is less reliant on excessive parameter tuning. Ultimately, we demonstrate the superiority of our approaches in comparison with state-of-the-art methods through extensive numerical experiments. Madeleine S. Kotzagiannidis, Carola-Bibiane Schönlieb |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Three-Stage Self-Training Framework for Semi-Supervised Semantic SegmentationabstractSemantic segmentation has been widely investigated in the community, in which state-of-the-art techniques are based on supervised models. Those models have reported unprecedented performance at the cost of requiring a large set of high quality segmentation masks for training. Obtaining such annotations is highly expensive and time consuming, in particular, in semantic segmentation where pixel-level annotations are required. In this work, we address this problem by proposing a holistic solution framed as a self-training framework for semi-supervised semantic segmentation. The key idea of our technique is the extraction of the pseudo-mask information on unlabelled data whilst enforcing segmentation consistency in a multi-task fashion. We achieve this through a three-stage solution. Firstly, a segmentation network is trained using the labelled data only and rough pseudo-masks are generated for all images. Secondly, we decrease the uncertainty of the pseudo-mask by using a multi-task model that enforces consistency and that exploits the rich statistical information of the data. Finally, the segmentation model is trained by taking into account the information of the higher quality pseudo-masks. We compare our approach against existing semi-supervised semantic segmentation methods and demonstrate state-of-the-art performance with extensive experiments. Rihuan Ke, Angelica I. Avilés-Rivero, Saurabh Pandey, Saikumar Reddy, Carola-Bibiane Schönlieb |
IEEE Trans. Image Process. | 5 |
| 2021 | End-to-end reconstruction meets data-driven regularization for inverse problemsabstractWe propose a new approach for learning end-to-end reconstruction operators based on unpaired training data for ill-posed inverse problems. The proposed method combines the classical variational framework with iterative unrolling and essentially seeks to minimize a weighted combination of the expected distortion in the measurement space and the Wasserstein-1 distance between the distributions of the reconstruction and the ground-truth. More specifically, the regularizer in the variational setting is parametrized by a deep neural network and learned simultaneously with the unrolled reconstruction operator. The variational problem is then initialized with the output of the reconstruction network and solved iteratively till convergence. Notably, it takes significantly fewer iterations to converge as compared to variational methods, thanks to the excellent initialization obtained via the unrolled operator. The resulting approach combines the computational efficiency of end-to-end unrolled reconstruction with the well-posedness and noise-stability guarantees of the variational setting. Moreover, we demonstrate with the example of image reconstruction in X-ray computed tomography (CT) that our approach outperforms state-of-the-art unsupervised methods and that it outperforms or is at least on par with state-of-the-art supervised data-driven reconstruction approaches. Subhadip Mukherjee, Marcello Carioni, Ozan Öktem, Carola-Bibiane Schönlieb |
NeurIPS | 4 |
| 2021 | Compressed sensing plus motion (CS + M): A new perspective for improving undersampled MR image reconstruction
Angelica I. Avilés-Rivero, Noémie Debroux, Guy B. Williams, Martin J. Graves, Carola-Bibiane Schönlieb |
Medical Image Anal. | 5 |
| 2021 | Variational multi-task MRI reconstruction: Joint reconstruction, registration and super-resolution
Veronica Corona, Angelica I. Avilés-Rivero, Noémie Debroux, Carole Le Guyader, Carola-Bibiane Schönlieb |
Medical Image Anal. | 5 |
| 2021 | Rethinking medical image reconstruction via shape prior, going deeper and faster: Deep joint indirect registration and reconstruction
Jiulong Liu, Angelica I. Avilés-Rivero, Hui Ji 0002, Carola-Bibiane Schönlieb |
Medical Image Anal. | 4 |
| 2021 | Dynamic spectral residual superpixels
Jianchao Zhang, Angelica I. Avilés-Rivero, Daniel Heydecker, Xiaosheng Zhuang, Raymond Chan 0001, Carola-Bibiane Schönlieb |
Pattern Recognit. | 6 |
| 2021 | Choose Your Path Wisely: Gradient Descent in a Bregman Distance FrameworkabstractWe propose an extension of a special form of gradient descent---in the literature known as linearized Bregman iteration---to a larger class of nonconvex functions. We replace the classical (squared) two norm metric in the gradient descent setting with a generalized Bregman distance, based on a proper, convex, and lower semicontinuous function. The algorithm's global convergence is proven for functions that satisfy the Kurdyka--Łojasiewicz property. Examples illustrate that features of different scale are being introduced throughout the iteration, transitioning from coarse to fine. This coarse-to-fine approach with respect to scale allows us to recover solutions of nonconvex optimization problems that are superior to those obtained with conventional gradient descent, or even projected and proximal gradient descent. The effectiveness of the linearized Bregman iteration in combination with early stopping is illustrated for the applications of parallel magnetic resonance imaging, blind deconvolution, as well as image classification with neural networks. Martin Benning, Marta M. Betcke, Matthias J. Ehrhardt, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 4 |
| 2021 | A Stochastic Proximal Alternating Minimization for Nonsmooth and Nonconvex OptimizationabstractIn this work, we introduce a novel stochastic proximal alternating linearized minimization algorithm [J. Bolte, S. Sabach, and M. Teboulle, Math. Program., 146 (2014), pp. 459--494] for solving a class of nonsmooth and nonconvex optimization problems. Large-scale imaging problems are becoming increasingly prevalent due to the advances in data acquisition and computational capabilities. Motivated by the success of stochastic optimization methods, we propose a stochastic variant of proximal alternating linearized minimization. We provide global convergence guarantees, demonstrating that our proposed method with variance-reduced stochastic gradient estimators, such as SAGA [A. Defazio, F. Bach, and S. Lacoste-Julien, Advances in Neural Information Processing Systems, 2014, pp. 1646--1654] and SARAH [L. M. Nguyen, J. Liu, K. Scheinberg, and M. Takáĉ, Proceedings of the 34th International Conference on Machine Learning, PMLR 70, 2017, pp. 2613--2621], achieves state-of-the-art oracle complexities. We also demonstrate the efficacy of our algorithm via several numerical examples including sparse nonnegative matrix factorization, sparse principal component analysis, and blind image-deconvolution. Derek Driggs, Junqi Tang, Jingwei Liang, Mike E. Davies 0001, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 5 |
| 2021 | On Learned Operator Correction in Inverse ProblemsabstractWe discuss the possibility of learning a data-driven explicit model correction for inverse problems and whether such a model correction can be used within a variational framework to obtain regularized reconstructions. This paper discusses the conceptual difficulty of learning such a forward model correction and proceeds to present a possible solution as a forward-adjoint correction that explicitly corrects in both data and solution spaces. We then derive conditions under which solutions to the variational problem with a learned correction converge to solutions obtained with the correct operator. The proposed approach is evaluated on an application to limited view photoacoustic tomography and compared to the established framework of the Bayesian approximation error method. Sebastian Lunz, Andreas Hauptmann, Tanja Tarvainen, Carola-Bibiane Schönlieb, Simon R. Arridge |
SIAM J. Imaging Sci. | 4 |
| 2021 | Multi-Task Deep Learning for Image Segmentation Using Recursive Approximation TasksabstractFully supervised deep neural networks for segmentation usually require a massive amount of pixel-level labels which are manually expensive to create. In this work, we develop a multi-task learning method to relax this constraint. We regard the segmentation problem as a sequence of approximation subproblems that are recursively defined and in increasing levels of approximation accuracy. The subproblems are handled by a framework that consists of 1) a segmentation task that learns from pixel-level ground truth segmentation masks of a small fraction of the images, 2) a recursive approximation task that conducts partial object regions learning and data-driven mask evolution starting from partial masks of each object instance, and 3) other problem oriented auxiliary tasks that are trained with sparse annotations and promote the learning of dedicated features. Most training images are only labeled by (rough) partial masks, which do not contain exact object boundaries, rather than by their full segmentation masks. During the training phase, the approximation task learns the statistics of these partial masks, and the partial regions are recursively increased towards object boundaries aided by the learned information from the segmentation task in a fully data-driven fashion. The network is trained on an extremely small amount of precisely segmented images and a large set of coarse labels. Annotations can thus be obtained in a cheap way. We demonstrate the efficiency of our approach in three applications with microscopy images and ultrasound images. Rihuan Ke, Aurélie Bugeau, Nicolas Papadakis, Mark Kirkland, Peter Schütz, Carola-Bibiane Schönlieb |
IEEE Trans. Image Process. | 6 |
| 2020 | Tuning-free Plug-and-Play Proximal Algorithm for Inverse Imaging ProblemsabstractPlug-and-play (PnP) is a non-convex framework that combines ADMM or other proximal algorithms with advanced denoiser priors. Recently, PnP has achieved great empirical success, especially with the integration of deep learning-based denoisers. However, a key problem of PnP based approaches is that they require manual parameter tweaking. It is necessary to obtain high-quality results across the high discrepancy in terms of imaging conditions and varying scene content. In this work, we present a tuning-free PnP proximal algorithm, which can automatically determine the internal parameters including the penalty parameter, the denoising strength and the terminal time. A key part of our approach is to develop a policy network for automatic search of parameters, which can be effectively learned via mixed model-free and model-based deep reinforcement learning. We demonstrate, through numerical and visual experiments, that the learned policy can customize different parameters for different states, and often more efficient and effective than existing handcrafted criteria. Moreover, we discuss the practical considerations of the plugged denoisers, which together with our learned policy yield state-of-the-art results. This is prevalent on both linear and nonlinear exemplary inverse imaging problems, and in particular, we show promising results on Compressed Sensing MRI and phase retrieval. Kaixuan Wei, Angelica I. Avilés-Rivero, Jingwei Liang, Ying Fu 0001, Carola-Bibiane Schönlieb, Hua Huang 0001 |
ICML | 5 |
| 2020 | Deeply Learned Spectral Total Variation DecompositionabstractNon-linear spectral decompositions of images based on one-homogeneous functionals such as total variation have gained considerable attention in the last few years. Due to their ability to extract spectral components corresponding to objects of different size and contrast, such decompositions enable filtering, feature transfer, image fusion and other applications. However, obtaining this decomposition involves solving multiple non-smooth optimisation problems and is therefore computationally highly intensive. In this paper, we present a neural network approximation of a non-linear spectral decomposition. We report up to four orders of magnitude (×10,000) speedup in processing of mega-pixel size images, compared to classical GPU implementations. Our proposed network, TVspecNET, is able to implicitly learn the underlying PDE and, despite being entirely data driven, inherits invariances of the model based transform. To the best of our knowledge, this is the first approach towards learning a non-linear spectral decomposition of images. Not only do we gain a staggering computational advantage, but this approach can also be seen as a step towards studying neural networks that can decompose an image into spectral components defined by a user rather than a handcrafted functional. Tamara G. Grossmann, Yury Korolev, Guy Gilboa, Carola-Bibiane Schönlieb |
NeurIPS | 4 |
| 2020 | A Variational Model Dedicated to Joint Segmentation, Registration, and Atlas Generation for Shape AnalysisabstractIn medical image analysis, constructing an atlas, i.e., a mean representative of an ensemble of images, is a critical task for practitioners to estimate variability of shapes inside a population, and to characterize and understand how structural shape changes have an impact on health. This involves identifying significant shape constituents of a set of images, a process called segmentation, and mapping this group of images to an unknown mean image, a task called registration, making a statistical analysis of the image population possible. To achieve this goal, we propose treating these operations jointly to leverage their positive mutual influence, in a hyperelasticity setting, by viewing the shapes to be matched as Ogden materials. The approach is complemented by novel hard constraints on the $L^\infty$ norm of both the Jacobian and its inverse, ensuring that the deformation is a bi-Lipschitz homeomorphism. Segmentation is based on the Potts model, which allows for a partition into more than two regions, i.e., more than one shape. The connection to the registration problem is ensured by the dissimilarity measure that aims to align the segmented shapes. A representation of the deformation field in a linear space equipped with a scalar product is then computed in order to perform a geometry-driven Principal Component Analysis (PCA) and to extract the main modes of variations inside the image population. Theoretical results emphasizing the mathematical soundness of the model are provided, among which are existence of minimizers, analysis of a numerical method, asymptotic results, and a PCA analysis, as well as numerical simulations demonstrating the ability of the model to produce an atlas exhibiting sharp edges, high contrast, and a consistent shape. Noémie Debroux, John A. D. Aston, Fabien Bonardi, Alistair Forbes, Carole Le Guyader, Marina Romanchikova, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 7 |
| 2020 | Higher-Order Total Directional Variation: Imaging ApplicationsabstractWe introduce a class of higher-order anisotropic total variation regularizers, which are defined for possibly inhomogeneous, smooth elliptic anisotropies, that extends the total generalized variation regularizer and its variants. We propose a primal-dual hybrid gradient approach to approximating numerically the associated gradient flow. This choice of regularizers allows us to preserve and enhance intrinsic anisotropic features in images. This is illustrated on various examples from different imaging applications: image denoising, wavelet-based image zooming, and reconstruction of surfaces from scattered height measurements. Simone Parisotto, Jan Lellmann, Simon Masnou, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 4 |
| 2020 | Higher-Order Total Directional Variation: AnalysisabstractWe analyze a new notion of total anisotropic higher-order variation which, differently from total generalized variation in [K. Bredies, K. Kunisch, and T. Pock, SIAM J. Imaging Sci., 3 (2010), pp. 492--526], quantifies for possibly nonsymmetric tensor fields their variations at arbitrary order weighted by possibly inhomogeneous, smooth elliptic anisotropies. We prove some properties of this total variation and of the associated spaces of tensors with finite variations. We show the existence of solutions to a related regularity-fidelity optimization problem. We also prove a decomposition formula which appears to be helpful for the design of numerical schemes, as shown in a companion paper, where several applications to image processing are studied. Simone Parisotto, Simon Masnou, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 3 |
| 2020 | Superpixel Contracted Graph-Based Learning for Hyperspectral Image ClassificationabstractA central problem in hyperspectral image (HSI) classification is obtaining high classification accuracy when using a limited amount of labeled data. In this article we present a novel graph-based semi-supervised framework to tackle this problem. Our framework uses a superpixel approach, allowing it to define meaningful local regions in HSIs, which with high probability share the same classification label. We then extract spectral and spatial features from these regions and use them to produce a contracted weighted graph-representation, where each node represents a region rather than a pixel. The graph is then fed into a graph-based semi-supervised classifier which gives the final classification. We show that using superpixels in a graph representation is an effective tool for speeding up graphical classifiers applied to HSIs. We demonstrate through exhaustive quantitative and qualitative results that our proposed method produces accurate classifications when an incredibly small amount of labeled data is used. We show that our approach mitigates the major drawbacks of existing approaches, resulting in our approach outperforming several comparative state-of-the-art techniques. Philip Sellars, Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | 3D Segmentation of Trees Through a Flexible Multiclass Graph Cut AlgorithmabstractDeveloping a robust algorithm for automatic individual tree crown (ITC) detection from airborne laser scanning (ALS) data sets is important for tracking the responses of trees to anthropogenic change. Such approaches allow the size, growth, and mortality of individual trees to be measured, enabling forest carbon stocks and dynamics to be tracked and understood. Many algorithms exist for structurally simple forests, including coniferous forests and plantations. Finding a robust solution for structurally complex, species-rich tropical forests remains a challenge; existing segmentation algorithms often perform less well than simple area-based approaches when estimating plot-level biomass. Here, we describe a multiclass graph cut (MCGC) approach to tree crown delineation. This uses local 3D geometry and density information, alongside knowledge of crown allometries, to segment ITCs from airborne light detection and ranging point clouds. Our approach robustly identifies trees in the top and intermediate layers of the canopy, but cannot recognize small trees. From these 3D crowns, we are able to measure individual tree biomass. Comparing these estimates with those from permanent inventory plots, our algorithm can produce robust estimates of hectare-scale carbon density, demonstrating the power of ITC approaches in monitoring forests. The flexibility of our method to add additional dimensions of information, such as spectral reflectance, make this approach an obvious avenue for future development and extension to other sources of 3D data, such as structure from motion data sets. Jonathan Williams 0003, Carola-Bibiane Schönlieb, Tom Swinfield, Juheon Lee, Xiaohao Cai, Lan Qie, David Coomes |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Variational Osmosis for Non-Linear Image FusionabstractWe propose a new variational model for non-linear image fusion. Our approach is based on the use of an osmosis energy term related to the one studied in Vogel et al. [44] and Weickert et al. [45]. The minimization of the proposed non-convex energy realizes visually plausible image data fusion, invariant to multiplicative brightness changes. On the practical side, it requires minimal supervision and parameter tuning and can encode prior information on the structure of the images to be fused. For the numerical solution of the proposed model, we develop a primal-dual algorithm and we apply the resulting minimization scheme to solve multi-modal face fusion, color transfer and cultural heritage conservation problems. Visual and quantitative comparisons to state-of-the-art approaches prove the out-performance and the flexibility of our method. Simone Parisotto, Luca Calatroni, Aurélie Bugeau, Nicolas Papadakis, Carola-Bibiane Schönlieb |
IEEE Trans. Image Process. | 5 |
| 2020 | Learning the Sampling Pattern for MRIabstractThe discovery of the theory of compressed sensing brought the realisation that many inverse problems can be solved even when measurements are "incomplete". This is particularly interesting in magnetic resonance imaging (MRI), where long acquisition times can limit its use. In this work, we consider the problem of learning a sparse sampling pattern that can be used to optimally balance acquisition time versus quality of the reconstructed image. We use a supervised learning approach, making the assumption that our training data is representative enough of new data acquisitions. We demonstrate that this is indeed the case, even if the training data consists of just 7 training pairs of measurements and ground-truth images; with a training set of brain images of size 192 by 192, for instance, one of the learned patterns samples only 35% of k-space, however results in reconstructions with mean SSIM 0.914 on a test set of similar images. The proposed framework is general enough to learn arbitrary sampling patterns, including common patterns such as Cartesian, spiral and radial sampling. Ferdia Sherry, Martin Benning, Juan Carlos de los Reyes, Martin J. Graves, Georg Maierhofer, Guy B. Williams, Carola-Bibiane Schönlieb, Matthias J. Ehrhardt |
IEEE Trans. Medical Imaging | 7 |
| 2019 | RainFlow: Optical Flow Under Rain Streaks and Rain Veiling EffectabstractOptical flow in heavy rainy scenes is challenging due to the presence of both rain steaks and rain veiling effect, which break the existing optical flow constraints. Concerning this, we propose a deep-learning based optical flow method designed to handle heavy rain. We introduce a feature multiplier in our network that transforms the features of an image affected by the rain veiling effect into features that are less affected by it, which we call veiling-invariant features. We establish a new mapping operation in the feature space to produce streak-invariant features. The operation is based on a feature pyramid structure of the input images, and the basic idea is to preserve the chromatic features of the background scenes while canceling the rain-streak patterns. Both the veiling-invariant and streak-invariant features are computed and optimized automatically based on the the accuracy of our optical flow estimation. Our network is end-to-end, and handles both rain streaks and the veiling effect in an integrated framework. Extensive experiments show the effectiveness of our method, which outperforms the state of the art method and other baseline methods. We also show that our network can robustly maintain good performance on clean (no rain) images even though it is trained under rain image data. Ruoteng Li, Robby T. Tan, Loong Fah Cheong, Angelica I. Avilés-Rivero, Qingnan Fan, Carola-Bibiane Schönlieb |
ICCV | 6 |
| 2019 | On the Connection Between Adversarial Robustness and Saliency Map InterpretabilityabstractRecent studies on the adversarial vulnerability of neural networks have shown that models trained to be more robust to adversarial attacks exhibit more interpretable saliency maps than their non-robust counterparts. We aim to quantify this behaviour by considering the alignment between input image and saliency map. We hypothesize that as the distance to the decision boundary grows, so does the alignment. This connection is strictly true in the case of linear models. We confirm these theoretical findings with experiments based on models trained with a local Lipschitz regularization and identify where the nonlinear nature of neural networks weakens the relation. Christian Etmann, Sebastian Lunz, Peter Maass, Carola-Bibiane Schönlieb |
ICML | 4 |
| 2019 | Semi-Supervised Learning with Graphs: Covariance Based Superpixels For Hyperspectral Image ClassificationabstractIn this paper, we present a graph-based semi-supervised framework for hyperspectral image classification. We first introduce a novel superpixel algorithm based on the spectral covariance matrix representation of pixels to provide a better representation of our data. We then construct a superpixel graph, based on carefully considered feature vectors, before performing classification. We demonstrate, through a set of experimental results using two benchmarking datasets, that our approach outperforms three state-of-the-art classification frameworks, especially when a extremely small amount of labelled data is used. Philip Sellars, Angelica I. Avilés-Rivero, Nicolas Papadakis, David Coomes, Anita Faul, Carola-Bibiane Schönlieb |
IGARSS | 6 |
| 2019 | GraphX $$^\mathbf{\small NET } -$$ -Chest X-Ray Classification Under Extreme Minimal Supervision
Angelica I. Avilés-Rivero, Nicolas Papadakis, Ruoteng Li, Philip Sellars, Qingnan Fan, Robby T. Tan, Carola-Bibiane Schönlieb |
MICCAI (6) | 7 |
| 2019 | GANReDL: Medical Image Enhancement Using a Generative Adversarial Network with Real-Order Derivative Induced Loss Functions
Pan Liu 0004, Chao Li 0031, Carola-Bibiane Schönlieb |
MICCAI (3) | 3 |
| 2019 | Mirror, Mirror, on the Wall, Who's Got the Clearest Image of Them All? - A Tailored Approach to Single Image Reflection RemovalabstractRemoving reflection artefacts from a single image is a problem of both theoretical and practical interest, which still presents challenges because of the massively ill-posed nature of the problem. In this paper, we propose a technique based on a novel optimization problem. First, we introduce a simple user interaction scheme, which helps minimize information loss in the reflection-free regions. Second, we introduce an H2fidelity term, which preserves fine detail while enforcing the global color similarity. We show that this combination allows us to mitigate the shortcomings in structure and color preservation, which presents some of the most prominent drawbacks in the existing methods for reflection removal. We demonstrate, through numerical and visual experiments, that our method is able to outperform the state-of-the-art model-based methods and compete with recent deep-learning approaches. Daniel Heydecker, Georg Maierhofer, Angelica I. Avilés-Rivero, Qingnan Fan, Dongdong Chen 0001, Carola-Bibiane Schönlieb, Sabine Süsstrunk |
IEEE Trans. Image Process. | 6 |
| 2018 | Peekaboo-Where are the Objects? Structure Adjusting SuperpixelsabstractThis paper addresses the search for a fast and meaningful image segmentation in the context of k-means clustering. The proposed method builds on a widely-used local version of Lloyd's algorithm, called Simple Linear Iterative Clustering (SLIC). We propose an algorithm which extends SLIC to dynamically adjust the local search, adopting superpixel resolution dynamically to structure existent in the image, and thus provides for more meaningful superpixels in the same linear runtime as standard SLIC. The proposed method is evaluated against state-of-the-art techniques and improved boundary adherence and undersegmentation error are observed, whilst still remaining among the fastest algorithms which are tested. Georg Maierhofer, Daniel Heydecker, Angelica I. Avilés-Rivero, Samar M. Alsaleh, Carola-Bibiane Schönlieb |
ICIP | 5 |
| 2018 | Local Convergence Properties of SAGA/Prox-SVRG and AccelerationabstractIn this paper, we present a local convergence anal- ysis for a class of stochastic optimisation meth- ods: the proximal variance reduced stochastic gradient methods, and mainly focus on SAGA (Defazio et al., 2014) and Prox-SVRG (Xiao & Zhang, 2014). Under the assumption that the non-smooth component of the optimisation prob- lem is partly smooth relative to a smooth mani- fold, we present a unified framework for the local convergence analysis of SAGA/Prox-SVRG: (i) the sequences generated by the methods are able to identify the smooth manifold in a finite num- ber of iterations; (ii) then the sequence enters a local linear convergence regime. Furthermore, we discuss various possibilities for accelerating these algorithms, including adapting to better lo- cal parameters, and applying higher-order deter- ministic/stochastic optimisation methods which can achieve super-linear convergence. Several concrete examples arising from machine learning are considered to demonstrate the obtained result. Clarice Poon, Jingwei Liang, Carola-Bibiane Schönlieb |
ICML | 3 |
| 2018 | Adversarial Regularizers in Inverse ProblemsabstractInverse Problems in medical imaging and computer vision are traditionally solved using purely model-based methods. Among those variational regularization models are one of the most popular approaches. We propose a new framework for applying data-driven approaches to inverse problems, using a neural network as a regularization functional. The network learns to discriminate between the distribution of ground truth images and the distribution of unregularized reconstructions. Once trained, the network is applied to the inverse problem by solving the corresponding variational problem. Unlike other data-based approaches for inverse problems, the algorithm can be applied even if only unsupervised training data is available. Experiments demonstrate the potential of the framework for denoising on the BSDS dataset and for computer tomography reconstruction on the LIDC dataset. Sebastian Lunz, Carola-Bibiane Schönlieb, Ozan Öktem |
NeurIPS | 2 |
| 2018 | A Variational Model for Joint Motion Estimation and Image ReconstructionabstractThe aim of this paper is to derive and analyze a variational model for the joint estimation of motion and reconstruction of image sequences, which is based on a time-continuous Eulerian motion model. The model can be set up in terms of the continuity equation or the brightness constancy equation. The analysis in this paper focuses on the latter for robust motion estimation on sequences of two-dimensional images. We rigorously prove the existence of a minimizer in a suitable function space setting. Moreover, we discuss the numerical solution of the model based on primal-dual algorithms and investigate several examples. Finally, the benefits of our model compared to existing techniques, such as sequential image reconstruction and motion estimation, are shown. Martin Burger 0001, Hendrik Dirks, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 3 |
| 2018 | Variational Image Regularization with Euler's Elastica Using a Discrete Gradient SchemeabstractThis paper concerns an optimization algorithm for unconstrained nonconvex problems where the objective function has sparse connections between the unknowns. The algorithm is based on applying a dissipation preserving numerical integrator, the Itoh--Abe discrete gradient scheme, to the gradient flow of an objective function, guaranteeing energy decrease regardless of step size. We introduce the algorithm, prove a convergence rate estimate for nonconvex problems with Lipschitz continuous gradients, and show an improved convergence rate if the objective function has sparse connections between unknowns. The algorithm is presented in serial and parallel versions. Numerical tests show its use in Euler's elastica regularized imaging problems and its convergence rate and compare the execution time of the method to that of the iPiano algorithm and the gradient descent and heavy-ball algorithms. Torbjørn Ringholm, Jasmina Lazic, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 3 |
| 2017 | Infimal Convolution of Data Discrepancies for Mixed Noise RemovalabstractWe consider the problem of image denoising in the presence of noise whose statistical properties are a combination of two different distributions. We focus on noise distributions frequently considered in applications, such as salt & pepper and Gaussian, and Gaussian and Poisson noise mixtures. We derive a variational image denoising model that features a total variation regularization term and a data discrepancy encoding the mixed noise as an infimal convolution of discrepancy terms of the single-noise distributions. We give a statistical derivation of this model by joint maximum a posteriori (MAP) estimation. Classical single-noise models are recovered asymptotically as the weighting parameters go to infinity. The numerical solution of the model is computed using second order Newton-type methods. Numerical results show the decomposition of the noise into its constituting components. The paper is furnished with several numerical experiments, and comparisons with other methods dealing with the mixed noise case are shown. Luca Calatroni, Juan Carlos de los Reyes, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 3 |
| 2017 | Guidefill: GPU Accelerated, Artist Guided Geometric Inpainting for 3D Conversion of FilmabstractThe conversion of traditional film into stereo 3D has become an important problem in the past decade. One of the main bottlenecks is a disocclusion step, which in commercial 3D conversion is usually done by teams of artists armed with a toolbox of inpainting algorithms. A current difficulty in this is that most available algorithms either are too slow for interactive use or provide no intuitive means for users to tweak the output. In this paper we present a new fast inpainting algorithm based on transporting along automatically detected splines, which the user may edit. Our algorithm is implemented on the GPU and fills the inpainting domain in successive shells that adapt their shape on the fly. In order to allocate GPU resources as efficiently as possible, we propose a parallel algorithm to track the inpainting interface as it evolves, ensuring that no resources are wasted on pixels that are not currently being worked on. Theoretical analyses of the time and processor complexity of our algorithm without and with tracking (as well as numerous numerical experiments) demonstrate the merits of the latter. Our transport mechanism is similar to the one used in coherence transport [F. Bornemann and T. März, J. Math. Imaging Vision, 28 (2007), pp. 259--278; T. März, SIAM J. Imaging Sci., 4 (2011), pp. 981--1000] but improves upon it by correcting a “kinking” phenomenon whereby extrapolated isophotes may bend at the boundary of the inpainting domain. Theoretical results explaining this phenomenon and its resolution are presented. Although our method ignores texture, in many cases this is not a problem due to the thin inpainting domains in 3D conversion. Experimental results show that our method can achieve a visual quality that is competitive with the state of the art while maintaining interactive speeds and providing the user with an intuitive interface to tweak the results. Rob Hocking, Russell MacKenzie, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 3 |
| 2017 | Learning to Diversify Deep Belief Networks for Hyperspectral Image ClassificationabstractIn the literature of remote sensing, deep models with multiple layers have demonstrated their potentials in learning the abstract and invariant features for better representation and classification of hyperspectral images. The usual supervised deep models, such as convolutional neural networks, need a large number of labeled training samples to learn their model parameters. However, the real-world hyperspectral image classification task provides only a limited number of training samples. This paper adopts another popular deep model, i.e., deep belief networks (DBNs), to deal with this problem. The DBNs allow unsupervised pretraining over unlabeled samples at first and then a supervised fine-tuning over labeled samples. But the usual pretraining and fine-tuning method would make many hidden units in the learned DBNs tend to behave very similarly or perform as “dead” (never responding) or “potential over-tolerant” (always responding) latent factors. These results could negatively affect description ability and thus classification performance of DBNs. To further improve DBN's performance, this paper develops a new diversified DBN through regularizing pretraining and fine-tuning procedures by a diversity promoting prior over latent factors. Moreover, the regularized pretraining and fine-tuning can be efficiently implemented through usual recursive greedy and back-propagation learning framework. The experiments over real-world hyperspectral images demonstrated that the diversity promoting prior in both pretraining and fine-tuning procedure lead to the learned DBNs with more diverse latent factors, which directly make the diversified DBNs obtain much better results than original DBNs and comparable or even better performances compared with other recent hyperspectral image classification methods. Ping Zhong 0001, Zhiqiang Gong, Shutao Li 0001, Carola-Bibiane Schönlieb |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2016 | A DBN-crf for spectral-spatial classification of hyperspectral dataabstractThis work shows how to improve hyperspectral image classification through using both a deep representation and contextual information. To implement this objective, this work proposes a new Conditional Random Field (CRF) model (named DBN-CRF) with potentials defined over deep features produced by the Deep Belief Networks (DBNs). The newly formulated DBN-CRF model takes advantage of strength of the DBNs in learning a good representation and the ability of CRFs to model contextual (spatial) information in both observations and labels. Within a piecewise training framework, an efficient training method is proposed to train the whole DBN-CRF model end-to-end. This means that parameters in DBN and CRF can be jointly trained and thus the proposed method can fully use the strength of both DBN and CRF. Moreover, in the proposed training method, the end-to-end training can be implemented with a standard back-propagation algorithm, avoiding the repeated inference usually involved in CRF training and thus is computationally efficient. Experiments on real-world hyperspectral data show that our method outperforms the most recent approaches in hyperspectral image classification. Ping Zhong 0001, Zhiqiang Gong, Carola-Bibiane Schönlieb |
ICPR | 3 |
| 2015 | Mapping individual trees from airborne multi-sensor imageryabstractIndividual tree species mapping is important to understand forest dynamics and species distribution patterns. Airborne LiDAR with hyperspectral imaging has been extensively used to extract biophysical traits of vegetation and detect species. However, its application for individual tree mapping is limited due to technical problems. To address the problems, this paper presents effective and efficient algorithms in terms of tackling co-alingment of LiDAR and hyperspectral datasets, classifying individual trees, thus detecting tree species and leaf chemistry from the tree mapping. Juheon Lee, Xiaohao Cai, Carola-Bibiane Schönlieb, David Coomes |
IGARSS | 3 |
| 2015 | Analysis and Application of a Nonlocal HessianabstractIn this work we introduce a formulation for a nonlocal Hessian that combines the ideas of higher-order and nonlocal regularization for image restoration, extending the idea of nonlocal gradients to higher-order derivatives. By intelligently choosing the weights, the model allows us to improve on the current state of the art higher-order method, total generalized variation, with respect to overall quality and preservation of jumps in the data. In the spirit of recent work by Brezis et al., our formulation also has analytic implications: for a suitable choice of weights it can be shown to converge to classical second-order regularizers, and in fact it allows a novel characterization of higher-order Sobolev and BV spaces. Jan Lellmann, Konstantinos Papafitsoros, Carola-Bibiane Schönlieb, Daniel Spector |
SIAM J. Imaging Sci. | 3 |
| 2015 | Nonparametric Image Registration of Airborne LiDAR, Hyperspectral and Photographic Imagery of Wooded LandscapesabstractThere is much current interest in using multisensor airborne remote sensing to monitor the structure and biodiversity of woodlands. This paper addresses the application of nonparametric (NP) image-registration techniques to precisely align images obtained from multisensor imaging, which is critical for the successful identification of individual trees using object recognition approaches. NP image registration, in particular, the technique of optimizing an objective function, containing similarity and regularization terms, provides a flexible approach for image registration. Here, we develop a NP registration approach, in which a normalized gradient field is used to quantify similarity, and curvature is used for regularization (NGF-Curv method). Using a survey of woodlands in southern Spain as an example, we show that NGF-Curv can be successful at fusing data sets when there is little prior knowledge about how the data sets are interrelated (i.e., in the absence of ground control points). The validity of NGF-Curv in airborne remote sensing is demonstrated by a series of experiments. We show that NGF-Curv is capable of aligning images precisely, making it a valuable component of algorithms designed to identify objects, such as trees, within multisensor data sets. Juheon Lee, Xiaohao Cai, Carola-Bibiane Schönlieb, David Coomes |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2015 | Variational Depth From Focus ReconstructionabstractThis paper deals with the problem of reconstructing a depth map from a sequence of differently focused images, also known as depth from focus (DFF) or shape from focus. We propose to state the DFF problem as a variational problem, including a smooth but nonconvex data fidelity term and a convex nonsmooth regularization, which makes the method robust to noise and leads to more realistic depth maps. In addition, we propose to solve the nonconvex minimization problem with a linearized alternating directions method of multipliers, allowing to minimize the energy very efficiently. A numerical comparison to classical methods on simulated as well as on real data is presented. Michael Möller 0001, Martin Benning, Carola-Bibiane Schönlieb, Daniel Cremers |
IEEE Trans. Image Process. | 3 |
| 2014 | Imaging with Kantorovich-Rubinstein DiscrepancyabstractWe propose the use of the Kantorovich--Rubinstein norm from optimal transport in imaging problems. In particular, we discuss a variational regularization model endowed with a Kantorovich--Rubinstein discrepancy term and total variation regularization in the context of image denoising and cartoon-texture decomposition. We point out connections of this approach to several other recently proposed methods such as total generalized variation and norms capturing oscillating patterns. We also show that the respective optimization problem can be turned into a convex-concave saddle point problem with simple constraints and hence can be solved by standard tools. Numerical examples exhibit interesting features and favorable performance for denoising and cartoon-texture decomposition. Jan Lellmann, Dirk A. Lorenz, Carola-Bibiane Schönlieb, Tuomo Valkonen |
SIAM J. Imaging Sci. | 3 |
| 2012 | Wavelet Decomposition Method for L2//TV-Image DeblurringabstractIn this paper, we show additional properties of the limit of a sequence produced by the subspace correction algorithm proposed by Fornasier and Schönlieb [SIAM J. Numer. Anal., 47 (2009), pp. 3397--3428] for $L_2/$TV-minimization problems. An important but missing property of such a limiting sequence in that paper is the convergence to a minimizer of the original minimization problem, which was obtained in [M. Fornasier, A. Langer, and C.-B. Schönlieb, Numer. Math., 116 (2010), pp. 645--685] with an additional condition of overlapping subdomains. We can now determine when the limit is indeed a minimizer of the original problem. Inspired by the work of Vonesch and Unser [IEEE Trans. Image Process., 18 (2009), pp. 509--523], we adapt and specify this algorithm to the case of an orthogonal wavelet space decomposition for deblurring problems and provide an equivalence condition to the convergence of such a limiting sequence to a minimizer. We also provide a counterexample of a limiting sequence by the algorithm that does not converge to a minimizer, which shows the necessity of our analysis of the minimizing algorithm. Massimo Fornasier, Andreas Langer, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 4 |
| 2009 | Cahn--Hilliard Inpainting and a Generalization for Grayvalue ImagesabstractThe Cahn–Hilliard equation is a nonlinear fourth order diffusion equation originating in material science for modeling phase separation and phase coarsening in binary alloys. The inpainting of binary images using the Cahn–Hilliard equation is a new approach in image processing. In this paper we discuss the stationary state of the proposed model and introduce a generalization for grayvalue images of bounded variation. This is realized by using subgradients of the total variation functional within the flow, which leads to structure inpainting with smooth curvature of level sets. Martin Burger 0001, Lin He 0010, Carola-Bibiane Schönlieb |
SIAM J. Imaging Sci. | 3 |