Angelica I. Avilés-Rivero

dblp:138/9507 · also Angelica I. Avilés · DBLP profile ↗
← Back
55ranked-venue papers
7as first author
44since 2021 · last 2026
0000-0002-8878-0325ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 2 first-author · 20 since 2021Artificial intelligence and machine learning · 24 · 4 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 3 first-author · 19 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 Sliding-Window Merging for Compacting Patch-Redundant Layers in LLMs
abstract
Depth-wise pruning accelerates LLM inference in resource-constrained scenarios but suffers from performance degradation due to indiscriminate removal of entire Transformer layers. This paper reveals ``Patch-Like'' redundancy across layers via correlation analysis of the outputs of different layers in reproducing kernel Hilbert space, demonstrating consecutive layers exhibit high functional similarity. Building on this observation, this paper proposes Sliding-Window Merging (SWM) - a dynamic compression method that selects consecutive layers from top to bottom using a pre-defined similarity threshold, and compacts patch-redundant layers through a parameter consolidation, thereby simplifying the model structure while maintaining its performance. Extensive experiments on LLMs with various architectures and different parameter scales show that our method outperforms existing pruning techniques in both zero-shot inference performance and retraining recovery quality after pruning. In particular, in the experiment with 35\% pruning on the Vicuna-7B model, our method achieved a 1.654\% improvement in average performance on zero-shot tasks compared to the existing method. Moreover, we further reveal the potential of combining depth pruning with width pruning to enhance the pruning effect.
Xiu Yan, Yueqi Zhou 0001, Kaihao Huang, Suzhong Fu, Angelica I. Avilés-Rivero, Chuanlong Xie, Yao Zhu 0003
AAAI8
2026 3D Wavelet-Based Structural Priors for Controlled Diffusion in Whole-Body Low-Dose PET Denoising
Peiyuan Jing, Chun-Wun Cheng, Zhenxuan Zhang, Liutao Yang, Thiago Lima 0001, Klaus Strobel, Antoine Leimgruber, Angelica I. Avilés-Rivero, Guang Yang 0006, Javier A. Montoya-Zegarra
ICPR (9)9
2026 MF2MR2: Multi-frequency fusion for accelerated multi-contrast MRI reconstruction
Lanqing Liu, Xiaohan Xing, Angelica I. Avilés-Rivero, Harry Qin
Expert Syst. Appl.4
2026 Brain foundation models with hypergraph dynamic adapter for brain disease analysis
abstract
Brain diseases, such as Alzheimer’s disease and brain tumors, present profound challenges due to their complexity and societal impact. Recent advancements in brain foundation models have shown significant promise in addressing a range of brain-related tasks. However, current brain foundation models are limited by task and data homogeneity, restricted generalization beyond segmentation or classification, and inefficient adaptation to diverse clinical tasks. In this work, we propose SAM-Brain3D, a brain-specific foundation model trained on over 66,000 brain image-label pairs across 14 MRI sub-modalities, and Hypergraph Dynamic Adapter (HyDA), a lightweight adapter for efficient and effective downstream adaptation. SAM-Brain3D captures detailed brain-specific anatomical and modality priors for segmenting diverse brain targets and broader downstream tasks. HyDA leverages hypergraphs to fuse complementary multi-modal data and dynamically generate patient-specific convolutional kernels for multi-scale feature fusion and personalized patient-wise adaptation. Together, our framework excels across a broad spectrum of brain disease segmentation and classification tasks. Extensive experiments demonstrate that our method consistently outperforms existing state-of-the-art approaches, offering a new paradigm for brain disease analysis through multi-modal, multi-scale, and dynamic foundation modeling.
Zhongying Deng, Ziyan Huang, Lipei Zhang, Angelica I. Avilés-Rivero, Chaoyu Liu, Junjun He, Zoe Kourtzi, Carola-Bibiane Schönlieb
Pattern Recognit.5
2026 From Coarse to Continuous: Progressive Refinement Implicit Neural Representation for Motion-Robust Anisotropic MRI Reconstruction
abstract
In motion-robust magnetic resonance imaging (MRI), slice-to-volume reconstruction is critical for recovering anatomically consistent 3D brain volumes from 2D slices, especially under accelerated acquisitions or patient motion. However, this task remains challenging due to hierarchical structural disruptions. It includes local detail loss from k-space undersampling, global structural aliasing caused by motion, and volumetric anisotropy. Therefore, we propose a progressive refinement implicit neural representation (PR-INR) framework. Our PR-INR unifies motion correction, structural refinement, and volumetric synthesis within a geometry-aware coordinate space. Specifically, a motion-aware diffusion module is first employed to generate coarse volumetric reconstructions that suppress motion artifacts and preserve global anatomical structures. Then, we introduce an implicit detail restoration module that performs residual refinement by aligning spatial coordinates with visual features. It corrects local structures and enhances boundary precision. Further, a voxel continuous-aware representation module represents the image as a continuous function over 3D coordinates. It enables accurate inter-slice completion and high-frequency detail recovery. We evaluate PR-INR on five public MRI datasets under various motion conditions (3% and 5% displacement), undersampling rates (4x and 8x) and slice resolutions (scale = 5). Experimental results demonstrate that PR-INR outperforms state-of-the-art methods in both quantitative reconstruction metrics and visual quality. It further shows generalization and robustness across diverse unseen domains.
Zhenxuan Zhang, Lipei Zhang, Yanqi Cheng, Zi Wang 0005, Fanwen Wang, Haosen Zhang, Yinzhe Wu 0001, Angelica I. Avilés-Rivero, Zhifan Gao, Guang Yang 0006, Peter J. Lally
IEEE Trans. Image Process.10
2026 HIBMatch: Hypergraph Information Bottleneck for Semi-Supervised Alzheimer's Progression
abstract
Alzheimer's disease progression prediction is critical for patients with early Mild Cognitive Impairment (MCI) to enable timely intervention and improve their quality of life. While existing progression prediction techniques demonstrate potential with multimodal data, they are highly limited by their reliance on labelled data and fail to account for a key element of future progression prediction: not all features extracted at the current moment may be relevant for predicting progression several years later. To address these limitations in the literature, we design a novel semi-supervised multimodal learning hypergraph architecture, termed HIBMatch, by harnessing hypergraph knowledge based on information bottleneck and consistency regularisation strategies. Firstly, our framework utilises hypergraphs to represent multimodal data, encompassing both imaging and non-imaging modalities. Secondly, to harmonise relevant information from the currently captured data for future MCI conversion prediction, we propose a Hypergraph Information Bottleneck (HIB) that discriminates against irrelevant information, thereby focusing exclusively on harmonising relevant information for future MCI conversion prediction. Thirdly, our method enforces consistency regularisation between the HIB and a discriminative classifier to enhance the robustness and generalisation capabilities of HIBMatch under both topological and feature perturbations. Finally, to fully exploit the unlabeled data, HIBMatch incorporates a cross-modal contrastive loss for data efficiency. Extensive experiments on the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset demonstrate that our proposed HIBMatch framework surpasses existing state-of-the-art methods in Alzheimer's disease prognosis.
Zhongying Deng, Angelica I. Avilés-Rivero, Zoe Kourtzi, Carola-Bibiane Schönlieb
IEEE J. Biomed. Health Informatics3
2026 Moving Beyond Functional Connectivity: Time-Series Modeling for fMRI-Based Brain Disorder Classification
abstract
Functional magnetic resonance imaging (fMRI) enables non-invasive brain disorder classification by capturing blood-oxygen-level-dependent (BOLD) signals. However, most existing methods rely on functional connectivity (FC) via Pearson correlation, which reduces 4D BOLD signals to static 2D matrices-discarding temporal dynamics and capturing only linear inter-regional relationships. In this work, we benchmark state-of-the-art temporal models (e.g., time-series models: PatchTST, TimesNet, TimeMixer) on raw BOLD signals across five public datasets. Results show these models consistently outperform traditional FC-based approaches, highlighting the value of directly modeling temporal information such as cycle-like oscillatory fluctuations and drift-like slow baseline trends. Building on this insight, we propose DeCI, a simple yet effective framework that integrates two key principles: (i) Cycle and Drift Decomposition to disentangle cycle and drift within each ROI (Region of Interest); and (ii) Channel-Independence to model each ROI separately, improving robustness and reducing overfitting. Extensive experiments demonstrate that DeCI achieves superior classification accuracy and generalization compared to both FC-based and temporal baselines. Our findings advocate for a shift toward end-to-end temporal modeling in fMRI analysis to better capture complex brain dynamics. The code is available at https://github.com/Levi-Ackman/DeCI.
Guoqi Yu, Xiaowei Hu 0001, Angelica I. Avilés-Rivero, Anqi Qiu
IEEE Trans. Medical Imaging3
2026 Training-Free Dual Hyperbolic Adapters for Better Cross-Modal Reasoning
abstract
Recent research in Vision-Language Models (VLMs) has significantly advanced our capabilities in cross-modal reasoning. However, existing methods suffer from performance degradation with domain changes or require substantial computational resources for fine-tuning in new domains. To address this issue, we develop a new adaptation method for large vision-language models, calledTraining-free Dual Hyperbolic Adapters(T-DHA). We characterize vision-language relationship between semantic concepts, which typically has a hierarchical tree structure, in the hyperbolic space instead of the traditional Euclidean space. Hyperbolic spaces exhibit exponential volume growth with radius, unlike the polynomial growth in Euclidean space. We find that this unique property is particularly effective for embedding hierarchical data structures using the Poincaré ball model, achieving significantly improved representation and discrimination power. Coupled with negative learning, it provides more accurate and robust classifications with fewer feature dimensions. Our extensive experimental results on various datasets demonstrate that the T-DHA method significantly outperforms existing state-of-the-art methods in few-shot image recognition and domain generalization tasks.
Yi Zhang 0109, Chun-Wun Cheng, Ke Yu 0004, Yushun Tang, Carola-Bibiane Schönlieb, Zhihai He, Angelica I. Avilés-Rivero
IEEE Trans. Multim.8
2025 Cross-Modal Few-Shot Learning with Second-Order Neural Ordinary Differential Equations
abstract
We introduce SONO, a novel method leveraging Second-Order Neural Ordinary Differential Equations (Second-Order NODEs) to enhance cross-modal few-shot learning. By employing a simple yet effective architecture consisting of a Second-Order NODEs model paired with a cross-modal classifier, SONO addresses the significant challenge of overfitting, which is common in few-shot scenarios due to limited training examples. Our second-order approach can approximate a broader class of functions, enhancing the model's expressive power and feature generalization capabilities. We initialize our cross-modal classifier with text embeddings derived from class-relevant prompts, streamlining training efficiency by avoiding the need for frequent text encoder processing. Additionally, we utilize text-based image augmentation, exploiting CLIP’s robust image-text correlation to enrich training data significantly. Extensive experiments across multiple datasets demonstrate that SONO outperforms existing state-of-the-art methods in few-shot learning performance.
Yi Zhang 0109, Chun-Wun Cheng, Zhihai He, Carola-Bibiane Schönlieb, Yuyan Chen, Angelica I. Avilés-Rivero
AAAI7
2025 VesSAM: Efficient Multi-Prompting for Segmenting Complex Vessel
abstract
Precise vessel segmentation is vital for clinical applications such as diagnosis and surgical planning but remains challenging due to thin, branching geometries and low texture contrast. Although foundation models such as the Segment Anything Model (SAM) show strong performance in general segmentation tasks, they remain suboptimal for vascular structures. In this work, we present VesSAM, a powerful and efficient framework tailored for 2D vessel segmentation. VesSAM integrates three core modules: a convolutional adapter that enhances local texture features, a multi-prompt encoder that fuses anatomical cues via hierarchical cross-attention, and a lightweight mask decoder that reduces jagged artifacts. We also introduce an automated pipeline to generate structured multi-prompt annotations, and curate a diverse benchmark dataset spanning 8 datasets across 5 imaging modalities. Extensive experiments show that VesSAM surpasses state-of-the-art PEFT-based SAM variants by over$\text{1 0 \%}$Dice and 13% IoU, while maintaining competitive accuracy to fully fine-tuned methods with far fewer parameters. VesSAM also generalizes well to out-of-distribution (OoD) settings, outperforming all baselines in average OoD Dice and IoU.
Suzhong Fu, Jingqi Dong, Yiming Yang 0001, Yao Zhu 0003, Min Chang Jordan Ren, Delin Deng, Angelica I. Avilés-Rivero, Shuguang Cui, Zhen Li 0026
BIBM9
2025 Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Models
abstract
Point cloud videos can faithfully capture real-world spatial geometries and temporal dynamics, which are essential for enabling intelligent agents to understand the dynamically changing world. However, designing an effective 4D backbone remains challenging, mainly due to the irregular and unordered distribution of points and temporal inconsistencies across frames. Also, recent transformer-based 4D backbones commonly suffer from large computational costs due to their quadratic complexity, particularly for long video sequences. To address these challenges, we propose a novel point cloud video understanding backbone purely based on the State Space Models (SSMs). Specifically, we first disentangle space and time in 4D video sequences and then establish the spatio-temporal correlation with the unified spatial-temporal Mamba blocks. The Intra-frame Spatial Mamba module is developed to encode locally similar geometric structures within a certain temporal stride. Subsequently, locally correlated tokens are delivered to the Inter-frame Temporal Mamba module, which integrates long-term point features across the entire video with linear complexity. Our proposed Mamba4D achieves competitive performance on the MSR-Action3D action recognition (+10.4% accuracy), HOI4D action segmentation (+0.7 F1 Score), and Synthia4D semantic segmentation (+0.19 mIoU) datasets. Mamba4D also has a significant efficiency improvement, especially for long video sequences, with 87.5% GPU memory reduction and × 5.36 speed-up. Codes are released at https://github.com/IRMVLab/Mamba4D.
Jiuming Liu, Jinru Han, Angelica I. Avilés-Rivero, Chaokang Jiang, Zhe Liu 0022, Hesheng Wang 0001
CVPR4
2025 Implicit U-KAN2.0: Dynamic, Efficient and Interpretable Medical Image Segmentation
Chun-Wun Cheng, Yanqi Cheng, Javier A. Montoya-Zegarra, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero
MICCAI (11)6
2025 D2SA: Dual-Stage Distribution and Slice Adaptation for Efficient Test-Time Adaptation in MRI Reconstruction
abstract
Variations in Magnetic resonance imaging (MRI) scanners and acquisition protocols cause distribution shifts that degrade reconstruction performance on unseen data. Test-time adaptation (TTA) offers a promising solution to address this discrepancies. However, previous single-shot TTA approaches are inefficient due to repeated training and suboptimal distributional models. Self-supervised learning methods may risk over-smoothing in scarce data scenarios. To address these challenges, we propose a novel Dual-Stage Distribution and Slice Adaptation (D2SA) via MRI implicit neural representation (MR-INR) to improve MRI reconstruction performance and efficiency, which features two stages. In the first stage, an MR-INR branch performs patient-wise distribution adaptation by learning shared representations across slices and modelling patient-specific shifts with mean and variance adjustments. In the second stage, single-slice adaptation refines the output from frozen convolutional layers with a learnable anisotropic diffusion module, preventing over-smoothing and reducing computation. Experiments across five MRI distribution shifts demonstrate that our method can integrate well with various self-supervised learning (SSL) framework, improving performance and accelerating convergence under diverse conditions.
Lipei Zhang, Zhongying Deng, Yanqi Cheng, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero
NeurIPS6
2025 Enhancing global sensitivity and uncertainty quantification in medical image reconstruction with Monte Carlo arbitrary-masked mamba
abstract
Deep learning has been extensively applied in medical image reconstruction, where Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) represent the predominant paradigms, each possessing distinct advantages and inherent limitations: CNNs exhibit linear complexity with local sensitivity, whereas ViTs demonstrate quadratic complexity with global sensitivity. The emerging Mamba has shown superiority in learning visual representation, which combines the advantages of linear scalability and global sensitivity. In this study, we introduce MambaMIR, an Arbitrary-Masked Mamba-based model with wavelet decomposition for joint medical image reconstruction and uncertainty estimation. A novel Arbitrary Scan Masking (ASM) mechanism "masks out" redundant information to introduce randomness for further uncertainty estimation. Compared to the commonly used Monte Carlo (MC) dropout, our proposed MC-ASM provides an uncertainty map without the need for hyperparameter tuning and mitigates the performance drop typically observed when applying dropout to low-level tasks. For further texture preservation and better perceptual quality, we employ the wavelet transformation into MambaMIR and explore its variant based on the Generative Adversarial Network, namely MambaMIR-GAN. Comprehensive experiments have been conducted for multiple representative medical image reconstruction tasks, demonstrating that the proposed MambaMIR and MambaMIR-GAN outperform other baseline and state-of-the-art methods in different reconstruction tasks, where MambaMIR achieves the best reconstruction fidelity and MambaMIR-GAN has the best perceptual quality. In addition, our MC-ASM provides uncertainty maps as an additional tool for clinicians, while mitigating the typical performance drop caused by the commonly used dropout.
Liutao Yang, Fanwen Wang, Yinzhe Wu 0001, Yang Nan 0002, Weiwen Wu, Chengyan Wang, Kuangyu Shi, Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb, Daoqiang Zhang, Guang Yang 0006
Medical Image Anal.9
2025 Learning homeomorphic image registration via conformal-invariant hyperelastic regularisation
Noémie Debroux, Harry Qin, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero
Medical Image Anal.6
2025 TrafficCAM: A Versatile Dataset for Traffic Flow Segmentation
abstract
Traffic flow analysis is revolutionising traffic management. By leveraging traffic flow data, traffic control bureaus could provide drivers with real-time alerts, advising the fastest routes and therefore optimising transportation logistics and reducing congestion. The existing traffic flow datasets have two major limitations. They feature a limited number of classes, usually limited to one type of vehicle, and the scarcity of unlabelled data. In this paper, we introduce a new benchmark traffic flow image dataset called TrafficCAM. Our dataset distinguishes itself by two major highlights. Firstly, TrafficCAM provides both pixel-level and instance-level semantic labelling along with a large range of types of vehicles and pedestrians. It is composed of a large and diverse set of video sequences recorded in streets from eight Indian cities with stationary cameras. Secondly, TrafficCAM aims to establish a new benchmark for developing fully-supervised tasks, and importantly, semi-supervised learning techniques. It is the first dataset that provides a vast amount of unlabelled data, helping to better capture traffic flow qualification under a low-cost annotation requirement. More precisely, our dataset has 4,364 image frames with semantic and instance annotations along with 58,689 unlabelled image frames. We validate our new dataset through a large and comprehensive range of experiments on several state-of-the-art approaches under four different settings: fully-supervised semantic and instance segmentation, and semi-supervised semantic and instance segmentation tasks. Our benchmark dataset and official toolkit are released athttps://math-ml-x.github.io/TrafficCAM/.
Zhongying Deng, Yanqi Cheng, Rihuan Ke, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero
IEEE Trans. Intell. Transp. Syst.7
2025 DiffMIC-v2: Medical Image Classification via Improved Diffusion Network
abstract
Recently, Denoising Diffusion Models have achieved outstanding success in generative image modeling and attracted significant attention in the computer vision community. Although a substantial amount of diffusion-based research has focused on generative tasks, few studies apply diffusion models to medical diagnosis. In this paper, we propose a diffusion-based network (named DiffMIC-v2) to address general medical image classification by eliminating unexpected noise and perturbations in image representations. To achieve this goal, we first devise an improved dual-conditional guidance strategy that conditions each diffusion step with multiple granularities to enhance step-wise regional attention. Furthermore, we design a novel Heterologous diffusion process that achieves efficient visual representation learning in the latent space. We evaluate the effectiveness of our DiffMIC-v2 on four medical classification tasks with different image modalities, including thoracic diseases classification on chest X-ray, placental maturity grading on ultrasound images, skin lesion classification using dermatoscopic images, and diabetic retinopathy grading using fundus images. Experimental results demonstrate that our DiffMIC-v2 outperforms state-of-the-art methods by a significant margin, which indicates the universality and effectiveness of the proposed model on multi-class and multi-label classification tasks. DiffMIC-v2 can use fewer iterations than our previous DiffMIC to obtain accurate estimations, and also achieves greater runtime efficiency with superior results. The code will be publicly available at https://github.com/scott-yjyang/DiffMICv2.
Huazhu Fu, Angelica I. Avilés-Rivero, Zhaohu Xing, Lei Zhu 0003
IEEE Trans. Medical Imaging3
2025 Contrastive Registration for Unsupervised Medical Image Segmentation
abstract
Medical image segmentation is an important task in medical imaging, as it serves as the first step for clinical diagnosis and treatment planning. While major success has been reported using deep learning supervised techniques, they assume a large and well-representative labeled set. This is a strong assumption in the medical domain where annotations are expensive, time-consuming, and inherent to human bias. To address this problem, unsupervised segmentation techniques have been proposed in the literature. Yet, none of the existing unsupervised segmentation techniques reach accuracies that come even near to the state-of-the-art of supervised segmentation methods. In this work, we present a novel optimization model framed in a new convolutional neural network (CNN)-based contrastive registration architecture for unsupervised medical image segmentation called CLMorph. The core idea of our approach is to exploit image-level registration and feature-level contrastive learning, to perform registration-based segmentation. First, we propose an architecture to capture the image-to-image transformation mapping via registration for unsupervised medical image segmentation. Second, we embed a contrastive learning mechanism in the registration architecture to enhance the discriminative capacity of the network at the feature level. We show that our proposed CLMorph technique mitigates the major drawbacks of existing unsupervised techniques. We demonstrate, through numerical and visual experiments, that our technique substantially outperforms the current state-of-the-art unsupervised segmentation methods on two major medical image datasets.
Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb
IEEE Trans. Neural Networks Learn. Syst.2
2024 Genuine Knowledge from Practice: Diffusion Test-Time Adaptation for Video Adverse Weather Removal
abstract
Real-world vision tasks frequently suffer from the appearance of unexpected adverse weather conditions, including rain, haze, snow, and raindrops. In the last decade, convolutional neural networks and vision transformers have yielded outstanding results in single-weather video removal. However, due to the absence of appropriate adaptation, most of them fail to generalize to other weather conditions. Although ViWS-Net is proposed to remove ad-verse weather conditions in videos with a single set of pre-trained weights, it is seriously blinded by seen weather at train-time and degenerates when coming to unseen weather during test-time. In this work, we introduce test-time adaptation into adverse weather removal in videos, and propose the first framework that integrates test-time adaptation into the iterative diffusion reverse process. Specifically, we devise a diffusion-based network with a novel temporal noise model to efficiently explore frame-correlated information in degraded video clips at training stage. During inference stage, we introduce a proxy task named Diffusion Tubelet Self-Calibration to learn the primer distribution of test video stream and optimize the model by approx-imating the temporal noise model for online adaptation. Experimental results, on benchmark datasets, demonstrate that our Test-Time Adaptation method with Diffusion-based network(Diff- TTA) outperforms state-of-the-art methods in terms of restoring videos degraded by seen weather conditions. Its generalizable capability is validated with unseen weather conditions in synthesized and real-world videos.
Angelica I. Avilés-Rivero, Yulun Zhang 0001, Harry Qin, Lei Zhu 0003
CVPR3
2024 Semi-supervised Video Desnowing Network via Temporal Decoupling Experts and Distribution-Driven Contrastive Regularization
Angelica I. Avilés-Rivero, Sixiang Chen, Haoyu Chen 0003, Lei Zhu 0003
ECCV (10)3
2024 HAMLET: Graph Transformer Neural Operator for Partial Differential Equations
abstract
We present a novel graph transformer framework, HAMLET, designed to address the challenges in solving partial differential equations (PDEs) using neural networks. The framework uses graph transformers with modular input encoders to directly incorporate differential equation information into the solution process. This modularity enhances parameter correspondence control, making HAMLET adaptable to PDEs of arbitrary geometries and varied input formats. Notably, HAMLET scales effectively with increasing data complexity and noise, showcasing its robustness. HAMLET is not just tailored to a single type of physical simulation, but can be applied across various domains. Moreover, it boosts model resilience and performance, especially in scenarios with limited data. We demonstrate, through extensive experiments, that our framework is capable of outperforming current techniques for PDEs.
Andrey Bryutkin, Zhongying Deng, Guang Yang 0006, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero
ICML6
2024 Revitalizing Multivariate Time Series Forecasting: Learnable Decomposition with Inter-Series Dependencies and Intra-Series Variations Modeling
abstract
Predicting multivariate time series is crucial, demanding precise modeling of intricate patterns, including inter-series dependencies and intra-series variations. Distinctive trend characteristics in each time series pose challenges, and existing methods, relying on basic moving average kernels, may struggle with the non-linear structure and complex trends in real-world data. Given that, we introduce a learnable decomposition strategy to capture dynamic trend information more reasonably. Additionally, we propose a dual attention module tailored to capture inter-series dependencies and intra-series variations simultaneously for better time series forecasting, which is implemented by channel-wise self-attention and autoregressive self-attention. To evaluate the effectiveness of our method, we conducted experiments across eight open-source datasets and compared it with the state-of-the-art methods. Through the comparison results, our $\textbf{Leddam}$ ($\textbf{LE}arnable$ $\textbf{D}ecomposition$ and $\textbf{D}ual $ $\textbf{A}ttention$ $\textbf{M}odule$) not only demonstrates significant advancements in predictive performance but also the proposed decomposition strategy can be plugged into other methods with a large performance-boosting, from 11.87% to 48.56% MSE error degradation. Code is available at this link: https://github.com/Levi-Ackman/Leddam.
Guoqi Yu, Xiaowei Hu 0001, Angelica I. Avilés-Rivero, Harry Qin
ICML4
2024 LGRNet: Local-Global Reciprocal Network for Uterine Fibroid Segmentation in Ultrasound Videos
Angelica I. Avilés-Rivero, Guang Yang 0006, Harry Qin, Lei Zhu 0003
MICCAI (4)3
2024 Biophysics Informed Pathological Regularisation for Brain Tumour Segmentation
Lipei Zhang, Yanqi Cheng, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero
MICCAI (12)5
2024 TrafficMOT: A Challenging Dataset for Multi-Object Tracking in Complex Traffic Scenarios
abstract
ACM Multimedia 2024, Melbourne, Australia, Oct 28 - Nov 1, 2024
Yanqi Cheng, Zhongying Deng, Dongdong Chen 0001, Xiaowei Hu 0001, Pietro Liò, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero
ACM Multimedia9
2024 Training-Free Feature Reconstruction with Sparse Optimization for Vision-Language Models
abstract
In this paper, we address the challenge of adapting vision-language models (VLMs) to few-shot image recognition in a training-free manner. We observe that existing methods are not able to effectively characterize the semantic relationship between support and query samples in a training-free setting. We recognize that, in the semantic feature space, the feature of the query image is a linear and sparse combination of support image features since support-query pairs are from the class and share the same small set of distinctive visual attributes. Motivated by this interesting observation, we propose a novel method called Training-free Feature ReConstruction with Sparse optimization (TaCo), which formulates the few-shot image recognition task as a feature reconstruction and sparse optimization problem. Specifically, we exploit the VLM to encode the query and support images into features. We utilize sparse optimization to reconstruct the query feature from the corresponding support features. The feature reconstruction error is then used to define the reconstruction similarity. Coupled with the text-image similarity provided by the VLM, our reconstruction similarity analysis accurately characterizes the relationship between support and query images. This results in significantly improved performance in few-shot image recognition. Our extensive experimental results on few-shot recognition demonstrate that our method outperforms existing state-of-the-art approaches by substantial margins.
Yi Zhang 0109, Ke Yu 0004, Angelica I. Avilés-Rivero, Jiyuan Jia, Yushun Tang, Zhihai He
ACM Multimedia3
2024 Guest Editorial: Special Issue on the British Machine Vision Conference 2022
Guang Yang 0006, Angelica I. Avilés-Rivero, Yingying Fang, Zhenhua Feng 0001, Gianluigi Ciocca, Yulia Hicks, Constantino Carlos Reyes-Aldasoro
Int. J. Comput. Vis.2
2024 CoNIC Challenge: Pushing the frontiers of nuclear detection, segmentation, classification and counting
abstract
Nuclear detection, segmentation and morphometric profiling are essential in helping us further understand the relationship between histology and patient outcome. To drive innovation in this area, we setup a community-wide challenge using the largest available dataset of its kind to assess nuclear segmentation and cellular composition. Our challenge, named CoNIC, stimulated the development of reproducible algorithms for cellular recognition with real-time result inspection on public leaderboards. We conducted an extensive post-challenge analysis based on the top-performing models using 1,658 whole-slide images of colon tissue. With around 700 million detected nuclei per model, associated features were used for dysplasia grading and survival analysis, where we demonstrated that the challenge's improvement over the previous state-of-the-art led to significant boosts in downstream performance. Our findings also suggest that eosinophils and neutrophils play an important role in the tumour microevironment. We release challenge models and WSI-level results to foster the development of further methods for biomarker discovery.
Simon Graham, Quoc Dang Vu, Mostafa Jahanifar, Martin Weigert 0001, Jun Zhang 0018, Sen Yang 0006, Jinxi Xiang, Josef Lorenz Rumberger, Elias Baumann, Peter Hirsch 0001, Chenyang Hong, Angelica I. Avilés-Rivero, Ayushi Jain, Heeyoung Ahn, Yiyu Hong, Hussam Azzuni, Min Xu 0009, Mohammad Yaqub, Marie-Claire Blache, Benoît Piégu, Bertrand Vernay, Tim Scherr, Moritz Böhland, Katharina Löffler, Weiqin Ying, Chixin Wang, David R. J. Snead, Shan E Ahmed Raza, Fayyaz ul Amir Afsar Minhas, Nasir M. Rajpoot
Medical Image Anal.16
2024 STADNet: Spatial-Temporal Attention-Guided Dual-Path Network for cardiac cine MRI super-resolution
Shuo Wang 0011, Yapeng Tian, Shunjie Dong, Chengyan Wang, Angelica I. Avilés-Rivero, Harry Qin
Medical Image Anal.7
2024 LaplaceNet: A Hybrid Graph-Energy Neural Network for Deep Semisupervised Classification
abstract
Semisupervised learning (SSL) has received a lot of recent attention as it alleviates the need for large amounts of labeled data which can often be expensive, requires expert knowledge, and be time consuming to collect. Recent developments in deep semisupervised classification have reached unprecedented performance and the gap between supervised and SSL is ever-decreasing. This improvement in performance has been based on the inclusion of numerous technical tricks, strong augmentation techniques, and costly optimization schemes with multiterm loss functions. We propose a new framework, LaplaceNet, for deep semisupervised classification that has a greatly reduced model complexity. We utilize a hybrid approach where pseudolabels are produced by minimizing the Laplacian energy on a graph. These pseudolabels are then used to iteratively train a neural-network backbone. Our model outperforms state-of-the-art methods for deep semisupervised classification, over several benchmark datasets. Furthermore, we consider the application of strong augmentations to neural networks theoretically and justify the use of a multisampling approach for SSL. We demonstrate, through rigorous experimentation, that a multisampling augmentation approach improves generalization and reduces the sensitivity of the network to augmentation.
Philip Sellars, Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb
IEEE Trans. Neural Networks Learn. Syst.2
2023 SCOTCH and SODA: A Transformer Video Shadow Detection Framework
abstract
Shadows in videos are difficult to detect because of the large shadow deformation between frames. In this work, we argue that accounting for shadow deformation is essential when designing a video shadow detection method. To this end, we introduce the shadow deformation attention trajectory (SODA), a new type of video self-attention module, specially designed to handle the large shadow deformations in videos. Moreover, we present a new shadow contrastive learning mechanism (SCOTCH) which aims at guiding the network to learn a unified shadow representation from massive positive shadow pairs across different videos. We demonstrate empirically the effectiveness of our two contributions in an ablation study. Furthermore, we show that SCOTCH and SODA significantly outperforms existing techniques for video shadow detection. Code is available at the project page: https://lihaoliu-cambridge.github.io/scotch_and_soda/
Jean Prost, Lei Zhu 0003, Nicolas Papadakis, Pietro Liò, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero
CVPR7
2023 Video Adverse-Weather-Component Suppression Network via Weather Messenger and Adversarial Backpropagation
abstract
Although convolutional neural networks (CNNs) have been proposed to remove adverse weather conditions in single images using a single set of pre-trained weights, they fail to restore weather videos due to the absence of temporal information. Furthermore, existing methods for removing adverse weather conditions (e.g., rain, fog, and snow) from videos can only handle one type of adverse weather. In this work, we propose the first framework for restoring videos from all adverse weather conditions by developing a video adverse-weather-component suppression network (ViWS-Net). To achieve this, we first devise a weather-agnostic video transformer encoder with multiple transformer stages. Moreover, we design a long short-term temporal modeling mechanism for weather messenger to early fuse input adjacent video frames and learn weather-specific information. We further introduce a weather discriminator with gradient reversion, to maintain the weather-invariant common information and suppress the weather-specific information in pixel features, by adversarially predicting weather types. Finally, we develop a messenger-driven video transformer decoder to retrieve the residual weather-specific feature, which is spatiotemporally aggregated with hierarchical pixel features and refined to predict the clean target frame of input videos. Experimental results, on benchmark datasets and real-world weather videos, demonstrate that our ViWS-Net outperforms current state-of-the-art methods in terms of restoring videos degraded by any weather condition.
Angelica I. Avilés-Rivero, Huazhu Fu, Weiming Wang 0002, Lei Zhu 0003
ICCV2
2023 CDiffMR: Can We Replace the Gaussian Noise with K-Space Undersampling for Fast MRI?
Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb, Guang Yang 0006
MICCAI (10)2
2023 DiffMIC: Dual-Guidance Diffusion Network for Medical Image Classification
Huazhu Fu, Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb, Lei Zhu 0003
MICCAI (6)3
2022 Multi-modal Hypergraph Diffusion Network with Dual Prior for Alzheimer Classification
Angelica I. Avilés-Rivero, Christina Runkel, Nicolas Papadakis, Zoe Kourtzi, Carola-Bibiane Schönlieb
MICCAI (3)1
2022 HERS Superpixels: Deep Affinity Learning for Hierarchical Entropy Rate Segmentation
abstract
Superpixels serve as a powerful preprocessing tool in numerous computer vision tasks. By using superpixel representation, the number of image primitives can be largely reduced by orders of magnitudes. With the rise of deep learning in recent years, a few works have attempted to feed deeply learned features / graphs into existing classical superpixel techniques. However, none of them are able to produce superpixels in near real-time, which is crucial to the applicability of superpixels in practice. In this work, we propose a two-stage graph-based framework for superpixel segmentation. In the first stage, we introduce an efficient Deep Affinity Learning (DAL) network that learns pairwise pixel affinities by aggregating multi-scale information. In the second stage, we propose a highly efficient superpixel method called Hierarchical Entropy Rate Segmentation (HERS). Using the learned affinities from the first stage, HERS builds a hierarchical tree structure that can produce any number of highly adaptive superpixels instantaneously. We demonstrate, through visual and numerical experiments, the effectiveness and efficiency of our method compared to various state-of-the-art superpixel methods.1
Hankui Peng, Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb
WACV2
2022 TFPnP: Tuning-free Plug-and-Play Proximal Algorithms with Applications to Inverse Imaging Problems
abstract
Plug-and-Play (PnP) is a non-convex optimization framework that combines proximal algorithms, for example, the alternating direction method of multipliers (ADMM), with advanced denoising priors. Over the past few years, great empirical success has been obtained by PnP algorithms, especially for the ones that integrate deep learning-based denoisers. However, a key problem of PnP approaches is the need for manual parameter tweaking which is essential to obtain high-quality results across the high discrepancy in imaging conditions and varying scene content. In this work, we present a class of tuning-free PnP proximal algorithms that can determine parameters such as denoising strength, termination time, and other optimization-specific parameters automatically. A core part of our approach is a policy network for automated parameter search which can be effectively learned via a mixture of model-free and model-based deep reinforcement learning strategies. We demonstrate, through rigorous numerical and visual experiments, that the learned policy can customize parameters to different settings, and is often more efficient and effective than existing handcrafted criteria. Moreover, we discuss several practical considerations of PnP denoisers, which together with our learned policy yield state-of-the-art results. This advanced performance is prevalent on both linear and nonlinear exemplar inverse imaging problems, and in particular shows promising results on compressed sensing MRI, sparse-view CT, single-photon imaging, and phase retrieval.
Kaixuan Wei, Angelica I. Avilés-Rivero, Jingwei Liang, Ying Fu 0001, Hua Huang 0001, Carola-Bibiane Schönlieb
J. Mach. Learn. Res.2
2022 Beyond fine-tuning: Classifying high resolution mammograms using function-preserving transformations
abstract
The task of classifying mammograms is very challenging because the lesion is usually small in the high resolution image. The current state-of-the-art approaches for medical image classification rely on using the de-facto method for convolutional neural networks-fine-tuning. However, there are fundamental differences between natural images and medical images, which based on existing evidence from the literature, limits the overall performance gain when designed with algorithmic approaches. In this paper, we propose to go beyond fine-tuning by introducing a novel framework called MorphHR, in which we highlight a new transfer learning scheme. The idea behind the proposed framework is to integrate function-preserving transformations, for any continuous non-linear activation neurons, to internally regularise the network for improving mammograms classification. The proposed solution offers two major advantages over the existing techniques. Firstly and unlike fine-tuning, the proposed approach allows for modifying not only the last few layers but also several of the first ones on a deep ConvNet. By doing this, we can design the network front to be suitable for learning domain specific features. Secondly, the proposed scheme is scalable to hardware. Therefore, one can fit high resolution images on standard GPU memory. We show that by using high resolution images, one prevents losing relevant information. We demonstrate, through numerical and visual experiments, that the proposed approach yields to a significant improvement in the classification performance over state-of-the-art techniques, and is indeed on a par with radiology experts. Moreover and for generalisation purposes, we show the effectiveness of the proposed learning scheme on another large dataset, the ChestX-ray14, surpassing current state-of-the-art techniques.
Angelica I. Avilés-Rivero, Shuo Wang 0011, Yuan Huang 0009, Fiona J. Gilbert, Carola-Bibiane Schönlieb, Chang Wen Chen
Medical Image Anal.2
2022 GraphXCOVID: Explainable deep graph diffusion pseudo-Labelling for identifying COVID-19 on chest X-rays
Angelica I. Avilés-Rivero, Philip Sellars, Carola-Bibiane Schönlieb, Nicolas Papadakis
Pattern Recognit.1
2022 A Three-Stage Self-Training Framework for Semi-Supervised Semantic Segmentation
abstract
Semantic segmentation has been widely investigated in the community, in which state-of-the-art techniques are based on supervised models. Those models have reported unprecedented performance at the cost of requiring a large set of high quality segmentation masks for training. Obtaining such annotations is highly expensive and time consuming, in particular, in semantic segmentation where pixel-level annotations are required. In this work, we address this problem by proposing a holistic solution framed as a self-training framework for semi-supervised semantic segmentation. The key idea of our technique is the extraction of the pseudo-mask information on unlabelled data whilst enforcing segmentation consistency in a multi-task fashion. We achieve this through a three-stage solution. Firstly, a segmentation network is trained using the labelled data only and rough pseudo-masks are generated for all images. Secondly, we decrease the uncertainty of the pseudo-mask by using a multi-task model that enforces consistency and that exploits the rich statistical information of the data. Finally, the segmentation model is trained by taking into account the information of the higher quality pseudo-masks. We compare our approach against existing semi-supervised semantic segmentation methods and demonstrate state-of-the-art performance with extensive experiments.
Rihuan Ke, Angelica I. Avilés-Rivero, Saurabh Pandey, Saikumar Reddy, Carola-Bibiane Schönlieb
IEEE Trans. Image Process.2
2021 Compressed sensing plus motion (CS + M): A new perspective for improving undersampled MR image reconstruction
Angelica I. Avilés-Rivero, Noémie Debroux, Guy B. Williams, Martin J. Graves, Carola-Bibiane Schönlieb
Medical Image Anal.1
2021 Variational multi-task MRI reconstruction: Joint reconstruction, registration and super-resolution
Veronica Corona, Angelica I. Avilés-Rivero, Noémie Debroux, Carole Le Guyader, Carola-Bibiane Schönlieb
Medical Image Anal.2
2021 Rethinking medical image reconstruction via shape prior, going deeper and faster: Deep joint indirect registration and reconstruction
Jiulong Liu, Angelica I. Avilés-Rivero, Hui Ji 0002, Carola-Bibiane Schönlieb
Medical Image Anal.2
2021 Dynamic spectral residual superpixels
Jianchao Zhang, Angelica I. Avilés-Rivero, Daniel Heydecker, Xiaosheng Zhuang, Raymond Chan 0001, Carola-Bibiane Schönlieb
Pattern Recognit.2
2020 Tuning-free Plug-and-Play Proximal Algorithm for Inverse Imaging Problems
abstract
Plug-and-play (PnP) is a non-convex framework that combines ADMM or other proximal algorithms with advanced denoiser priors. Recently, PnP has achieved great empirical success, especially with the integration of deep learning-based denoisers. However, a key problem of PnP based approaches is that they require manual parameter tweaking. It is necessary to obtain high-quality results across the high discrepancy in terms of imaging conditions and varying scene content. In this work, we present a tuning-free PnP proximal algorithm, which can automatically determine the internal parameters including the penalty parameter, the denoising strength and the terminal time. A key part of our approach is to develop a policy network for automatic search of parameters, which can be effectively learned via mixed model-free and model-based deep reinforcement learning. We demonstrate, through numerical and visual experiments, that the learned policy can customize different parameters for different states, and often more efficient and effective than existing handcrafted criteria. Moreover, we discuss the practical considerations of the plugged denoisers, which together with our learned policy yield state-of-the-art results. This is prevalent on both linear and nonlinear exemplary inverse imaging problems, and in particular, we show promising results on Compressed Sensing MRI and phase retrieval.
Kaixuan Wei, Angelica I. Avilés-Rivero, Jingwei Liang, Ying Fu 0001, Carola-Bibiane Schönlieb, Hua Huang 0001
ICML2
2020 Superpixel Contracted Graph-Based Learning for Hyperspectral Image Classification
abstract
A central problem in hyperspectral image (HSI) classification is obtaining high classification accuracy when using a limited amount of labeled data. In this article we present a novel graph-based semi-supervised framework to tackle this problem. Our framework uses a superpixel approach, allowing it to define meaningful local regions in HSIs, which with high probability share the same classification label. We then extract spectral and spatial features from these regions and use them to produce a contracted weighted graph-representation, where each node represents a region rather than a pixel. The graph is then fed into a graph-based semi-supervised classifier which gives the final classification. We show that using superpixels in a graph representation is an effective tool for speeding up graphical classifiers applied to HSIs. We demonstrate through exhaustive quantitative and qualitative results that our proposed method produces accurate classifications when an incredibly small amount of labeled data is used. We show that our approach mitigates the major drawbacks of existing approaches, resulting in our approach outperforming several comparative state-of-the-art techniques.
Philip Sellars, Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb
IEEE Trans. Geosci. Remote. Sens.2
2020 Controllable Image Processing via Adaptive FilterBank Pyramid
abstract
Traditional image processing operators often provide some control parameters to tweak the final results. Recently, different convolutional neural networks have been used to approximate or improve these operators. However, in those methods, one single model can only handle one operator of a specific parameter value and does not support parameter tuning. In this paper, we propose a new plugin module, “Adaptive Filterbank Pyramid”, which can be inserted into a backbone network to support multiple operators and continuous parameter tuning. Our module explicitly represents one operator with one filterbank pyramid. To generate the results of a specific operator, the corresponding filterbank pyramid is convolved with the intermediate feature pyramid produced by the backbone network. The weights of the filterbank pyramid are directly regressed by another sub-network, which is jointly trained with the backbone network and adapted to the input parameter, thus enabling continuous parameter tuning. We applied the proposed module for a large variety of image processing tasks, including image smoothing, image denoising, image deblocking, image enhancement and neural style transfer. Experiments show that our method is generalized to different types of image processing tasks and different backbone network structures. Compared to the single-operator-single-parameter baseline, our method can produce comparable results but is significantly more efficient in both training and testing.
Dongdong Chen 0001, Qingnan Fan, Jing Liao 0001, Angelica I. Avilés-Rivero, Lu Yuan 0001, Nenghai Yu, Gang Hua 0001
IEEE Trans. Image Process.4
2019 RainFlow: Optical Flow Under Rain Streaks and Rain Veiling Effect
abstract
Optical flow in heavy rainy scenes is challenging due to the presence of both rain steaks and rain veiling effect, which break the existing optical flow constraints. Concerning this, we propose a deep-learning based optical flow method designed to handle heavy rain. We introduce a feature multiplier in our network that transforms the features of an image affected by the rain veiling effect into features that are less affected by it, which we call veiling-invariant features. We establish a new mapping operation in the feature space to produce streak-invariant features. The operation is based on a feature pyramid structure of the input images, and the basic idea is to preserve the chromatic features of the background scenes while canceling the rain-streak patterns. Both the veiling-invariant and streak-invariant features are computed and optimized automatically based on the the accuracy of our optical flow estimation. Our network is end-to-end, and handles both rain streaks and the veiling effect in an integrated framework. Extensive experiments show the effectiveness of our method, which outperforms the state of the art method and other baseline methods. We also show that our network can robustly maintain good performance on clean (no rain) images even though it is trained under rain image data.
Ruoteng Li, Robby T. Tan, Loong Fah Cheong, Angelica I. Avilés-Rivero, Qingnan Fan, Carola-Bibiane Schönlieb
ICCV4
2019 Semi-Supervised Learning with Graphs: Covariance Based Superpixels For Hyperspectral Image Classification
abstract
In this paper, we present a graph-based semi-supervised framework for hyperspectral image classification. We first introduce a novel superpixel algorithm based on the spectral covariance matrix representation of pixels to provide a better representation of our data. We then construct a superpixel graph, based on carefully considered feature vectors, before performing classification. We demonstrate, through a set of experimental results using two benchmarking datasets, that our approach outperforms three state-of-the-art classification frameworks, especially when a extremely small amount of labelled data is used.
Philip Sellars, Angelica I. Avilés-Rivero, Nicolas Papadakis, David Coomes, Anita Faul, Carola-Bibiane Schönlieb
IGARSS2
2019 GraphX $$^\mathbf{\small NET } -$$ -Chest X-Ray Classification Under Extreme Minimal Supervision
Angelica I. Avilés-Rivero, Nicolas Papadakis, Ruoteng Li, Philip Sellars, Qingnan Fan, Robby T. Tan, Carola-Bibiane Schönlieb
MICCAI (6)1
2019 Mirror, Mirror, on the Wall, Who's Got the Clearest Image of Them All? - A Tailored Approach to Single Image Reflection Removal
abstract
Removing reflection artefacts from a single image is a problem of both theoretical and practical interest, which still presents challenges because of the massively ill-posed nature of the problem. In this paper, we propose a technique based on a novel optimization problem. First, we introduce a simple user interaction scheme, which helps minimize information loss in the reflection-free regions. Second, we introduce an H2fidelity term, which preserves fine detail while enforcing the global color similarity. We show that this combination allows us to mitigate the shortcomings in structure and color preservation, which presents some of the most prominent drawbacks in the existing methods for reflection removal. We demonstrate, through numerical and visual experiments, that our method is able to outperform the state-of-the-art model-based methods and compete with recent deep-learning approaches.
Daniel Heydecker, Georg Maierhofer, Angelica I. Avilés-Rivero, Qingnan Fan, Dongdong Chen 0001, Carola-Bibiane Schönlieb, Sabine Süsstrunk
IEEE Trans. Image Process.3
2018 Peekaboo-Where are the Objects? Structure Adjusting Superpixels
abstract
This paper addresses the search for a fast and meaningful image segmentation in the context of k-means clustering. The proposed method builds on a widely-used local version of Lloyd's algorithm, called Simple Linear Iterative Clustering (SLIC). We propose an algorithm which extends SLIC to dynamically adjust the local search, adopting superpixel resolution dynamically to structure existent in the image, and thus provides for more meaningful superpixels in the same linear runtime as standard SLIC. The proposed method is evaluated against state-of-the-art techniques and improved boundary adherence and undersegmentation error are observed, whilst still remaining among the fastest algorithms which are tested.
Georg Maierhofer, Daniel Heydecker, Angelica I. Avilés-Rivero, Samar M. Alsaleh, Carola-Bibiane Schönlieb
ICIP3
2018 Sensory Substitution for Force Feedback Recovery: A Perception Experimental Study
abstract
Robotic-assisted surgeries are commonly used today as a more efficient alternative to traditional surgical options. Both surgeons and patients benefit from those systems, as they offer many advantages, including less trauma and blood loss, fewer complications, and better ergonomics. However, a remaining limitation of currently available surgical systems is the lack of force feedback due to the teleoperation setting, which prevents direct interaction with the patient. Once the force information is obtained by either a sensing device or indirectly through vision-based force estimation, a concern arises on how to transmit this information to the surgeon. An attractive alternative is sensory substitution, which allows transcoding information from one sensory modality to present it in a different sensory modality. In the current work, we used visual feedback to convey interaction forces to the surgeon. Our overarching goal was to address the following question: How should interaction forces be displayed to support efficient comprehension by the surgeon without interfering with the surgeon’s perception and workflow during surgery? Until now, the use the visual modality for force feedback has not been carefully evaluated. For this reason, we conducted an experimental study with two aims: (1) to demonstrate the potential benefits of using this modality and (2) to understand the surgeons’ perceptual preferences. The results derived from our study of 28 surgeons revealed a strong positive acceptance of the users (96%) using this modality. Moreover, we found that for surgeons to easily interpret the information, their mental model must be considered, meaning that the design of the visualizations should fit the perceptual and cognitive abilities of the end user. To our knowledge, this is the first time that these principles have been analyzed for exploring sensory substitution in medical robotics. Finally, we provide user-centered recommendations for the design of visual displays for robotic surgical systems.
Angelica I. Avilés-Rivero, Samar M. Alsaleh, John Philbeck, Stella P. Raventos, Naji Younes, James K. Hahn, Alicia Casals
ACM Trans. Appl. Percept.1
2017 Sight to touch: 3D diffeomorphic deformation recovery with mixture components for perceiving forces in robotic-assisted surgery
abstract
Robotic-assisted minimally invasive surgical systems suffer from one major limitation which is the lack of interaction forces feedback. The restricted sense of touch hinders the surgeons' performance and reduces their dexterity and precision during a procedure. In this work, we present a sensory substitution approach that relies on visual stimuli to transmit the tool-tissue interaction forces to the operating surgeon. Our approach combines a 3D diffeomorphic deformation mapping with a generative model to precisely label the force level. The main highlights of our approach are that the use of diffeomorphic transformation ensures anatomical structure preservation and the label assignment is based on a parametric form of several mixture elements. We performed experimentations on both ex-vivo and in-vivo datasets and offer careful numerical results evaluating our approach. The results show that our solution has an error measure less than 1mm in all directions and an average labeling error of 2.05%. It can also be applicable to other scenarios that require force feedback such as microsurgery, knot tying or needle-based procedures.
Angelica I. Avilés-Rivero, Samar M. Alsaleh, Alicia Casals
IROS1
2016 A Deep-Neuro-Fuzzy approach for estimating the interaction forces in Robotic surgery
abstract
Fuzzy theory was motivated by the need to create human-like solutions that allow representing vagueness and uncertainty that exist in the real-world. These capabilities have been recently further enhanced by deep learning since it allows converting complex relation between data into knowledge. In this paper, we present a novel Deep-Neuro-Fuzzy strategy for unsupervised estimation of the interaction forces in Robotic Assisted Minimally Invasive scenarios. In our approach, the capability of Neuro-Fuzzy systems for handling visual uncertainty, as well as the inherent imprecision of real physical problems, is reinforced by the advantages provided by Deep Learning methods. Experiments conducted in a realistic setting have demonstrated the superior performance of the proposed approach over existing alternatives. More precisely, our method increased the accuracy of the force estimation and compared favorably to existing state of the art approaches, offering a percentage of improvement that ranges from about 35% to 85%.
Angelica I. Avilés-Rivero, Samar M. Alsaleh, Eduard Montseny, Pilar Sobrevilla, Alicia Casals
FUZZ-IEEE1