Raphael Sznitman

dblp:29/8000 · DBLP profile ↗
← Back
53ranked-venue papers
10as first author
23since 2021 · last 2025
0000-0001-6791-4753ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 7 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 29 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Iterative Deployment Exposure for Unsupervised Out-of-Distribution Detection
Lars Doorenbos, Raphael Sznitman, Pablo Márquez-Neila
MICCAI (6)2
2025 WetCat: Enabling Automated Skill Assessment in Wet-Lab Cataract Surgery Videos
abstract
To meet the growing demand for systematic surgical training, wet-lab environments have become indispensable platforms for hands-on practice in ophthalmology. Yet, traditional wet-lab training depends heavily on manual performance evaluations, which are labor-intensive, time-consuming, and often subject to variability. Recent advances in computer vision offer promising avenues for automated skill assessment, enhancing both the efficiency and objectivity of surgical education. Despite notable progress in ophthalmic surgical datasets, existing resources predominantly focus on real surgeries or isolated tasks, falling short of supporting comprehensive skill evaluation in controlled wet-lab settings. To address these limitations, we introduce WetCat, the first dataset of wet-lab cataract surgery videos specifically curated for automated skill assessment. WetCat comprises high-resolution recordings of surgeries performed by trainees on artificial eyes, featuring comprehensive phase annotations and semantic segmentations of key anatomical structures. These annotations are meticulously designed to facilitate skill assessment during the critical capsulorhexis and phacoemulsification phases, adhering to standardized surgical skill assessment frameworks. By focusing on these essential phases, WetCat enables the development of interpretable, AI-driven evaluation tools aligned with established clinical metrics. This dataset lays a strong foundation for advancing objective, scalable surgical education and sets a new benchmark for automated workflow analysis and skill assessment in ophthalmology training. The dataset and annotations are publicly available in Synapse (https://www.synapse.org/Synapse:syn66401174/files/).
Negin Ghamsarian, Raphael Sznitman, Klaus Schöffmann, Jens Kowal
ACM Multimedia2
2025 GynSurg: A Comprehensive Gynecology Laparoscopic Surgery Dataset
abstract
Recent advances in deep learning have transformed computer-assisted intervention and surgical video analysis, driving improvements not only in surgical training, intraoperative decision support, and patient outcomes, but also in postoperative documentation and surgical discovery. Central to these developments is the availability of large, high-quality annotated datasets. In gynecologic laparoscopy, surgical scene understanding and action recognition are fundamental for building intelligent systems that assist surgeons during operations and provide deeper analysis after surgery. However, existing datasets are often limited by small scale, narrow task focus, or insufficiently detailed annotations, limiting their utility for comprehensive, end-to-end workflow analysis. To address these limitations, we introduce GynSurg, the largest and most diverse multi-task dataset for gynecologic laparoscopic surgery to date. GynSurg provides rich annotations across multiple tasks, supporting applications in action recognition, semantic segmentation, surgical documentation, and discovery of novel procedural insights. We demonstrate the dataset's quality and versatility by benchmarking state-of-the-art models under a standardized training protocol. To accelerate progress in the field, we publicly release the GynSurg dataset and its annotations (https://ftp.itec.aau.at/datasets/GynSurge/).
Sahar Nasirihaghighi, Negin Ghamsarian, Leonie Peschek, Matteo Munari, Heinrich Husslein, Raphael Sznitman, Klaus Schöffmann
ACM Multimedia6
2025 SAM-DA: Decoder Adapter for Efficient Medical Domain Adaptation
abstract
This paper addresses the domain adaptation challenge for semantic segmentation in medical imaging. Despite the impressive performance of recent foundational segmentation models like SAM on natural images, they struggle with medical domain images. Beyond this, recent approaches that perform end-to-end fine-tuning of models are simply not computationally tractable. To address this, we propose a novel SAM adapter approach that minimizes the number of trainable parameters while achieving comparable performances to full fine-tuning. The proposed SAM adapter is strategically placed in the mask decoder, offering excellent and broad generalization capabilities and improved segmentation across both fully supervised and test-time domain adaptation tasks. Extensive validation on four datasets showcases the adapter's efficacy, outperforming existing methods while training less than 1% of SAM's total parameters.
Javier Gamazo Tejero, Moritz Schmid, Pablo Márquez-Neila, Martin Zinkernagel, Sebastian Wolf 0005, Raphael Sznitman
WACV6
2025 Active Learning with Context Sampling and One-vs-Rest Entropy for Semantic Segmentation
abstract
Multi-class semantic segmentation remains a corner-stone challenge in computer vision. Yet, dataset creation remains excessively demanding in time and effort, especially for specialized domains. Active Learning (AL) mit-igates this challenge by selecting data points for annotation strategically. However, existing patch-based AL methods often overlook boundary pixels' critical information, essential for accurate segmentation. We present OREAL, a novel patch-based AL method designed for multi-class semantic segmentation. OREAL enhances boundary detection by employing maximum aggregation of pixel-wise uncertainty scores. Additionally, we introduce one-vs-rest entropy, a novel uncertainty score function that computes class-wise uncertainties while achieving implicit class balancing during dataset creation. Comprehensive experiments across di-verse datasets and model architectures validate our hypothesis.
Pablo Márquez-Neila, Hedyeh Rafii-Tari, Raphael Sznitman
WACV4
2024 Learning Non-linear Invariants for Unsupervised Out-of-Distribution Detection
Lars Doorenbos, Raphael Sznitman, Pablo Márquez-Neila
ECCV (52)2
2024 StofNet: Super-Resolution Time of Flight Network
abstract
Time of Flight (ToF) is a prevalent depth sensing technology in the fields of robotics, medical imaging, and non-destructive testing. Yet, ToF sensing faces challenges from complex ambient conditions making an inverse modelling from the sparse temporal information intractable. This paper highlights the potential of modern super-resolution techniques to learn varying surroundings for a reliable and accurate ToF detection. Unlike existing models, we tailor an architecture for sub-sample precise semi-global signal localization by combining super-resolution with an efficient residual contraction block to balance between fine signal details and large scale contextual information. We consolidate research on ToF by conducting a benchmark comparison against six state-of-the-art methods for which we employ two publicly available datasets. This includes the release of our SToF-Chirp dataset captured by an airborne ultrasound transducer. Results showcase the superior performance of our proposed StofNet in terms of precision, reliability and model complexity. Our code is available at https://github.com/hahnec/stofnet.
Christopher Hahne, Michel Hayoz, Raphael Sznitman
ICASSP3
2024 Online 3D Reconstruction and Dense Tracking in Endoscopic Videos
Michel Hayoz, Christopher Hahne, Thomas Kurmann, Maximilian Allan, Guido Beldi, Daniel Candinas, Pablo Márquez-Neila, Raphael Sznitman
MICCAI (6)8
2024 Correlation-aware active learning for surgery video segmentation
abstract
Semantic segmentation is a complex task that relies heavily on large amounts of annotated image data. However, annotating such data can be time-consuming and resource-intensive, especially in the medical domain. Active Learning (AL) is a popular approach that can help to reduce this burden by iteratively selecting images for annotation to improve the model performance. In the case of video data, it is important to consider the model uncertainty and the temporal nature of the sequences when selecting images for annotation. This work proposes a novel AL strategy for surgery video segmentation, COWAL, COrrelationaWare Active Learning. Our approach involves projecting images into a latent space that has been fine-tuned using contrastive learning and then selecting a fixed number of representative images from local clusters of video frames. We demonstrate the effectiveness of this approach on two video datasets of surgical instruments and three real-world video datasets. The datasets and code will be made publicly available upon receiving necessary approvals.
Pablo Márquez-Neila, Mingyi Zheng, Hedyeh Rafii-Tari, Raphael Sznitman
WACV5
2024 RF-ULM: Ultrasound Localization Microscopy Learned From Radio-Frequency Wavefronts
abstract
In Ultrasound Localization Microscopy (ULM), achieving high-resolution images relies on the precise localization of contrast agent particles across a series of beamformed frames. However, our study uncovers an enormous potential: The process of delay-and-sum beamforming leads to an irreversible reduction of Radio-Frequency (RF) channel data, while its implications for localization remain largely unexplored. The rich contextual information embedded within RF wavefronts, including their hyperbolic shape and phase, offers great promise for guiding Deep Neural Networks (DNNs) in challenging localization scenarios. To fully exploit this data, we propose to directly localize scatterers in RF channel data. Our approach involves a custom super-resolution DNN using learned feature channel shuffling, non-maximum suppression, and a semi-global convolutional block for reliable and accurate wavefront localization. Additionally, we introduce a geometric point transformation that facilitates seamless mapping to the B-mode coordinate space. To understand the impact of beamforming on ULM, we validate the effectiveness of our method by conducting an extensive comparison with State-Of-The-Art (SOTA) techniques. We present the inaugural in vivo results from a wavefront-localizing DNN, highlighting its real-world practicality. Our findings show that RF-ULM bridges the domain shift between synthetic and real datasets, offering a considerable advantage in terms of precision and complexity. To enable the broader research community to benefit from our findings, our code and the associated SOTA methods are made available at https://github.com/hahnec/rf-ulm.
Christopher Hahne, Georges Chabouh, Arthur Chavignon, Olivier Couture, Raphael Sznitman
IEEE Trans. Medical Imaging5
2023 Logical Implications for Visual Question Answering Consistency
abstract
Despite considerable recent progress in Visual Question Answering (VQA) models, inconsistent or contradictory answers continue to cast doubt on their true reasoning capabilities. However, most proposed methods use indirect strategies or strong assumptions on pairs of questions and answers to enforce model consistency. Instead, we propose a novel strategy intended to improve model performance by directly reducing logical inconsistencies. To do this, we introduce a new consistency loss term that can be used by a wide range of the VQA models and which relies on knowing the logical relation between pairs of questions and answers. While such information is typically not available in VQA datasets, we propose to infer these logical relations using a dedicated language model and use these in our proposed consistency loss function. We conduct extensive experiments on the VQA Introspect and DME datasets and show that our method brings improvements to state-of-the-art VQA models while being robust across different architectures and settings.
Sergio Tascon-Morales, Pablo Márquez-Neila, Raphael Sznitman
CVPR3
2023 Full or Weak Annotations? An Adaptive Strategy for Budget-Constrained Annotation Campaigns
abstract
Annotating new datasets for machine learning tasks is tedious, time-consuming, and costly. For segmentation applications, the burden is particularly high as manual delin-eations of relevant image content are often extremely expensive or can only be done by experts with domain-specific knowledge. Thanks to developments in transfer learning and training with weak supervision, segmentation models can now also greatly benefit from annotations of different kinds. However, for any new domain application looking to use weak supervision, the dataset builder still needs to define a strategy to distribute full segmentation and other weak annotations. Doing so is challenging, however, as it is a priori unknown how to distribute an annotation budget for a given new dataset. To this end, we propose a novel approach to determine annotation strategies for segmentation datasets, whereby estimating what proportion of segmentation and classification annotations should be collected given a fixed budget. To do so, our method sequentially determines proportions of segmentation and classification annotations to collect for budget-fractions by modeling the expected improvement of the final segmentation model. We show in our experiments that our approach yields annotations that perform very close to the optimal for a number of different annotation budgets and datasets.
Javier Gamazo Tejero, Martin Zinkernagel, Sebastian Wolf 0005, Raphael Sznitman, Pablo Márquez-Neila
CVPR4
2023 Stochastic Segmentation with Conditional Categorical Diffusion Models
abstract
Semantic segmentation has made significant progress in recent years thanks to deep neural networks, but the common objective of generating a single segmentation output that accurately matches the image's content may not be suitable for safety-critical domains such as medical diagnostics and autonomous driving. Instead, multiple possible correct segmentation maps may be required to reflect the true distribution of annotation maps. In this context, stochastic semantic segmentation methods must learn to predict conditional distributions of labels given the image, but this is challenging due to the typically multimodal distributions, high-dimensional output spaces, and limited annotation data. To address these challenges, we propose a conditional categorical diffusion model (CCDM) for semantic segmentation based on Denoising Diffusion Probabilistic Models. Our model is conditioned to the input image, enabling it to generate multiple segmentation label maps that account for the aleatoric uncertainty arising from divergent ground truth annotations. Our experimental results show that CCDM achieves state-of-the-art performance on LIDC, a stochastic semantic segmentation dataset, and outperforms established baselines on the classical segmentation dataset Cityscapes.
Lukas Zbinden, Lars Doorenbos, Theodoros Pissas, Adrian Thomas Huber, Raphael Sznitman, Pablo Márquez-Neila
ICCV5
2023 Domain Adaptation for Medical Image Segmentation Using Transformation-Invariant Self-training
Negin Ghamsarian, Javier Gamazo Tejero, Pablo Márquez-Neila, Sebastian Wolf 0005, Martin Zinkernagel, Klaus Schöffmann, Raphael Sznitman
MICCAI (1)7
2023 Geometric Ultrasound Localization Microscopy
Christopher Hahne, Raphael Sznitman
MICCAI (10)2
2023 Localized Questions in Medical Visual Question Answering
Sergio Tascon-Morales, Pablo Márquez-Neila, Raphael Sznitman
MICCAI (2)3
2023 A reinforcement learning approach for VQA validation: An application to diabetic macular edema grading
Tatiana Fountoukidou, Raphael Sznitman
Medical Image Anal.2
2022 Data Invariants to Understand Unsupervised Out-of-Distribution Detection
Lars Doorenbos, Raphael Sznitman, Pablo Márquez-Neila
ECCV (31)2
2022 DeepPyramid: Enabling Pyramid View and Deformable Pyramid Reception for Semantic Segmentation in Cataract Surgery Videos
Negin Ghamsarian, Mario Taschwer, Raphael Sznitman, Klaus Schöffmann
MICCAI (5)3
2022 Consistency-Preserving Visual Question Answering in Medical Imaging
Sergio Tascon-Morales, Pablo Márquez-Neila, Raphael Sznitman
MICCAI (8)3
2022 Surgical data science - from concepts toward clinical translation
abstract
Recent developments in data science in general and machine learning in particular have transformed the way experts envision the future of surgery. Surgical Data Science (SDS) is a new research field that aims to improve the quality of interventional healthcare through the capture, organization, analysis and modeling of data. While an increasing number of data-driven approaches and clinical applications have been studied in the fields of radiological and clinical data science, translational success stories are still lacking in surgery. In this publication, we shed light on the underlying reasons and provide a roadmap for future advances in the field. Based on an international workshop involving leading researchers in the field of SDS, we review current practice, key achievements and initiatives as well as available standards and tools for a number of topics relevant to the field, namely (1) infrastructure for data acquisition, storage and access in the presence of regulatory constraints, (2) data annotation and sharing and (3) data analytics. We further complement this technical perspective with (4) a review of currently available SDS products and the translational progress from academia and (5) a roadmap for faster clinical translation and exploitation of the full potential of SDS, based on an international multi-round Delphi process.
Lena Maier-Hein, Matthias Eisenmann, Duygu Sarikaya, Keno März, Toby Collins, Anand Malpani, Johannes Fallert, Hubertus Feußner, Stamatia Giannarou, Pietro Mascagni, Hirenkumar Nakawala, Adrian Park 0001, Carla M. Pugh, Danail Stoyanov, S. Swaroop Vedula, Kevin Cleary, Gabor Fichtinger, Germain Forestier, Bernard Gibaud, Teodor P. Grantcharov, Makoto Hashizume, Doreen Heckmann-Nötzel, Hannes Kenngott, Ron Kikinis, Lars Mündermann, Nassir Navab, Sinan Onogur, Tobias Roß, Raphael Sznitman, Russell H. Taylor, Minu Tizabi, Martin Wagner 0001, Gregory D. Hager, Thomas Neumuth, Nicolas Padoy, Justin Collins, Ines Gockel, Jan Goedeke, Daniel A. Hashimoto, Luc Joyeux, Kyle Lam, Daniel Richard Leff, Amin Madani, Hani J. Marcus, Ozanan R. Meireles, Alexander Seitel, Dogu Teber, Frank Ückert, Beat P. Müller-Stich, Pierre Jannin, Stefanie Speidel
Medical Image Anal.29
2021 CataNet: Predicting Remaining Cataract Surgery Duration
Andrés Marafioti, Michel Hayoz, Mathias Gallardo, Pablo Márquez-Neila, Sebastian Wolf 0005, Martin Zinkernagel, Raphael Sznitman
MICCAI (4)7
2021 A positive/unlabeled approach for the segmentation of medical sequences using point-wise supervision
Laurent Lejeune, Raphael Sznitman
Medical Image Anal.2
2020 A Question-Centric Model for Visual Question Answering in Medical Imaging
abstract
Deep learning methods have proven extremely effective at performing a variety of medical image analysis tasks. With their potential use in clinical routine, their lack of transparency has however been one of their few weak points, raising concerns regarding their behavior and failure modes. While most research to infer model behavior has focused on indirect strategies that estimate prediction uncertainties and visualize model support in the input image space, the ability to explicitly query a prediction model regarding its image content offers a more direct way to determine the behavior of trained models. To this end, we present a novel Visual Question Answering approach that allows an image to be queried by means of a written question. Experiments on a variety of medical and natural image datasets show that by fusing image and question features in a novel way, the proposed approach achieves an equal or higher accuracy compared to current methods.
Minh H. Vu, Tommy Löfstedt, Tufve Nyholm, Raphael Sznitman
IEEE Trans. Medical Imaging4
2019 Concept-Centric Visual Turing Tests for Method Validation
Tatiana Fountoukidou, Raphael Sznitman
MICCAI (5)2
2019 Deep Multi-label Classification in Affine Subspaces
Thomas Kurmann, Pablo Márquez-Neila, Sebastian Wolf 0005, Raphael Sznitman
MICCAI (1)4
2019 Fused Detection of Retinal Biomarkers in OCT Volumes
Thomas Kurmann, Pablo Márquez-Neila, Siqing Yu, Marion Munk, Sebastian Wolf 0005, Raphael Sznitman
MICCAI (1)6
2019 Image Data Validation for Medical Systems
Pablo Márquez-Neila, Raphael Sznitman
MICCAI (4)2
2019 Geometry in active learning for binary and multi-class image segmentation
Ksenia Konyushkova, Raphael Sznitman, Pascal Fua
Comput. Vis. Image Underst.2
2019 Patient-attentive sequential strategy for perimetry-based visual field acquisition
Serife Seda Kucur, Pablo Márquez-Neila, Mathias Abegg, Raphael Sznitman
Medical Image Anal.4
2018 A Combined Simulation and Machine Learning Approach for Image-Based Force Classification During Robotized Intravitreal Injections
Andrea Mendizabal, Tatiana Fountoukidou, Jan Hermann, Raphael Sznitman, Stephane Cotin
MICCAI (4)4
2018 Iterative multi-path tracking for video and volume segmentation with sparse point supervision
Laurent Lejeune, Jan Grossrieder, Raphael Sznitman
Medical Image Anal.3
2018 Reconstructing Evolving Tree Structures in Time Lapse Sequences by Enforcing Time-Consistency
abstract
We propose a novel approach to reconstructing curvilinear tree structures evolving over time, such as road networks in 2D aerial images or neural structures in 3D microscopy stacks acquired in vivo. To enforce temporal consistency, we simultaneously process all images in a sequence, as opposed to reconstructing structures of interest in each image independently. We formulate the problem as a Quadratic Mixed Integer Program and demonstrate the additional robustness that comes from using all available visual clues at once, instead of working frame by frame. Furthermore, when the linear structures undergo local changes over time, our approach automatically detects them.
Przemyslaw Glowacki, Miguel Amável Pinheiro, Agata Mosinska, Engin Türetken, Daniel Lebrecht, Raphael Sznitman, Anthony Holtmaat, Jan Kybic, Pascal Fua
IEEE Trans. Pattern Anal. Mach. Intell.6
2018 Articulated Multi-Instrument 2-D Pose Estimation Using Fully Convolutional Networks
abstract
Instrument detection, pose estimation, and tracking in surgical videos are an important vision component for computer-assisted interventions. While significant advances have been made in recent years, articulation detection is still a major challenge. In this paper, we propose a deep neural network for articulated multi-instrument 2-D pose estimation, which is trained on detailed annotations of endoscopic and microscopic data sets. Our model is formed by a fully convolutional detection-regression network. Joints and associations between joint pairs in our instrument model are located by the detection subnetwork and are subsequently refined through a regression subnetwork. Based on the output from the model, the poses of the instruments are inferred using maximum bipartite graph matching. Our estimation framework is powered by deep learning techniques without any direct kinematic information from a robot. Our framework is tested on single-instrument RMIT data, and also on multi-instrument EndoVis and in vivo data with promising results. In addition, the data set annotations are publicly released along with our code and model.
Xiaofei Du 0001, Thomas Kurmann, Ping-Lin Chang, Maximilian Allan, Sébastien Ourselin, Raphael Sznitman, John D. Kelly, Danail Stoyanov
IEEE Trans. Medical Imaging6
2017 Pathological OCT Retinal Layer Segmentation Using Branch Residual U-Shape Networks
Stefanos Apostolopoulos, Sandro De Zanet, Carlos Ciller, Sebastian Wolf 0005, Raphael Sznitman
MICCAI (3)5
2017 Simultaneous Recognition and Pose Estimation of Instruments in Minimally Invasive Surgery
Thomas Kurmann, Pablo Márquez-Neila, Xiaofei Du 0001, Pascal Fua, Danail Stoyanov, Sebastian Wolf 0005, Raphael Sznitman
MICCAI (2)7
2017 Learning Active Learning from Data
abstract
In this paper, we suggest a novel data-driven approach to active learning (AL). The key idea is to train a regressor that predicts the expected error reduction for a candidate sample in a particular learning state. By formulating the query selection procedure as a regression problem we are not restricted to working with existing AL heuristics; instead, we learn strategies based on experience from previous AL outcomes. We show that a strategy can be learnt either from simple synthetic 2D datasets or from a subset of domain-specific data. Our method yields strategies that work well on real data from a wide range of domains.
Ksenia Konyushkova, Raphael Sznitman, Pascal Fua
NIPS2
2016 Active Learning for Delineation of Curvilinear Structures
abstract
Many recent delineation techniques owe much of their increased effectiveness to path classification algorithms that make it possible to distinguish promising paths from others. The downside of this development is that they require annotated training data, which is tedious to produce. In this paper, we propose an Active Learning approach that considerably speeds up the annotation process. Unlike standard ones, it takes advantage of the specificities of the delineation problem. It operates on a graph and can reduce the training set size by up to 80% without compromising the reconstruction quality. We will show that our approach outperforms conventional ones on various biomedical and natural image datasets, thus showing that it is broadly applicable.
Agata Mosinska, Raphael Sznitman, Przemyslaw Glowacki, Pascal Fua
CVPR2
2015 Introducing Geometry in Active Learning for Image Segmentation
abstract
We propose an Active Learning approach to training a segmentation classifier that exploits geometric priors to streamline the annotation process in 3D image volumes. To this end, we use these priors not only to select voxels most in need of annotation but to guarantee that they lie on 2D planar patch, which makes it much easier to annotate than if they were randomly distributed in the volume. A simplified version of this approach is effective in natural 2D images. We evaluated our approach on Electron Microscopy and Magnetic Resonance image volumes, as well as on natural images. Comparing our approach against several accepted baselines demonstrates a marked performance increase.
Ksenia Konyushkova, Raphael Sznitman, Pascal Fua
ICCV2
2015 Bayesian Multiple Target Localization
abstract
We consider the problem of quickly localizing multiple targets by asking questions of the form “How many targets are within this set" while obtaining noisy answers. This setting is a generalization to multiple targets of the game of 20 questions in which only a single target is queried. We assume that the targets are points on the real line, or in a two dimensional plane for the experiments, drawn independently from a known distribution. We evaluate the performance of a policy using the expected entropy of the posterior distribution after a fixed number of questions with noisy answers. We derive a lower bound for the value of this problem and study a specific policy, named the dyadic policy. We show that this policy achieves a value which is no more than twice this lower bound when answers are noise-free, and show a more general constant factor approximation guarantee for the noisy setting. We present an empirical evaluation of this policy on simulated data for the problem of detecting multiple instances of the same object in an image. Finally, we present experiments on localizing multiple faces simultaneously on real images.
Purnima Rajan, Weidong Han 0004, Raphael Sznitman, Peter I. Frazier, Bruno Jedynak
ICML3
2015 Non-Rigid Graph Registration Using Active Testing Search
abstract
We present a new approach for matching sets of branching curvilinear structures that form graphs embedded in R2 or R3 and may be subject to deformations. Unlike earlier methods, ours does not rely on local appearance similarity nor does require a good initial alignment. Furthermore, it can cope with non-linear deformations, topological differences, and partial graphs. To handle arbitrary non-linear deformations, we use Gaussian process regressions to represent the geometrical mapping relating the two graphs. In the absence of appearance information, we iteratively establish correspondences between points, update the mapping accordingly, and use it to estimate where to find the most likely correspondences that will be used in the next step. To make the computation tractable for large graphs, the set of new potential matches considered at each iteration is not selected at random as with many RANSAC-based algorithms. Instead, we introduce a so-called Active Testing Search strategy that performs a priority search to favor the most likely matches and speed-up the process. We demonstrate the effectiveness of our approach first on synthetic cases and then on angiography data, retinal fundus images, and microscopy image stacks acquired at very different resolutions.
Eduard Serradell, Miguel Amável Pinheiro, Raphael Sznitman, Jan Kybic, Francesc Moreno-Noguer, Pascal Fua
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Reconstructing Evolving Tree Structures in Time Lapse Sequences
abstract
We propose an approach to reconstructing tree structures that evolve over time in 2D images and 3D image stacks such as neuronal axons or plant branches. Instead of reconstructing structures in each image independently, we do so for all images simultaneously to take advantage of temporal-consistency constraints. We show that this problem can be formulated as a Quadratic Mixed Integer Program and solved efficiently. The outcome of our approach is a framework that provides substantial improvements in reconstructions over traditional single time-instance formulations. Furthermore, an added benefit of our approach is the ability to automatically detect places where significant changes have occurred over time, which is challenging when considering large amounts of data.
Przemyslaw Glowacki, Miguel Amável Pinheiro, Engin Türetken, Raphael Sznitman, Daniel Lebrecht, Jan Kybic, Anthony Holtmaat, Pascal Fua
CVPR4
2014 Fast Part-Based Classification for Instrument Detection in Minimally Invasive Surgery
Raphael Sznitman, Carlos J. Becker, Pascal Fua
MICCAI (2)1
2013 Fast Object Detection with Entropy-Driven Evaluation
abstract
Cascade-style approaches to implementing ensemble classifiers can deliver significant speed-ups at test time. While highly effective, they remain challenging to tune and their overall performance depends on the availability of large validation sets to estimate rejection thresholds. These characteristics are often prohibitive and thus limit their applicability. We introduce an alternative approach to speeding-up classifier evaluation which overcomes these limitations. It involves maintaining a probability estimate of the class label at each intermediary response and stopping when the corresponding uncertainty becomes small enough. As a result, the evaluation terminates early based on the sequence of responses observed. Furthermore, it does so independently of the type of ensemble classifier used or the way it was trained. We show through extensive experimentation that our method provides 2 to 10 fold speed-ups, over existing state-of-the-art methods, at almost no loss in accuracy on a number of object classification tasks.
Raphael Sznitman, Carlos J. Becker, François Fleuret, Pascal Fua
CVPR1
2013 An Optimal Policy for Target Localization with Application to Electron Microscopy
abstract
This paper considers the task of finding a target location by making a limited number of sequential observations. Each observation results from evaluating an imperfect classifier of a chosen cost and accuracy on an interval of chosen length and position. Within a Bayesian framework, we study the problem of minimizing an objective that combines the entropy of the posterior distribution with the cost of the questions asked. In this problem, we show that the one-step lookahead policy is Bayes-optimal for any arbitrary time horizon. Moreover, this one-step lookahead policy is easy to compute and implement. We then use this policy in the context of localizing mitochondria in electron microscope images, and experimentally show that significant speed ups in acquisition can be gained, while maintaining near equal image quality at target locations, when compared to current policies.
Raphael Sznitman, Aurélien Lucchi, Peter I. Frazier, Bruno Jedynak, Pascal Fua
ICML (1)1
2013 Flash Scanning Electron Microscopy
Raphael Sznitman, Aurélien Lucchi, Marco Cantoni, Graham Knott, Pascal Fua
MICCAI (3)1
2013 Unified Detection and Tracking of Instruments during Retinal Microsurgery
abstract
Methods for tracking an object have generally fallen into two groups: tracking by detection and tracking through local optimization. The advantage of detection-based tracking is its ability to deal with target appearance and disappearance, but it does not naturally take advantage of target motion continuity during detection. The advantage of local optimization is efficiency and accuracy, but it requires additional algorithms to initialize tracking when the target is lost. To bridge these two approaches, we propose a framework for unified detection and tracking as a time-series Bayesian estimation problem. The basis of our approach is to treat both detection and tracking as a sequential entropy minimization problem, where the goal is to determine the parameters describing a target in each frame. To do this we integrate the Active Testing (AT) paradigm with Bayesian filtering, and this results in a framework capable of both detecting and tracking robustly in situations where the target object enters and leaves the field of view regularly. We demonstrate our approach on a retinal tool tracking problem and show through extensive experiments that our method provides an efficient and robust tracking solution.
Raphael Sznitman, Rogério Richa, Russell H. Taylor, Bruno Jedynak, Gregory D. Hager
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Data-Driven Visual Tracking in Retinal Microsurgery
Raphael Sznitman, Karim Ali 0002, Rogério Richa, Russell H. Taylor, Gregory D. Hager, Pascal Fua
MICCAI (2)1
2012 Efficient Scanning for EM Based Target Localization
Raphael Sznitman, Aurélien Lucchi, Natasa Pjescic-Emedji, Graham Knott, Pascal Fua
MICCAI (3)1
2011 Visual tracking using the sum of conditional variance
abstract
The goal of this paper is to introduce a direct visual tracking method based on an image similarity measure called the sum of conditional variance (SCV). The SCV was originally proposed in the medical imaging domain for registering multi-modal images. In the context of visual tracking, the SCV is invariant to non-linear illumination variations, multi-modal and computationally inexpensive. Compared to information theoretic tracking methods, it requires less iterations to converge and has a significantly larger convergence radius. The novelty in this paper is a generalization of the efficient second-order minimization formulation for tracking using the SCV, allowing us to combine the efficient second-order approximation of the Hessian with a similarity metric invariant to non-linear illumination variations. The result is a visual tracking method that copes with non-linear illumination variations without requiring the estimation of photometric correction parameters at every iteration. We demonstrate the superior performance of the proposed method through comparative studies and tracking experiments under challenging illumination conditions and rapid motions.
Rogério Richa, Raphael Sznitman, Russell H. Taylor, Gregory D. Hager
IROS2
2011 Unified Detection and Tracking in Retinal Microsurgery
Raphael Sznitman, Anasuya Basu, Rogério Richa, Jim Handa, Peter Gehlbach, Russell H. Taylor, Bruno Jedynak, Gregory D. Hager
MICCAI (1)1
2010 Adaptive Multispectral Illumination for Retinal Microsurgery
Raphael Sznitman, Diego Rother, James Handa, Peter Gehlbach, Gregory D. Hager, Russell H. Taylor
MICCAI (3)1
2010 Active Testing for Face Detection and Localization
abstract
We provide a novel search technique which uses a hierarchical model and a mutual information gain heuristic to efficiently prune the search space when localizing faces in images. We show exponential gains in computation over traditional sliding window approaches, while keeping similar performance levels.
Raphael Sznitman, Bruno Jedynak
IEEE Trans. Pattern Anal. Mach. Intell.1