EDBT 2026 Demo / reviewers in the wild / expert
Kevin Smith 0001
dblp:07/4241-1
· DBLP profile ↗
27ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-6163-191XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 17 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VORTEX: Challenging CNNs at texture recognition by using vision transformers with orderless and randomized token encodings
Leonardo F. S. Scabini, Kallil M. C. Zielinski, Emir Konuk, Ricardo T. Fares, Lucas Correia Ribas, Kevin Smith 0001, Odemir Martinez Bruno |
Neurocomputing | 6 |
| 2025 | On the Complexity-Faithfulness Trade-Off of Gradient-Based Explanations
Amir Mehrpanah, Matteo Gamba, Kevin Smith 0001, Hossein Azizpour |
ICCV | 3 |
| 2024 | Privacy Protection in MRI Scans Using 3D Masked Autoencoders
Lennart Alexander Van der Goten, Kevin Smith 0001 |
MICCAI (7) | 2 |
| 2024 | A Framework for Assessing Joint Human-AI Systems Based on Uncertainty Estimation
Emir Konuk, Robert Welch, Filip Christiansen, Elisabeth Epstein, Kevin Smith 0001 |
MICCAI (10) | 5 |
| 2024 | Learning from Offline Foundation Features with Tensor AugmentationsabstractWe introduce Learning from Offline Foundation Features with Tensor Augmentations (LOFF-TA), an efficient training scheme designed to harness the capabilities of foundation models in limited resource settings where their direct development is not feasible. LOFF-TA involves training a compact classifier on cached feature embeddings from a frozen foundation model, resulting in up to $37\times$ faster training and up to $26\times$ reduced GPU memory usage. Because the embeddings of augmented images would be too numerous to store, yet the augmentation process is essential for training, we propose to apply tensor augmentations to the cached embeddings of the original non-augmented images. LOFF-TA makes it possible to leverage the power of foundation models, regardless of their size, in settings with limited computational capacity. Moreover, LOFF-TA can be used to apply foundation models to high-resolution images without increasing compute. In certain scenarios, we find that training with LOFF-TA yields better results than directly fine-tuning the foundation model. Emir Konuk, Christos Matsoukas, Moein Sorkhei, Phitchapha Lertsiravarameth, Kevin Smith 0001 |
NeurIPS | 5 |
| 2024 | Bridging Generalization Gaps in High Content Imaging Through Online Self-Supervised Domain AdaptationabstractHigh Content Imaging (HCI) plays a vital role in modern drug discovery and development pipelines, facilitating various stages from hit identification to candidate drug characterization. Applying machine learning models to these datasets can prove challenging as they typically consist of multiple batches, affected by experimental variation, especially if different imaging equipment have been used. Moreover, as new data arrive, it is preferable that they are analyzed in an online fashion. To overcome this, we propose CODA, an online self-supervised domain adaptation approach. CODA divides the classifier’s role into a generic feature extractor and a task-specific model. We adapt the feature extractor’s weights to the new domain using crossbatch self-supervision while keeping the task-specific model unchanged. Our results demonstrate that this strategy significantly reduces the generalization gap, achieving up to a 300% improvement when applied to data from different labs utilizing different microscopes. CODA can be applied to new, unlabeled out-of-domain data sources of different sizes, from a single plate to multiple experimental batches. Johan Fredin Haslum, Christos Matsoukas, Karl-Johan Leuchowius, Kevin Smith 0001 |
WACV | 4 |
| 2024 | Are Natural Domain Foundation Models Useful for Medical Image Classification?abstractThe deep learning field is converging towards the use of general foundation models that can be easily adapted for diverse tasks. While this paradigm shift has become common practice within the field of natural language processing, progress has been slower in computer vision. In this paper we attempt to address this issue by investigating the transferability of various state-of-the-art foundation models to medical image classification tasks. Specifically, we evaluate the performance of five foundation models, namely Sam, Seem, Dinov2, BLIP, and OpenCLIP across four well-established medical imaging datasets. We explore different training settings to fully harness the potential of these models. Our study shows mixed results. Dinov2 consistently outperforms the standard practice of ImageNet pretraining. However, other foundation models failed to consistently beat this established baseline indicating limitations in their transferability to medical image classification tasks. Joana Palés Huix, Adithya Raju Ganeshan, Johan Fredin Haslum, Magnus Söderberg, Christos Matsoukas, Kevin Smith 0001 |
WACV | 6 |
| 2023 | PatchDropout: Economizing Vision Transformers Using Patch DropoutabstractVision transformers have demonstrated the potential to outperform CNNs in a variety of vision tasks. But the computational and memory requirements of these models prohibit their use in many applications, especially those that depend on high-resolution images, such as medical image classification. Efforts to train ViTs more efficiently are overly complicated, necessitating architectural changes or intricate training schemes. In this work, we show that standard ViT models can be efficiently trained at high resolution by randomly dropping input image patches. This simple approach, PatchDropout, reduces FLOPs and memory by at least 50% in standard natural image datasets such as ImageNet, and those savings only increase with image size. On CSAW, a high-resolution medical dataset, we observe a 5× savings in computation and memory using PatchDropout, along with a boost in performance. For practitioners with a fixed computational or memory budget, PatchDropout makes it possible to choose image resolution, hyperparameters, or model size to get the most performance out of their model. Christos Matsoukas, Fredrik Strand, Hossein Azizpour, Kevin Smith 0001 |
WACV | 5 |
| 2022 | Wide-Range MRI Artifact Removal with Transformers
Lennart Alexander Van der Goten, Kevin Smith 0001 |
BMVC | 2 |
| 2022 | What Makes Transfer Learning Work for Medical Images: Feature Reuse & Other FactorsabstractTransfer learning is a standard technique to transfer knowledge from one domain to another. For applications in medical imaging, transfer from ImageNet has become the de-facto approach, despite differences in the tasks and im-age characteristics between the domains. However, it is un-clear what factors determine whether - and to what extent- transfer learning to the medical domain is useful. The long- standing assumption that features from the source domain get reused has recently been called into question. Through a series of experiments on several medical image bench-mark datasets, we explore the relationship between transfer learning, data size, the capacity and inductive bias of the model, as well as the distance between the source and tar-get domain. Our findings suggest that transfer learning is beneficial in most cases, and we characterize the important role feature reuse plays in its success. Christos Matsoukas, Johan Fredin Haslum, Moein Sorkhei, Magnus Söderberg, Kevin Smith 0001 |
CVPR | 5 |
| 2021 | Conditional De-Identification of 3D Magnetic Resonance Images
Lennart Alexander Van der Goten, Tobias Hepp 0002, Zeynep Akata, Kevin Smith 0001 |
BMVC | 4 |
| 2020 | Explanation-Based Weakly-Supervised Learning of Visual Relations with Graph Networks
Federico Baldassarre, Kevin Smith 0001, Josephine Sullivan, Hossein Azizpour |
ECCV (28) | 2 |
| 2020 | Adding seemingly uninformative labels helps in low data regimesabstractEvidence suggests that networks trained on large datasets generalize well not solely because of the numerous training examples, but also class diversity which encourages learning of enriched features. This raises the question of whether this remains true when data is scarce - is there an advantage to learning with additional labels in low-data regimes? In this work, we consider a task that requires difficult-to-obtain expert annotations: tumor segmentation in mammography images. We show that, in low-data settings, performance can be improved by complementing the expert annotations with seemingly uninformative labels from non-expert annotators, turning the task into a multi-class problem. We reveal that these gains increase when less expert data is available, and uncover several interesting properties through further studies. We demonstrate our findings on CSAW-S, a new dataset that we introduce here, and confirm them on two public datasets. Christos Matsoukas, Albert Bou I Hernandez, Karin Dembrower, Gisele Helena Barboni Miranda, Emir Konuk, Johan Fredin Haslum, Athanasios Zouzos, Peter Lindholm, Fredrik Strand, Kevin Smith 0001 |
ICML | 11 |
| 2020 | Decoupling Inherent Risk and Early Cancer Signs in Image-Based Breast Cancer Risk Models
Hossein Azizpour, Fredrik Strand, Kevin Smith 0001 |
MICCAI (6) | 4 |
| 2018 | Bayesian Uncertainty Estimation for Batch Normalized Deep NetworksabstractWe show that training a deep network using batch normalization is equivalent to approximate inference in Bayesian models. We further demonstrate that this finding allows us to make meaningful estimates of the model uncertainty using conventional architectures, without modifications to the network or the training procedure. Our approach is thoroughly validated by measuring the quality of uncertainty in a series of empirical experiments on different tasks. It outperforms baselines with strong statistical significance, and displays competitive performance with recent Bayesian approaches. Mattias Teye, Hossein Azizpour, Kevin Smith 0001 |
ICML | 3 |
| 2015 | Learning Structured Models for Segmentation of 2-D and 3-D ImageryabstractEfficient and accurate segmentation of cellular structures in microscopic data is an essential task in medical imaging. Many state-of-the-art approaches to image segmentation use structured models whose parameters must be carefully chosen for optimal performance. A popular choice is to learn them using a large-margin framework and more specifically structured support vector machines (SSVM). Although SSVMs are appealing, they suffer from certain limitations. First, they are restricted in practice to linear kernels because the more powerful nonlinear kernels cause the learning to become prohibitively expensive. Second, they require iteratively finding the most violated constraints, which is often intractable for the loopy graphical models used in image segmentation. This requires approximation that can lead to reduced quality of learning. In this paper, we propose three novel techniques to overcome these limitations. We first introduce a method to "kernelize" the features so that a linear SSVM framework can leverage the power of nonlinear kernels without incurring much additional computational cost. Moreover, we employ a working set of constraints to increase the reliability of approximate subgradient methods and introduce a new way to select a suitable step size at each iteration. We demonstrate the strength of our approach on both 2-D and 3-D electron microscopic (EM) image data and show consistent performance improvement over state-of-the-art approaches. Aurélien Lucchi, Pablo Márquez-Neila, Carlos J. Becker, Yunpeng Li 0002, Kevin Smith 0001, Graham Knott, Pascal Fua |
IEEE Trans. Medical Imaging | 5 |
| 2012 | Structured Image Segmentation Using Kernelized Features
Aurélien Lucchi, Yunpeng Li 0002, Kevin Smith 0001, Pascal Fua |
ECCV (2) | 3 |
| 2012 | SLIC Superpixels Compared to State-of-the-Art Superpixel MethodsabstractComputer vision applications have come to rely increasingly on superpixels in recent years, but it is not always clear what constitutes a good superpixel algorithm. In an effort to understand the benefits and drawbacks of existing methods, we empirically compare five state-of-the-art superpixel algorithms for their ability to adhere to image boundaries, speed, memory efficiency, and their impact on segmentation performance. We then introduce a new superpixel algorithm, simple linear iterative clustering (SLIC), which adapts a k-means clustering approach to efficiently generate superpixels. Despite its simplicity, SLIC adheres to boundaries as well as or better than previous methods. At the same time, it is faster and more memory efficient, improves segmentation performance, and is straightforward to extend to supervoxel generation. Radhakrishna Achanta, Appu Shaji, Kevin Smith 0001, Aurélien Lucchi, Pascal Fua, Sabine Süsstrunk |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Supervoxel-Based Segmentation of Mitochondria in EM Image Stacks With Learned Shape FeaturesabstractIt is becoming increasingly clear that mitochondria play an important role in neural function. Recent studies show mitochondrial morphology to be crucial to cellular physiology and synaptic function and a link between mitochondrial defects and neuro-degenerative diseases is strongly suspected. Electron microscopy (EM), with its very high resolution in all three directions, is one of the key tools to look more closely into these issues but the huge amounts of data it produces make automated analysis necessary. State-of-the-art computer vision algorithms designed to operate on natural 2-D images tend to perform poorly when applied to EM data for a number of reasons. First, the sheer size of a typical EM volume renders most modern segmentation schemes intractable. Furthermore, most approaches ignore important shape cues, relying only on local statistics that easily become confused when confronted with noise and textures inherent in the data. Finally, the conventional assumption that strong image gradients always correspond to object boundaries is violated by the clutter of distracting membranes. In this work, we propose an automated graph partitioning scheme that addresses these issues. It reduces the computational complexity by operating on supervoxels instead of voxels, incorporates shape features capable of describing the 3-D shape of the target objects, and learns to recognize the distinctive appearance of true boundaries. Our experiments demonstrate that our approach is able to segment mitochondria at a performance level close to that of a human annotator, and outperforms a state-of-the-art 3-D segmentation technique. Aurélien Lucchi, Kevin Smith 0001, Radhakrishna Achanta, Graham Knott, Pascal Fua |
IEEE Trans. Medical Imaging | 2 |
| 2011 | Are spatial and global constraints really necessary for segmentation?abstractMany state-of-the-art segmentation algorithms rely on Markov or Conditional Random Field models designed to enforce spatial and global consistency constraints. This is often accomplished by introducing additional latent variables to the model, which can greatly increase its complexity. As a result, estimating the model parameters or computing the best maximum a posteriori (MAP) assignment becomes a computationally expensive task. In a series of experiments on the PASCAL and the MSRC datasets, we were unable to find evidence of a significant performance increase attributed to the introduction of such constraints. On the contrary, we found that similar levels of performance can be achieved using a much simpler design that essentially ignores these constraints. This more simple approach makes use of the same local and global features to leverage evidence from the image, but instead directly biases the preferences of individual pixels. While our investigation does not prove that spatial and consistency constraints are not useful in principle, it points to the conclusion that they should be validated in a larger context. Aurélien Lucchi, Yunpeng Li 0002, Xavier Boix, Kevin Smith 0001, Pascal Fua |
ICCV | 4 |
| 2010 | A Fully Automated Approach to Segmentation of Irregularly Shaped Cellular Structures in EM Images
Aurélien Lucchi, Kevin Smith 0001, Radhakrishna Achanta, Vincent Lepetit, Pascal Fua |
MICCAI (2) | 2 |
| 2009 | Fast Ray features for learning irregular shapesabstractWe introduce a new class of image features, the Ray feature set, that consider image characteristics at distant contour points, capturing information which is difficult to represent with standard feature sets. This property allows Ray features to efficiently and robustly recognize deformable or irregular shapes, such as cells in microscopic imagery. Experiments show Ray features clearly outperform other powerful features including Haar-like features and Histograms of Oriented Gradients when applied to detecting irregularly shaped neuron nuclei and mitochondria. Ray features can also provide important complementary information to Haar features for other tasks such as face detection, reducing the number of weak learners and computational cost. Ray features can be efficiently precomputed to reduce cost, just as precomputing integral images reduces the overall cost of Haar features. While Rays are slightly more expensive to precompute, their computational cost is less than that of Haar features for scanning an AdaBoost-based detector window across an image at run-time. Kevin Smith 0001, Alan Carleton, Vincent Lepetit |
ICCV | 1 |
| 2008 | General constraints for batch Multiple-Target Tracking applied to large-scale videomicroscopyabstractWhile there is a large class of Multiple-Target Tracking (MTT) problems for which batch processing is possible and desirable, batch MTT remains relatively unexplored in comparison to sequential approaches. In this paper, we give a principled probabilistic formalization of batch MTT in which we introduce two new, very general constraints that considerably help us in reaching the correct solution. First, we exploit the correlation between the appearance of a target and its motion. Second, entrances and departures of targets are encouraged to occur at the boundaries of the scene. We show how to implement these constraints in a formal and efficient manner. Our approach is applied to challenging 3-D biomedical imaging data where the number of targets is unknown and may vary, and numerous challenging tracking events occur. We demonstrate the ability of our model to simultaneously track the nuclei of over one hundred migrating neuron precursor cells in image stack series collected from a 2-photon microscope. Kevin Smith 0001, Alan Carleton, Vincent Lepetit |
CVPR | 1 |
| 2008 | Tracking the Visual Focus of Attention for a Varying Number of Wandering PeopleabstractWe define and address the problem of finding the visual focus of attention for a varying number of wandering people (VFOA-W), determining where the people's movement is unconstrained. VFOA-W estimation is a new and important problem with mplications for behavior understanding and cognitive science, as well as real-world applications. One such application, which we present in this article, monitors the attention passers-by pay to an outdoor advertisement. Our approach to the VFOA-W problem proposes a multi-person tracking solution based on a dynamic Bayesian network that simultaneously infers the (variable) number of people in a scene, their body locations, their head locations, and their head pose. For efficient inference in the resulting large variable-dimensional state-space we propose a Reversible Jump Markov Chain Monte Carlo (RJMCMC) sampling scheme, as well as a novel global observation model which determines the number of people in the scene and localizes them. We propose a Gaussian Mixture Model (GMM) and Hidden Markov Model (HMM)-based VFOA-W model which use head pose and location information to determine people's focus state. Our models are evaluated for tracking performance and ability to recognize people looking at an outdoor advertisement, with results indicating good performance on sequences where a moderate number of people pass in front of an advertisement. Kevin Smith 0001, Sileye O. Ba, Jean-Marc Odobez, Daniel Gatica-Perez |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | Tracking the multi person wandering visual focus of attentionabstractEstimating the wandering visual focus of attention (WVFOA) for multiple people is an important problem with many applications in human behavior understanding. One such application, addressed in this paper, monitors the attention of passers-by to outdoor advertisements. This paper investigates the problem of tracking the wandering visual focus-of-attention (VFOA) of multiple people, an important problem with many applications in human behavior understanding. We address the specific problem of monitoring attention to outdoor advertisements. To solve the WVFOA problem, we propose a multi-person tracking approach based on a hybrid Dynamic Bayesian Network that simultaneously infers the number of people in the scene, their body and head locations, and their head pose, in a joint state-space formulation that is amenable for person interaction modeling. The model exploits both global measurements and individual observations for the VFOA. For inference in the resulting high-dimensional state-space, we propose a trans-dimensional Markov Chain Monte Carlo (MCMC) sampling scheme, which not only handles a varying number of people, but also efficiently searches the state-space by allowing person-part state updates. Our model was rigorously evaluated for tracking and its ability to recognize when people look at an outdoor advertisement using a realistic data set. Kevin Smith 0001, Sileye O. Ba, Daniel Gatica-Perez, Jean-Marc Odobez |
ICMI | 1 |
| 2005 | Using Particles to Track Varying Numbers of Interacting PeopleabstractIn this paper, we present a Bayesian framework for the fully automatic tracking of a variable number of interacting targets using a fixed camera. This framework uses a joint multi-object state-space formulation and a trans-dimensional Markov Chain Monte Carlo (MCMC) particle filter to recursively estimates the multi-object configuration and efficiently search the state-space. We also define a global observation model comprised of color and binary measurements capable of discriminating between different numbers of objects in the scene. We present results which show that our method is capable of tracking varying numbers of people through several challenging real-world tracking situations such as full/partial occlusion and entering/leaving the scene. Kevin Smith 0001, Daniel Gatica-Perez, Jean-Marc Odobez |
CVPR (1) | 1 |
| 2004 | Order Matters: A Distributed Sampling Method for Multi-Object TrackingabstractMulti-Object tracking (MOT) is an important problem in a number of vision applications. For particle filter (PF) tracking, as the number of objects tracked increases, the search space for random sampling explodes in dimension. Partitioned sampling (PS) solves this problem by partitioning the search space, then searching each partition sequentially. However, sequential weighted resampling steps cause an impoverishment effect that increases with the number of objects. This effect depends on the specific order in which the partitions are explored, creating an erratic and undesirable performance. We propose a method to search the state space that fairly distributes these impoverishment effects between the objects by defining a set of mixture components and performing PS in each of these components using one of a small set of representative object orderings. Using synthetic and real data, we show that our method retains the overall performance and reduced computational cost of PS, while improving performance in scenes where the impoverishment effect is significant. 1 Kevin Smith 0001, Daniel Gatica-Perez |
BMVC | 1 |