VLDB 2026 Research / reviewers in the wild / expert
Alexander Krull
dblp:150/4220
· DBLP profile ↗
17ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-7778-7169ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 4 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
4 papers |
Image and video processing · 83% Computational photography and imaging · 17% | |
| Artificial intelligence
9 papers |
3D vision · 73% Generative modeling · 21% Learning paradigms · 3% |
Topics — the 23 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
object pose estimation |
1.5 | 6 | 2017 | Random forests versus Neural Networks - What's best for camera localization? · ICRA 2017 Global Hypothesis Generation for 6D Object Pose Estimation · CVPR 2017 PoseAgent: Budget-Constrained 6D Object Pose Estimation via Reinforcement Learning · CVPR 2017 |
Image and video processing › image restoration
denoising |
1.1 | 2 | 2024 | bit2bit: 1-bit quanta video reconstruction via self-supervised photon prediction · NeurIPS 2024 Noise2Void - Learning Denoising From Single Noisy Images · CVPR 2019 |
Image and video processing
image restoration |
1.0 | 2 | 2023 | μSplit: image decomposition for fluorescence microscopy · ICCV 2023 Noise2Void - Learning Denoising From Single Noisy Images · CVPR 2019 |
Computer vision › 3D vision
visual localization |
0.8 | 3 | 2017 | Random forests versus Neural Networks - What's best for camera localization? · ICRA 2017 DSAC - Differentiable RANSAC for Camera Localization · CVPR 2017 Uncertainty-Driven 6D Pose Estimation of Objects and Scenes from a Single RGB Image · CVPR 2016 |
Computer vision › 3D vision › object pose estimation
6d object pose estimation |
0.8 | 3 | 2017 | Global Hypothesis Generation for 6D Object Pose Estimation · CVPR 2017 PoseAgent: Budget-Constrained 6D Object Pose Estimation via Reinforcement Learning · CVPR 2017 Learning 6D Object Pose Estimation Using 3D Object Coordinates · ECCV (2) 2014 |
Image and video processing › image restoration › image denoising › non-gaussian noise removal
poisson noise removal |
0.8 | 1 | 2024 | bit2bit: 1-bit quanta video reconstruction via self-supervised photon prediction · NeurIPS 2024 |
Computational photography and imaging
single-photon imaging |
0.8 | 1 | 2024 | bit2bit: 1-bit quanta video reconstruction via self-supervised photon prediction · NeurIPS 2024 |
Machine learning › Generative modeling › generative model
hierarchical generative model |
0.7 | 1 | 2023 | μSplit: image decomposition for fluorescence microscopy · ICCV 2023 |
Image and video processing
image decomposition |
0.7 | 1 | 2023 | μSplit: image decomposition for fluorescence microscopy · ICCV 2023 |
Machine learning › Generative modeling
variational autoencoder |
0.5 | 1 | 2021 | Fully Unsupervised Diversity Denoising with Convolutional Variational Autoencoders · ICLR 2021 |
Image and video processing › image restoration
image denoising |
0.5 | 1 | 2021 | Fully Unsupervised Diversity Denoising with Convolutional Variational Autoencoders · ICLR 2021 |
Image and video processing › image restoration › image denoising
self-supervised image denoising |
0.4 | 1 | 2019 | Noise2Void - Learning Denoising From Single Noisy Images · CVPR 2019 |
Image and video processing › image restoration › image denoising
single image denoising |
0.4 | 1 | 2019 | Noise2Void - Learning Denoising From Single Noisy Images · CVPR 2019 |
Computer vision › 3D vision › pose estimation
robust pose estimation |
0.3 | 1 | 2017 | DSAC - Differentiable RANSAC for Camera Localization · CVPR 2017 |
Computer vision › 3D vision › visual localization
scene coordinate regression |
0.3 | 1 | 2017 | Random forests versus Neural Networks - What's best for camera localization? · ICRA 2017 |
Computational photography and imaging › event-based vision
high-speed video reconstruction |
0.2 | 1 | 2024 | bit2bit: 1-bit quanta video reconstruction via self-supervised photon prediction · NeurIPS 2024 |
Computer vision › 3D vision
analysis-by-synthesis |
0.2 | 1 | 2015 | Learning Analysis-by-Synthesis for 6D Pose Estimation in RGB-D Images · ICCV 2015 |
Bioinformatics and computational biology › bioimage informatics › bioimage analysis › microscopy image analysis
fluorescence microscopy |
0.2 | 1 | 2023 | μSplit: image decomposition for fluorescence microscopy · ICCV 2023 |
Machine learning › Learning paradigms
unsupervised learning |
0.1 | 1 | 2021 | Fully Unsupervised Diversity Denoising with Convolutional Variational Autoencoders · ICLR 2021 |
Medical and health informatics
medical imaging |
0.1 | 1 | 2019 | Noise2Void - Learning Denoising From Single Noisy Images · CVPR 2019 |
Machine learning › Reinforcement learning
policy learning |
0.1 | 1 | 2017 | PoseAgent: Budget-Constrained 6D Object Pose Estimation via Reinforcement Learning · CVPR 2017 |
Mathematical optimization › optimization under uncertainty
robust optimization |
0.1 | 1 | 2017 | DSAC - Differentiable RANSAC for Camera Localization · CVPR 2017 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2015 | Learning Analysis-by-Synthesis for 6D Pose Estimation in RGB-D Images · ICCV 2015 |
Methods — techniques the papers use, named apart from their topics
u-net · 2.0lateral contextualization · 2.0hierarchical autoencoder · 2.0hierarchical VAE · 2.0ELBO loss · 2.0convolutional neural network · 1.3convolutional variational autoencoder · 1.0reinforcement learning · 0.9RANSAC · 0.8self-supervised learning · 0.8masked loss · 0.8bernoulli lattice process · 0.8differentiable RANSAC · 0.6noise2void · 0.4noise2noise · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ShapeEmbed: a self-supervised learning framework for 2D contour quantificationabstractThe shape of objects is an important source of visual information in a wide range of applications. One of the core challenges of shape quantification is to ensure that the extracted measurements remain invariant to transformations that preserve an object’s intrinsic geometry, such as changing its size, orientation, and position in the image. In this work, we introduce ShapeEmbed, a self-supervised representation learning framework designed to encode the contour of objects in 2D images, represented as a Euclidean distance matrix, into a shape descriptor that is invariant to translation, scaling, rotation, reflection, and point indexing. Our approach overcomes the limitations of traditional shape descriptors while improving upon existing state-of-the-art autoencoder-based approaches. We demonstrate that the descriptors learned by our framework outperform their competitors in shape classification tasks on natural and biological images. We envision our approach to be of particular relevance to biological imaging applications. Anna Foix Romero, Craig Russell, Alexander Krull, Virginie Uhlmann |
NeurIPS | 3 |
| 2025 | Unsupervised Denoising for Signal-Dependent and Row-Correlated Imaging NoiseabstractAccurate analysis of microscopy images is hindered by the presence of noise. This noise is usually signaldependent and often additionally correlated along rows or columns of pixels. Current self- and unsupervised denoisers can address signal-dependent noise, but none can reliably remove noise that is also row- or column-correlated. Here, we present the first fully unsupervised deep learningbased denoiser capable of handling imaging noise that is row-correlated as well as signal-dependent. Our approach uses a Variational Autoencoder (VAE) with a specially designed autoregressive decoder. This decoder is capable of modeling row-correlated and signal-dependent noise but is incapable of independently modeling underlying clean signal. The VAE therefore produces latent variables containing only clean signal information, and these are mapped back into image space using a proposed second decoder network. Our method does not require a pre-trained noise model and can be trained from scratch using unpaired noisy data. We benchmark our approach on microscopy datatsets from a range of imaging modalities and sensor types, each with row- or column-correlated, signal-dependent noise, and show that it outperforms existing self- and unsupervised denoisers. Benjamin Salmon, Alexander Krull |
WACV | 2 |
| 2024 | WiNet: Wavelet-Based Incremental Learning for Efficient Medical Image Registration
Xinxing Cheng, Xi Jia, Wenqi Lu 0001, Qiufu Li, LinLin Shen, Alexander Krull, Jinming Duan 0001 |
MICCAI (2) | 6 |
| 2024 | bit2bit: 1-bit quanta video reconstruction via self-supervised photon predictionabstractQuanta image sensors, such as SPAD arrays, are an emerging sensor technology, producing 1-bit arrays representing photon detection events over exposures as short as a few nanoseconds. In practice, raw data are post-processed using heavy spatiotemporal binning to create more useful and interpretable images at the cost of degrading spatiotemporal resolution. In this work, we propose bit2bit, a new method for reconstructing high-quality image stacks at the original spatiotemporal resolution from sparse binary quanta image data. Inspired by recent work on Poisson denoising, we developed an algorithm that creates a dense image sequence from sparse binary photon data by predicting the photon arrival location probability distribution. However, due to the binary nature of the data, we show that the assumption of a Poisson distribution is inadequate. Instead, we model the process with a Bernoulli lattice process from the truncated Poisson. This leads to the proposal of a novel self-supervised solution based on a masked loss function. We evaluate our method using both simulated and real data. On simulated data from a conventional video, we achieve 34.35 mean PSNR with extremely photon-sparse binary input (<0.06 photons per pixel per frame). We also present a novel dataset containing a wide range of real SPAD high-speed videos under various challenging imaging conditions. The scenes cover strong/weak ambient light, strong motion, ultra-fast events, etc., which will be made available to the community, on which we demonstrate the promise of our approach. Both reconstruction quality and throughput substantially surpass the state-of-the-art methods (e.g., Quanta Burst Photography (QBP)). Our approach significantly enhances the visualization and usability of the data, enabling the application of existing analysis techniques. Yehe Liu, Alexander Krull, Hector Basevi, Ales Leonardis, Michael W. Jenkins |
NeurIPS | 2 |
| 2024 | Image Denoising and the Generative Accumulation of PhotonsabstractWe present a fresh perspective on shot noise corrupted images and noise removal. By viewing image formation as the sequential accumulation of photons on a detector grid, we show that a network trained to predict where the next photon could arrive is in fact solving the minimum mean square error (MMSE) denoising task. This new perspective allows us to make three contributions: i. We present a new strategy for self-supervised denoising, ii. We present a new method for sampling from the posterior of possible solutions by iteratively sampling and adding small numbers of photons to the image. iii. We derive a full generative model by starting this process from an empty canvas. We call this approach generative accumulation of photons (GAP). We evaluate our method quantitatively and qualitatively on 4 new fluorescence microscopy datasets, which will be made available to the community. We find that it outperforms its baselines or performs on-par. Alexander Krull, Hector Basevi, Benjamin Salmon, Andre Zeug, Franziska Müller 0004, Samuel Tonks, Leela Muppala, Ales Leonardis |
WACV | 1 |
| 2023 | μSplit: image decomposition for fluorescence microscopyabstractWe present μSplit, a dedicated approach for trained image decomposition in the context of fluorescence microscopy images. We find that best results using regular deep architectures are achieved when large image patches are used during training, making memory consumption the limiting factor to further improving performance. We therefore introduce lateral contextualization (LC), a novel meta-architecture that enables the memory efficient incorporation of large image-context, which we observe is a key ingredient to solving the image decomposition task at hand. We integrate LC with U-Nets, Hierarchical AEs, and Hierarchical VAEs, for which we formulate a modified ELBO loss. Additionally, LC enables training deeper hierarchical models than otherwise possible and, interestingly, helps to reduce tiling artefacts that are inherently impossible to avoid when using tiled VAE predictions. We apply μSplit to five decomposition tasks, one on a synthetic dataset, four others derived from real microscopy data. Our method consistently achieves best results (average improvements to the best baseline of 2.25 dB PSNR), while simultaneously requiring considerably less GPU memory. Our code and datasets can be found at https://github.com/juglab/uSplit. Ashesh Mishra, Alexander Krull, Moises Di Sante, Francesco Silvio Pasqualini, Florian Jug |
ICCV | 2 |
| 2021 | Fully Unsupervised Diversity Denoising with Convolutional Variational Autoencoders
Mangal Prakash, Alexander Krull, Florian Jug |
ICLR | 2 |
| 2019 | Noise2Void - Learning Denoising From Single Noisy ImagesabstractThe field of image denoising is currently dominated by discriminative deep learning methods that are trained on pairs of noisy input and clean target images. Recently it has been shown that such methods can also be trained without clean targets. Instead, independent pairs of noisy images can be used, in an approach known as Noise2Noise (N2N). Here, we introduce Noise2Void (N2V), a training scheme that takes this idea one step further. It does not require noisy image pairs, nor clean target images. Consequently, N2V allows us to train directly on the body of data to be denoised and can therefore be applied when other methods cannot. Especially interesting is the application to biomedical image data, where the acquisition of training targets, clean or noisy, is frequently not possible. We compare the performance of N2V to approaches that have either clean target images and/or noisy image pairs available. Intuitively, N2V cannot be expected to outperform methods that have more information available during training. Still, we observe that the denoising performance of Noise2Void drops in moderation and compares favorably to training-free denoising methods. Alexander Krull, Tim-Oliver Buchholz, Florian Jug |
CVPR | 1 |
| 2017 | DSAC - Differentiable RANSAC for Camera LocalizationabstractRANSAC is an important algorithm in robust optimization and a central building block for many computer vision applications. In recent years, traditionally hand-crafted pipelines have been replaced by deep learning pipelines, which can be trained in an end-to-end fashion. However, RANSAC has so far not been used as part of such deep learning pipelines, because its hypothesis selection procedure is non-differentiable. In this work, we present two different ways to overcome this limitation. The most promising approach is inspired by reinforcement learning, namely to replace the deterministic hypothesis selection by a probabilistic selection for which we can derive the expected loss w.r.t. to all learnable parameters. We call this approach DSAC, the differentiable counterpart of RANSAC. We apply DSAC to the problem of camera localization, where deep learning has so far failed to improve on traditional approaches. We demonstrate that by directly minimizing the expected loss of the output camera poses, robustly estimated by RANSAC, we achieve an increase in accuracy. In the future, any deep learning pipeline can use DSAC as a robust optimization component. Eric Brachmann, Alexander Krull, Sebastian Nowozin, Jamie Shotton, Frank Michel 0002, Stefan Gumhold, Carsten Rother |
CVPR | 2 |
| 2017 | PoseAgent: Budget-Constrained 6D Object Pose Estimation via Reinforcement LearningabstractState-of-the-art computer vision algorithms often achieve efficiency by making discrete choices about which hypotheses to explore next. This allows allocation of computational resources to promising candidates, however, such decisions are non-differentiable. As a result, these algorithms are hard to train in an end-to-end fashion. In this work we propose to learn an efficient algorithm for the task of 6D object pose estimation. Our system optimizes the parameters of an existing state-of-the art pose estimation system using reinforcement learning, where the pose estimation system now becomes the stochastic policy, parametrized by a CNN. Additionally, we present an efficient training algorithm that dramatically reduces computation time. We show empirically that our learned pose estimation procedure makes better use of limited resources and improves upon the state-of-the-art on a challenging dataset. Our approach enables differentiable end-to-end training of complex algorithmic pipelines and learns to make optimal use of a given computational budget. Alexander Krull, Eric Brachmann, Sebastian Nowozin, Frank Michel 0002, Jamie Shotton, Carsten Rother |
CVPR | 1 |
| 2017 | Global Hypothesis Generation for 6D Object Pose EstimationabstractThis paper addresses the task of estimating the 6D-pose of a known 3D object from a single RGB-D image. Most modern approaches solve this task in three steps: i) compute local features, ii) generate a pool of pose-hypotheses, iii) select and refine a pose from the pool. This work focuses on the second step. While all existing approaches generate the hypotheses pool via local reasoning, e.g. RANSAC or Hough-Voting, we are the first to show that global reasoning is beneficial at this stage. In particular, we formulate a novel fully-connected Conditional Random Field (CRF) that outputs a very small number of pose-hypotheses. Despite the potential functions of the CRF being non-Gaussian, we give a new, efficient two-step optimization procedure, with some guarantees for optimality. We utilize our global hypotheses generation procedure to produce results that exceed state-of-the-art for the challenging "Occluded Object Dataset". Frank Michel 0002, Alexander Kirillov, Eric Brachmann, Alexander Krull, Stefan Gumhold, Bogdan Savchynskyy, Carsten Rother |
CVPR | 4 |
| 2017 | Random forests versus Neural Networks - What's best for camera localization?abstractThis work addresses the task of camera localization in a known 3D scene given a single input RGB image. State-of-the-art approaches accomplish this in two steps: firstly, regressing for every pixel in the image its 3D scene coordinate and subsequently, using these coordinates to estimate the final 6D camera pose via RANSAC. To solve the first step. Random Forests (RFs) are typically used. On the other hand. Neural Networks (NNs) reign in many dense regression tasks, but are not test-time efficient. We ask the question: which of the two is best for camera localization? To address this, we make two method contributions: (1) a test-time efficient NN architecture which we term a ForestNet that is derived and initialized from a RF, and (2) a new fully-differentiable robust averaging technique for regression ensembles which can be trained end-to-end with a NN. Our experimental findings show that for scene coordinate regression, traditional NN architectures are superior to test-time efficient RFs and ForestNets, however, this does not translate to final 6D camera pose accuracy where RFs and ForestNets perform slightly better. To summarize, our best method, a ForestNet with a robust average, which has an equivalent fast and lightweight RF, improves over the state-of-the-art for camera localization on the 7-Scenes dataset [1]. While this work focuses on scene coordinate regression for camera localization, our innovations may also be applied to other continuous regression tasks. Daniela Massiceti, Alexander Krull, Eric Brachmann, Carsten Rother, Philip Torr 0001 |
ICRA | 2 |
| 2016 | Uncertainty-Driven 6D Pose Estimation of Objects and Scenes from a Single RGB ImageabstractIn recent years, the task of estimating the 6D pose of object instances and complete scenes, i.e. camera localization, from a single input image has received considerable attention. Consumer RGB-D cameras have made this feasible, even for difficult, texture-less objects and scenes. In this work, we show that a single RGB image is sufficient to achieve visually convincing results. Our key concept is to model and exploit the uncertainty of the system at all stages of the processing pipeline. The uncertainty comes in the form of continuous distributions over 3D object coordinates and discrete distributions over object labels. We give three technical contributions. Firstly, we develop a regularized, auto-context regression framework which iteratively reduces uncertainty in object coordinate and object label predictions. Secondly, we introduce an efficient way to marginalize object coordinate distributions over depth. This is necessary to deal with missing depth information. Thirdly, we utilize the distributions over object labels to detect multiple objects simultaneously with a fixed budget of RANSAC hypotheses. We tested our system for object pose estimation and camera localization on commonly used data sets. We see a major improvement over competing systems. Eric Brachmann, Frank Michel 0002, Alexander Krull, Michael Ying Yang, Stefan Gumhold, Carsten Rother |
CVPR | 3 |
| 2015 | Pose Estimation of Kinematic Chain Instances via Object Coordinate RegressionabstractAccurate pose estimation of object instances is a key aspect in many applications, including augmented reality or robotics. For example, a task of a domestic robot could be to fetch an item from an open drawer. The poses of both, the drawer and the item have to be known by the robot in order to fulfil the task. 6D pose estimation of rigid objects has been addressed with great success in recent years. In large part, this has been due to the advent of consumer-level RGB-D cameras, which provide rich, robust input data. However, the practical use of state-of-the-art pose estimation approaches is limited by the assumption that objects are rigid. In cluttered, domestic environments this assumption does often not hold. Examples are doors, many types of furniture, certain electronic devices and toys. A robot might encounter these items in any state of articulation. This work considers the task of one-shot pose estimation of articulated object instances from an RGB-D image. In particular, we address objects with the topology of a kinematic chain of any length, i.e. objects are composed of a chain of parts interconnected by joints. We restrict joints to either revolute joints with 1 DOF (degrees of freedom) rotational movement or prismatic joints with 1 DOF translational movement. This topology covers a wide range of common objects (see our dataset for examples). However, our approach can easily be expanded to any topology, and to joints with higher degrees of freedom. Frank Michel 0002, Alexander Krull, Eric Brachmann, Michael Ying Yang, Stefan Gumhold, Carsten Rother |
BMVC | 2 |
| 2015 | Learning Analysis-by-Synthesis for 6D Pose Estimation in RGB-D ImagesabstractAnalysis-by-synthesis has been a successful approach for many tasks in computer vision, such as 6D pose estimation of an object in an RGB-D image which is the topic of this work. The idea is to compare the observation with the output of a forward process, such as a rendered image of the object of interest in a particular pose. Due to occlusion or complicated sensor noise, it can be difficult to perform this comparison in a meaningful way. We propose an approach that "learns to compare", while taking these difficulties into account. This is done by describing the posterior density of a particular object pose with a convolutional neural network (CNN) that compares observed and rendered images. The network is trained with the maximum likelihood paradigm. We observe empirically that the CNN does not specialize to the geometry or appearance of specific objects. It can be used with objects of vastly different shapes and appearances, and in different backgrounds. Compared to state-of-the-art, we demonstrate a significant improvement on two different datasets which include a total of eleven objects, cluttered background, and heavy occlusion. Alexander Krull, Eric Brachmann, Frank Michel 0002, Michael Ying Yang, Stefan Gumhold, Carsten Rother |
ICCV | 1 |
| 2014 | 6-DOF Model Based Tracking via Object Coordinate Regression
Alexander Krull, Frank Michel 0002, Eric Brachmann, Stefan Gumhold, Stephan Ihrke, Carsten Rother |
ACCV (4) | 1 |
| 2014 | Learning 6D Object Pose Estimation Using 3D Object Coordinates
Eric Brachmann, Alexander Krull, Frank Michel 0002, Stefan Gumhold, Jamie Shotton, Carsten Rother |
ECCV (2) | 2 |