VLDB 2026 Research / reviewers in the wild / expert
Stepan Tulyakov
dblp:141/9935
· DBLP profile ↗
7ranked-venue papers
6as first author
3since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Image and video processing · 81% Computational photography and imaging · 19% | |
| Artificial intelligence
3 papers |
3D vision · 52% Representation and self-supervised learning · 18% Efficient and distributed learning · 18% |
Topics — the 17 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing › video frame interpolation
event-based frame interpolation |
1.1 | 2 | 2022 | Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale Fusion · CVPR 2022 Time Lens: Event-Based Video Frame Interpolation · CVPR 2021 |
Image and video processing
motion estimation |
1.1 | 2 | 2022 | Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale Fusion · CVPR 2022 Time Lens: Event-Based Video Frame Interpolation · CVPR 2021 |
Image and video processing › motion estimation
optical flow |
1.1 | 2 | 2022 | Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale Fusion · CVPR 2022 Time Lens: Event-Based Video Frame Interpolation · CVPR 2021 |
Image and video processing
video frame interpolation |
1.1 | 2 | 2022 | Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale Fusion · CVPR 2022 Time Lens: Event-Based Video Frame Interpolation · CVPR 2021 |
Computer vision › 3D vision › 3d reconstruction › multi-view stereo
stereo reconstruction |
0.7 | 2 | 2019 | Learning an Event Sequence Embedding for Dense Event-Based Deep Stereo · ICCV 2019 Weakly Supervised Learning of Deep Metrics for Stereo Reconstruction · ICCV 2017 |
Image and video processing › image restoration
image deblurring |
0.7 | 1 | 2023 | EvShutter: Transforming Events for Unconstrained Rolling Shutter Correction · CVPR 2023 |
Computational photography and imaging › image signal processing
rolling shutter correction |
0.7 | 1 | 2023 | EvShutter: Transforming Events for Unconstrained Rolling Shutter Correction · CVPR 2023 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.4 | 2 | 2018 | Practical Deep Stereo (PDS): Toward applications-friendly deep stereo matching · NeurIPS 2018 Weakly Supervised Learning of Deep Metrics for Stereo Reconstruction · ICCV 2017 |
Computer vision › 3D vision › stereo vision
event-based stereo |
0.4 | 1 | 2019 | Learning an Event Sequence Embedding for Dense Event-Based Deep Stereo · ICCV 2019 |
Computer vision › 3D vision › stereo vision › stereo matching
deep stereo matching |
0.3 | 1 | 2018 | Practical Deep Stereo (PDS): Toward applications-friendly deep stereo matching · NeurIPS 2018 |
Machine learning › Efficient and distributed learning › inference efficiency
memory-efficient inference |
0.3 | 1 | 2018 | Practical Deep Stereo (PDS): Toward applications-friendly deep stereo matching · NeurIPS 2018 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2018 | Practical Deep Stereo (PDS): Toward applications-friendly deep stereo matching · NeurIPS 2018 |
Computational photography and imaging
event camera |
0.3 | 2 | 2022 | Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale Fusion · CVPR 2022 Time Lens: Event-Based Video Frame Interpolation · CVPR 2021 |
Machine learning › Representation and self-supervised learning › representation learning › metric learning
deep metric learning |
0.3 | 1 | 2017 | Weakly Supervised Learning of Deep Metrics for Stereo Reconstruction · ICCV 2017 |
Machine learning › Learning paradigms
weakly supervised learning |
0.3 | 1 | 2017 | Weakly Supervised Learning of Deep Metrics for Stereo Reconstruction · ICCV 2017 |
Computational photography and imaging
event-based vision |
0.2 | 1 | 2023 | EvShutter: Transforming Events for Unconstrained Rolling Shutter Correction · CVPR 2023 |
Computer vision › 3D vision › feature matching › local feature matching
patch matching |
0.1 | 1 | 2017 | Weakly Supervised Learning of Deep Metrics for Stereo Reconstruction · ICCV 2017 |
Methods — techniques the papers use, named apart from their topics
deep learning · 0.7filter and flip · 0.7double encoder hourglass network · 0.7adaptive interpolation simulator · 0.7nonlinear motion estimation · 0.6multi-scale feature fusion · 0.6synthesis-based interpolation · 0.5flow-based interpolation · 0.5event-based sensor · 0.4sub-pixel cross-entropy loss · 0.3MAP estimation · 0.3stochastic gradient descent · 0.3stereo constraints · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | EvShutter: Transforming Events for Unconstrained Rolling Shutter CorrectionabstractWidely used Rolling Shutter (RS) CMOS sensors capture high resolution images at the expense of introducing distortions and artifacts in the presence of motion. In such situations, RS distortion correction algorithms are critical. Recent methods rely on a constant velocity assumption and require multiple frames to predict the dense displacement field. In this work, we introduce a new method, called Eventful Shutter (EvShutter)11The evaluation code and the dataset can be found here https://github.com/juliuserbach/EvShutter, that corrects RS artifacts using a single RGB image and event information with high temporal resolution. The method firstly removes blur using a novel flow-based deblurring module and then compensates RS using a double encoder hourglass network. In contrast to previous methods, it does not rely on a constant velocity assumption and uses a simple architecture thanks to an event transformation dedicated to RS, called Filter and Flip (FnF), that transforms input events to encode only the changes between GS and RS images. To evaluate the proposed method and facilitate future research, we collect the first dataset with real events and high-quality RS images with optional blur, called RS-ERGB. We generate the RS images from GS images using a newly proposed simulator based on adaptive interpolation. The simulator permits the use of inexpensive cameras with long exposure to capture high-quality GS images. We show that on this realistic dataset the proposed method outperforms the state-of-the-art image-and event-based methods by 9.16 dB and 0.75 dB respectively in terms of PSNR and an improvement of 23 % and 21 % in LPIPS. Julius Erbach, Stepan Tulyakov, Patricia Vitoria, Alfredo Bochicchio, Yuanyou Li |
CVPR | 2 |
| 2022 | Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale FusionabstractRecently, video frame interpolation using a combination of frame- and event-based cameras has surpassed traditional image-based methods both in terms of performance and memory efficiency. However, current methods still suffer from (i) brittle image-level fusion of complementary interpolation results, that fails in the presence of artifacts in the fused image, (ii) potentially temporally inconsistent and inefficient motion estimation procedures, that run for every inserted frame and (iii) low contrast regions that do not trigger events, and thus cause events-only motion estimation to generate artifacts. Moreover, previous methods were only tested on datasets consisting of planar and far-away scenes, which do not capture the full complexity of the real world. In this work, we address the above problems by introducing multi-scale feature-level fusion and computing one-shot non-linear inter-frame motion-which can be efficiently sampled for image warping-from events and images. We also collect the first large-scale events and frames dataset consisting of more than 100 challenging scenes with depth variations, captured with a new experimental setup based on a beamsplitter. We show that our method improves the reconstruction quality by up to 0.2 dB in terms of PSNR and up to 15% in LPIPS score. Stepan Tulyakov, Alfredo Bochicchio, Daniel Gehrig, Stamatios Georgoulis, Yuanyou Li, Davide Scaramuzza 0001 |
CVPR | 1 |
| 2021 | Time Lens: Event-Based Video Frame InterpolationabstractState-of-the-art frame interpolation methods generate intermediate frames by inferring object motions in the image from consecutive key-frames. In the absence of additional information, first-order approximations, i.e. optical flow, must be used, but this choice restricts the types of motions that can be modeled, leading to errors in highly dynamic scenarios. Event cameras are novel sensors that address this limitation by providing auxiliary visual information in the blind-time between frames. They asynchronously measure per-pixel brightness changes and do this with high temporal resolution and low latency. Event-based frame interpolation methods typically adopt a synthesis-based approach, where predicted frame residuals are directly applied to the key-frames. However, while these approaches can capture non-linear motions they suffer from ghosting and perform poorly in low-texture regions with few events. Thus, synthesis-based and flow-based approaches are complementary. In this work, we introduce Time Lens, a novel method that leverages the advantages of both. We extensively evaluate our method on three synthetic and two real benchmarks where we show an up to 5.21 dB improvement in terms of PSNR over state-of-the-art frame-based and event-based methods. Finally, we release a new large-scale dataset in highly dynamic scenarios, aimed at pushing the limits of existing methods. Stepan Tulyakov, Daniel Gehrig, Stamatios Georgoulis, Julius Erbach, Mathias Gehrig, Yuanyou Li, Davide Scaramuzza 0001 |
CVPR | 1 |
| 2019 | Learning an Event Sequence Embedding for Dense Event-Based Deep StereoabstractToday, a frame-based camera is the sensor of choice for machine vision applications. However, these cameras, originally developed for acquisition of static images rather than for sensing of dynamic uncontrolled visual environments, suffer from high power consumption, data rate, latency and low dynamic range. An event-based image sensor addresses these drawbacks by mimicking a biological retina. Instead of measuring the intensity of every pixel in a fixed time-interval, it reports events of significant pixel intensity changes. Every such event is represented by its position, sign of change, and timestamp, accurate to the microsecond. Asynchronous event sequences require special handling, since traditional algorithms work only with synchronous, spatially gridded data. To address this problem we introduce a new module for event sequence embedding, for use in difference applications. The module builds a representation of an event sequence by firstly aggregating information locally across time, using a novel fully-connected layer for an irregularly sampled continuous domain, and then across discrete spatial domain. Based on this module, we design a deep learning-based stereo method for event-based cameras. The proposed method is the first learning-based stereo method for an event-based camera and the only method that produces dense results. We show that large performance increases on the Multi Vehicle Stereo Event Camera Dataset (MVSEC), which became the standard set for benchmarking of event-based stereo methods. Stepan Tulyakov, François Fleuret, Martin Kiefel, Peter V. Gehler |
ICCV | 1 |
| 2018 | Practical Deep Stereo (PDS): Toward applications-friendly deep stereo matchingabstractEnd-to-end deep-learning networks recently demonstrated extremely good performance for stereo matching. However, existing networks are difficult to use for practical applications since (1) they are memory-hungry and unable to process even modest-size images, (2) they have to be fully re-trained to handle a different disparity range. The Practical Deep Stereo (PDS) network that we propose addresses both issues: First, its architecture relies on novel bottleneck modules that drastically reduce the memory footprint in inference, and additional design choices allow to handle greater image size during training. This results in a model that leverages large image context to resolve matching ambiguities. Second, a novel sub-pixel cross-entropy loss combined with a MAP estimator make this network less sensitive to ambiguous matches, and applicable to any disparity range without re-training. We compare PDS to state-of-the-art methods published over the recent months, and demonstrate its superior performance on FlyingThings3D and KITTI sets. Stepan Tulyakov, Anton Ivanov, François Fleuret |
NeurIPS | 1 |
| 2017 | Weakly Supervised Learning of Deep Metrics for Stereo ReconstructionabstractDeep-learning metrics have recently demonstrated extremely good performance to match image patches for stereo reconstruction. However, training such metrics requires large amount of labeled stereo images, which can be difficult or costly to collect for certain applications (consider, for example, satellite stereo imaging). The main contribution of our work is a new weakly supervised method for learning deep metrics from unlabeled stereo images, given coarse information about the scenes and the optical system. Our method alternatively optimizes the metric with a standard stochastic gradient descent, and applies stereo constraints to regularize its prediction. Experiments on reference data-sets show that, for a given network architecture, training with this new method without ground-truth produces a metric with performance as good as state-of-the-art baselines trained with the said ground-truth. This work has three practical implications. Firstly, it helps to overcome limitations of training sets, in particular noisy ground truth. Secondly it allows to use much more training data during learning. Thirdly, it allows to tune deep metric for a particular stereo system, even if ground truth is not available. Stepan Tulyakov, Anton Ivanov, François Fleuret |
ICCV | 1 |
| 2013 | Quadratic formulation of disparity estimation problem for light-field cameraabstractNewly available light-field (LF) cameras are able to capture several views of a scene simultaneously. These views typically have small parallax, and thus can be easily registered. In this paper we exploit this property of the views captured by the LF camera to formulate disparity estimation problem as a quadratic energy minimization problem. Our problem formulation has three benefits. Firstly, it allows computation of continuous disparity with subpixel accuracy. Secondly it permits recovering disparity of loosely textured objects and ensures that the disparity boundaries are aligned with the object's boundaries. And, finally, it allows finding the solution very quickly. It takes 15-20s for our non-optimized Matlab code to compute the solution for 25 × 350 × 350 input views. Stepan Tulyakov, Tae Hee Lee, Heechul Han |
ICIP | 1 |