Stepan Tulyakov

dblp:141/9935 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Image and video processing · 81% Computational photography and imaging · 19%
Artificial intelligence
3 papers
3D vision · 52% Representation and self-supervised learning · 18% Efficient and distributed learning · 18%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › video frame interpolation
event-based frame interpolation
1.122022
Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale Fusion · CVPR 2022
Time Lens: Event-Based Video Frame Interpolation · CVPR 2021
Image and video processing
motion estimation
1.122022
Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale Fusion · CVPR 2022
Time Lens: Event-Based Video Frame Interpolation · CVPR 2021
Image and video processing › motion estimation
optical flow
1.122022
Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale Fusion · CVPR 2022
Time Lens: Event-Based Video Frame Interpolation · CVPR 2021
Image and video processing
video frame interpolation
1.122022
Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale Fusion · CVPR 2022
Time Lens: Event-Based Video Frame Interpolation · CVPR 2021
Computer vision › 3D vision › 3d reconstruction › multi-view stereo
stereo reconstruction
0.722019
Learning an Event Sequence Embedding for Dense Event-Based Deep Stereo · ICCV 2019
Weakly Supervised Learning of Deep Metrics for Stereo Reconstruction · ICCV 2017
Image and video processing › image restoration
image deblurring
0.712023
EvShutter: Transforming Events for Unconstrained Rolling Shutter Correction · CVPR 2023
Computational photography and imaging › image signal processing
rolling shutter correction
0.712023
EvShutter: Transforming Events for Unconstrained Rolling Shutter Correction · CVPR 2023
Computer vision › 3D vision › stereo vision
stereo matching
0.422018
Practical Deep Stereo (PDS): Toward applications-friendly deep stereo matching · NeurIPS 2018
Weakly Supervised Learning of Deep Metrics for Stereo Reconstruction · ICCV 2017
Computer vision › 3D vision › stereo vision
event-based stereo
0.412019
Learning an Event Sequence Embedding for Dense Event-Based Deep Stereo · ICCV 2019
Computer vision › 3D vision › stereo vision › stereo matching
deep stereo matching
0.312018
Practical Deep Stereo (PDS): Toward applications-friendly deep stereo matching · NeurIPS 2018
Machine learning › Efficient and distributed learning › inference efficiency
memory-efficient inference
0.312018
Practical Deep Stereo (PDS): Toward applications-friendly deep stereo matching · NeurIPS 2018
Machine learning › Efficient and distributed learning
model compression
0.312018
Practical Deep Stereo (PDS): Toward applications-friendly deep stereo matching · NeurIPS 2018
Computational photography and imaging
event camera
0.322022
Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale Fusion · CVPR 2022
Time Lens: Event-Based Video Frame Interpolation · CVPR 2021
Machine learning › Representation and self-supervised learning › representation learning › metric learning
deep metric learning
0.312017
Weakly Supervised Learning of Deep Metrics for Stereo Reconstruction · ICCV 2017
Machine learning › Learning paradigms
weakly supervised learning
0.312017
Weakly Supervised Learning of Deep Metrics for Stereo Reconstruction · ICCV 2017
Computational photography and imaging
event-based vision
0.212023
EvShutter: Transforming Events for Unconstrained Rolling Shutter Correction · CVPR 2023
Computer vision › 3D vision › feature matching › local feature matching
patch matching
0.112017
Weakly Supervised Learning of Deep Metrics for Stereo Reconstruction · ICCV 2017

Methods — techniques the papers use, named apart from their topics

deep learning · 0.7filter and flip · 0.7double encoder hourglass network · 0.7adaptive interpolation simulator · 0.7nonlinear motion estimation · 0.6multi-scale feature fusion · 0.6synthesis-based interpolation · 0.5flow-based interpolation · 0.5event-based sensor · 0.4sub-pixel cross-entropy loss · 0.3MAP estimation · 0.3stochastic gradient descent · 0.3stereo constraints · 0.3
YearPublicationVenuePosition
2023 EvShutter: Transforming Events for Unconstrained Rolling Shutter Correction
abstract
Widely used Rolling Shutter (RS) CMOS sensors capture high resolution images at the expense of introducing distortions and artifacts in the presence of motion. In such situations, RS distortion correction algorithms are critical. Recent methods rely on a constant velocity assumption and require multiple frames to predict the dense displacement field. In this work, we introduce a new method, called Eventful Shutter (EvShutter)11The evaluation code and the dataset can be found here https://github.com/juliuserbach/EvShutter, that corrects RS artifacts using a single RGB image and event information with high temporal resolution. The method firstly removes blur using a novel flow-based deblurring module and then compensates RS using a double encoder hourglass network. In contrast to previous methods, it does not rely on a constant velocity assumption and uses a simple architecture thanks to an event transformation dedicated to RS, called Filter and Flip (FnF), that transforms input events to encode only the changes between GS and RS images. To evaluate the proposed method and facilitate future research, we collect the first dataset with real events and high-quality RS images with optional blur, called RS-ERGB. We generate the RS images from GS images using a newly proposed simulator based on adaptive interpolation. The simulator permits the use of inexpensive cameras with long exposure to capture high-quality GS images. We show that on this realistic dataset the proposed method outperforms the state-of-the-art image-and event-based methods by 9.16 dB and 0.75 dB respectively in terms of PSNR and an improvement of 23 % and 21 % in LPIPS.
Julius Erbach, Stepan Tulyakov, Patricia Vitoria, Alfredo Bochicchio, Yuanyou Li
CVPR2
2022 Time Lens++: Event-based Frame Interpolation with Parametric Nonlinear Flow and Multi-scale Fusion
abstract
Recently, video frame interpolation using a combination of frame- and event-based cameras has surpassed traditional image-based methods both in terms of performance and memory efficiency. However, current methods still suffer from (i) brittle image-level fusion of complementary interpolation results, that fails in the presence of artifacts in the fused image, (ii) potentially temporally inconsistent and inefficient motion estimation procedures, that run for every inserted frame and (iii) low contrast regions that do not trigger events, and thus cause events-only motion estimation to generate artifacts. Moreover, previous methods were only tested on datasets consisting of planar and far-away scenes, which do not capture the full complexity of the real world. In this work, we address the above problems by introducing multi-scale feature-level fusion and computing one-shot non-linear inter-frame motion-which can be efficiently sampled for image warping-from events and images. We also collect the first large-scale events and frames dataset consisting of more than 100 challenging scenes with depth variations, captured with a new experimental setup based on a beamsplitter. We show that our method improves the reconstruction quality by up to 0.2 dB in terms of PSNR and up to 15% in LPIPS score.
Stepan Tulyakov, Alfredo Bochicchio, Daniel Gehrig, Stamatios Georgoulis, Yuanyou Li, Davide Scaramuzza 0001
CVPR1
2021 Time Lens: Event-Based Video Frame Interpolation
abstract
State-of-the-art frame interpolation methods generate intermediate frames by inferring object motions in the image from consecutive key-frames. In the absence of additional information, first-order approximations, i.e. optical flow, must be used, but this choice restricts the types of motions that can be modeled, leading to errors in highly dynamic scenarios. Event cameras are novel sensors that address this limitation by providing auxiliary visual information in the blind-time between frames. They asynchronously measure per-pixel brightness changes and do this with high temporal resolution and low latency. Event-based frame interpolation methods typically adopt a synthesis-based approach, where predicted frame residuals are directly applied to the key-frames. However, while these approaches can capture non-linear motions they suffer from ghosting and perform poorly in low-texture regions with few events. Thus, synthesis-based and flow-based approaches are complementary. In this work, we introduce Time Lens, a novel method that leverages the advantages of both. We extensively evaluate our method on three synthetic and two real benchmarks where we show an up to 5.21 dB improvement in terms of PSNR over state-of-the-art frame-based and event-based methods. Finally, we release a new large-scale dataset in highly dynamic scenarios, aimed at pushing the limits of existing methods.
Stepan Tulyakov, Daniel Gehrig, Stamatios Georgoulis, Julius Erbach, Mathias Gehrig, Yuanyou Li, Davide Scaramuzza 0001
CVPR1
2019 Learning an Event Sequence Embedding for Dense Event-Based Deep Stereo
abstract
Today, a frame-based camera is the sensor of choice for machine vision applications. However, these cameras, originally developed for acquisition of static images rather than for sensing of dynamic uncontrolled visual environments, suffer from high power consumption, data rate, latency and low dynamic range. An event-based image sensor addresses these drawbacks by mimicking a biological retina. Instead of measuring the intensity of every pixel in a fixed time-interval, it reports events of significant pixel intensity changes. Every such event is represented by its position, sign of change, and timestamp, accurate to the microsecond. Asynchronous event sequences require special handling, since traditional algorithms work only with synchronous, spatially gridded data. To address this problem we introduce a new module for event sequence embedding, for use in difference applications. The module builds a representation of an event sequence by firstly aggregating information locally across time, using a novel fully-connected layer for an irregularly sampled continuous domain, and then across discrete spatial domain. Based on this module, we design a deep learning-based stereo method for event-based cameras. The proposed method is the first learning-based stereo method for an event-based camera and the only method that produces dense results. We show that large performance increases on the Multi Vehicle Stereo Event Camera Dataset (MVSEC), which became the standard set for benchmarking of event-based stereo methods.
Stepan Tulyakov, François Fleuret, Martin Kiefel, Peter V. Gehler
ICCV1
2018 Practical Deep Stereo (PDS): Toward applications-friendly deep stereo matching
abstract
End-to-end deep-learning networks recently demonstrated extremely good performance for stereo matching. However, existing networks are difficult to use for practical applications since (1) they are memory-hungry and unable to process even modest-size images, (2) they have to be fully re-trained to handle a different disparity range. The Practical Deep Stereo (PDS) network that we propose addresses both issues: First, its architecture relies on novel bottleneck modules that drastically reduce the memory footprint in inference, and additional design choices allow to handle greater image size during training. This results in a model that leverages large image context to resolve matching ambiguities. Second, a novel sub-pixel cross-entropy loss combined with a MAP estimator make this network less sensitive to ambiguous matches, and applicable to any disparity range without re-training. We compare PDS to state-of-the-art methods published over the recent months, and demonstrate its superior performance on FlyingThings3D and KITTI sets.
Stepan Tulyakov, Anton Ivanov, François Fleuret
NeurIPS1
2017 Weakly Supervised Learning of Deep Metrics for Stereo Reconstruction
abstract
Deep-learning metrics have recently demonstrated extremely good performance to match image patches for stereo reconstruction. However, training such metrics requires large amount of labeled stereo images, which can be difficult or costly to collect for certain applications (consider, for example, satellite stereo imaging). The main contribution of our work is a new weakly supervised method for learning deep metrics from unlabeled stereo images, given coarse information about the scenes and the optical system. Our method alternatively optimizes the metric with a standard stochastic gradient descent, and applies stereo constraints to regularize its prediction. Experiments on reference data-sets show that, for a given network architecture, training with this new method without ground-truth produces a metric with performance as good as state-of-the-art baselines trained with the said ground-truth. This work has three practical implications. Firstly, it helps to overcome limitations of training sets, in particular noisy ground truth. Secondly it allows to use much more training data during learning. Thirdly, it allows to tune deep metric for a particular stereo system, even if ground truth is not available.
Stepan Tulyakov, Anton Ivanov, François Fleuret
ICCV1
2013 Quadratic formulation of disparity estimation problem for light-field camera
abstract
Newly available light-field (LF) cameras are able to capture several views of a scene simultaneously. These views typically have small parallax, and thus can be easily registered. In this paper we exploit this property of the views captured by the LF camera to formulate disparity estimation problem as a quadratic energy minimization problem. Our problem formulation has three benefits. Firstly, it allows computation of continuous disparity with subpixel accuracy. Secondly it permits recovering disparity of loosely textured objects and ensures that the disparity boundaries are aligned with the object's boundaries. And, finally, it allows finding the solution very quickly. It takes 15-20s for our non-optimized Matlab code to compute the solution for 25 × 350 × 350 input views.
Stepan Tulyakov, Tae Hee Lee, Heechul Han
ICIP1