Dmitry Rudoy

dblp:53/5166 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 2017
0000-0002-1916-8570ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Video understanding and tracking · 45% Deep learning architectures and training · 20% Generative modeling · 20%
Computer graphics and multimedia
1 paper
Virtual and augmented reality · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
video saliency prediction
0.522017
Learning Gaze Transitions from Depth to Improve Video Saliency Estimation · ICCV 2017
Learning Video Saliency from Human Gaze Using Candidate Selection · CVPR 2013
Machine learning › Deep learning architectures and training
convolutional neural network
0.312017
Learning Gaze Transitions from Depth to Improve Video Saliency Estimation · ICCV 2017
Machine learning › Generative modeling › energy-based model
generative convnet
0.312017
Learning Gaze Transitions from Depth to Improve Video Saliency Estimation · ICCV 2017
Computer vision › Video understanding and tracking
action recognition
0.112012
Viewpoint Selection for Human Actions · Int. J. Comput. Vis. 2012
Robotics › Robot navigation and mapping › view planning
next-best-view planning
0.112012
Viewpoint Selection for Human Actions · Int. J. Comput. Vis. 2012
Machine learning › Efficient and distributed learning
candidate selection
0.012013
Learning Video Saliency from Human Gaze Using Candidate Selection · CVPR 2013
Computer vision › Video understanding and tracking
human action analysis
0.012012
Viewpoint Selection for Human Actions · Int. J. Comput. Vis. 2012

Methods — techniques the papers use, named apart from their topics

generative convolutional neural network · 0.6fixation prediction · 0.6gaze tracking · 0.2conditional saliency prediction · 0.2viewpoint entropy · 0.1discriminative model · 0.1
YearPublicationVenuePosition
2017 Learning Gaze Transitions from Depth to Improve Video Saliency Estimation
abstract
In this paper we introduce a novel Depth-Aware Video Saliency approach to predict human focus of attention when viewing videos that contain a depth map (RGBD) on a 2D screen. Saliency estimation in this scenario is highly important since in the near future 3D video content will be easily acquired yet hard to display. Despite considerable progress in 3D display technologies, most are still expensive and require special glasses for viewing, so RGBD content is primarily viewed on 2D screens, removing the depth channel from the final viewing experience. We train a generative convolutional neural network that predicts the 2D viewing saliency map for a given frame using the RGBD pixel values and previous fixation estimates in the video. To evaluate the performance of our approach, we present a new comprehensive database of 2D viewing eye-fixation ground-truth for RGBD videos. Our experiments indicate that it is beneficial to integrate depth into video saliency estimates for content that is viewed on a 2D display. We demonstrate that our approach outperforms state-of-the-art methods for video saliency, achieving 15% relative improvement.
George Leifman, Dmitry Rudoy, Tristan Swedish, Eduardo Bayro-Corrochano, Ramesh Raskar
ICCV2
2013 Learning Video Saliency from Human Gaze Using Candidate Selection
abstract
During recent years remarkable progress has been made in visual saliency modeling. Our interest is in video saliency. Since videos are fundamentally different from still images, they are viewed differently by human observers. For example, the time each video frame is observed is a fraction of a second, while a still image can be viewed leisurely. Therefore, video saliency estimation methods should differ substantially from image saliency methods. In this paper we propose a novel method for video saliency estimation, which is inspired by the way people watch videos. We explicitly model the continuity of the video by predicting the saliency map of a given frame, conditioned on the map from the previous frame. Furthermore, accuracy and computation speed are improved by restricting the salient locations to a carefully selected candidate set. We validate our method using two gaze-tracked video datasets and show we outperform the state-of-the-art.
Dmitry Rudoy, Dan B. Goldman, Eli Shechtman, Lihi Zelnik-Manor
CVPR1
2013 Video inlays: a system for user-friendly matchmove
abstract
Digital editing technology is highly popular as it enables to easily change photos and add to them artificial objects. Conversely, video editing is still challenging and mainly left to the professionals. Even basic video manipulations involve complicated software tools that are typically not adopted by the amateur user. In this paper we propose a system that allows an amateur user to performs a basic matchmove by adding an inlay to a video. Our system does not require any previous experience and relies on a simple user interaction. We allow adding 3D objects and volumetric textures to virtually any video. We demonstrate the method's applicability on a variety of videos downloaded from the web.
Dmitry Rudoy, Lihi Zelnik-Manor
VRST1
2012 Viewpoint Selection for Human Actions
Dmitry Rudoy, Lihi Zelnik-Manor
Int. J. Comput. Vis.1
2010 Posing to the Camera: Automatic Viewpoint Selection for Human Actions
Dmitry Rudoy, Lihi Zelnik-Manor
ACCV (4)1