Sarthak Upadhyay

dblp:162/6178 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
3D vision · 44% Segmentation and scene understanding · 28% Robot navigation and mapping · 22%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
monocular 3d object detection
0.212016
Monocular reconstruction of vehicles: Combining SLAM with shape priors · ICRA 2016
Computer vision › Segmentation and scene understanding
shape prior
0.212016
Monocular reconstruction of vehicles: Combining SLAM with shape priors · ICRA 2016
Robotics › Robot navigation and mapping
SLAM
0.212016
Monocular reconstruction of vehicles: Combining SLAM with shape priors · ICRA 2016
Computer vision › 3D vision
structure from motion
0.212016
Monocular reconstruction of vehicles: Combining SLAM with shape priors · ICRA 2016
Computer vision › Segmentation and scene understanding
scene understanding
0.112016
Monocular reconstruction of vehicles: Combining SLAM with shape priors · ICRA 2016
Robotics › Autonomous driving › perception
vehicle detection and tracking
0.112016
Monocular reconstruction of vehicles: Combining SLAM with shape priors · ICRA 2016

Methods — techniques the papers use, named apart from their topics

optimization · 0.2multi-view detection · 0.23d model fitting · 0.2
YearPublicationVenuePosition
2016 Monocular reconstruction of vehicles: Combining SLAM with shape priors
abstract
Reasoning about objects in images and videos using 3D representations is re-emerging as a popular paradigm in computer vision. Specifically, in the context of scene understanding for roads, 3D vehicle detection and tracking from monocular videos still needs a lot of attention to enable practical applications. Current approaches leverage two kinds of information to deal with the vehicle detection and tracking problem: (1) 3D representations (eg. wireframe models or voxel based or CAD models) for diverse vehicle skeletal structures learnt from data, and (2) classifiers trained to detect vehicles or vehicle parts in single images built on top of a basic feature extraction step. In this paper, we propose to extend current approaches in two ways. First, we extend detection to a multiple view setting. We show that leveraging information given by feature or part detectors in multiple images can lead to more accurate detection results than single image detection. Secondly, we show that given multiple images of a vehicle, we can also leverage 3D information from the scene generated using a unique structure from motion algorithm. This helps us localize the vehicle in 3D, and constrain the parameters of optimization for fitting the 3D model to image data. We show results on the KITTI dataset, and demonstrate superior results compared with recent state-of-the-art methods, with upto 14.64 % improvement in localization error.
Falak Chhaya, N. Dinesh Reddy, Sarthak Upadhyay, Visesh Chari, M. Zeeshan Zia, K. Madhava Krishna
ICRA3