Dhaval Salvi

dblp:65/8652 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-authorArtificial intelligence and machine learning · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Image recognition and object detection · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object localization
0.112010
Free-shape subwindow search for object localization · CVPR 2010
Computer vision › Image recognition and object detection › object detection
subwindow search
0.112010
Free-shape subwindow search for object localization · CVPR 2010

Methods — techniques the papers use, named apart from their topics

ratio-contour graph algorithm · 0.1bag-of-visual-words · 0.1
YearPublicationVenuePosition
2015 Distance Transform Based Active Contour Approach for Document Image Rectification
abstract
Digitization of document images using OCR based systems is adversely affected if the image of the document contains distortion (warping). Often, costly and precisely calibrated special hardware such as stereo cameras, laser scanners, etc. are used to infer the 3D model of the distorted image which is used to remove the distortion. Recent methods focus on creating a 3D shape model based on the 2D document image. The performance of these methods is highly dependent on estimating an accurate 2D distortion grid. In the domain of printed document images, the white space between the text lines carries as much information about the 2D distortion as the text lines themselves. Based on this intuitive idea, we build a 2D distortion grid from white space lines, which can be used to rectify a printed document image by a Dewar ping algorithm. These white space lines are extracted using a propagation technique on the distance transform of the binarized document image, guided by an open active contour algorithm. We compare our proposed method against a state-of-the-art 2D distortion grid construction method and obtain better results. We also present qualitative and quantitative evaluations for the proposed method.
Dhaval Salvi, Youjie Zhou, Song Wang 0002
WACV1
2013 A graph-based algorithm for multi-target tracking with occlusion
abstract
Multi-target tracking plays a key role in many computer vision applications including robotics, human-computer interaction, event recognition, etc., and has received increasing attention in past several years. Starting with an object detector is one of many approaches used by existing multi-target tracking methods to create initial short tracks called tracklets. These tracklets are then gradually grouped into longer final tracks in a heirarchical framework. Although object detectors have greatly improved in recent years, these detectors are far from perfect and can fail to detect the object of interest or identify a false positive as the desired object. Due to the presence of false positives or mis-detections from the object detector, these tracking methods can suffer from track fragmentations and identity switches. To address this problem, we formulate multi-target tracking as a min-cost flow graph problem which we call the average shortest path. This average shortest path is designed to be less biased towards the track length. In our average shortest path framework, object misdetection is treated as an occlusion and is represented by the edges between track-let nodes across non consecutive frames. We evaluate our method on the publicly available ETH dataset. Camera motion and long occlusions in a busy street scene make ETH a challenging dataset. We achieve competitive results with lower identity switches on this dataset as compared to the state of the art methods.
Dhaval Salvi, Jarrell W. Waggoner, Andrew Temlyakov, Song Wang 0002
WACV1
2013 Handwritten text segmentation using average longest path algorithm
abstract
Offline handwritten text recognition is a very challenging problem. Aside from the large variation of different handwriting styles, neighboring characters within a word are usually connected, and we may need to segment a word into individual characters for accurate character recognition. Many existing methods achieve text segmentation by evaluating the local stroke geometry and imposing constraints on the size of each resulting character, such as the character width, height and aspect ratio. These constraints are well suited for printed texts, but may not hold for handwritten texts. Other methods apply holistic approach by using a set of lexicons to guide and correct the segmentation and recognition. This approach may fail when the lexicon domain is insufficient. In this paper, we present a new global non-holistic method for handwritten text segmentation, which does not make any limiting assumptions on the character size and the number of characters in a word. Specifically, the proposed method finds the text segmentation with the maximum average likeliness for the resulting characters. For this purpose, we use a graph model that describes the possible locations for segmenting neighboring characters, and we then develop an average longest path algorithm to identify the globally optimal segmentation. We conduct experiments on real images of handwritten texts taken from the IAM handwriting database and compare the performance of the proposed method against an existing text segmentation algorithm that uses dynamic programming.
Dhaval Salvi, Jarrell W. Waggoner, Song Wang 0002
WACV1
2013 Shape and image retrieval by organizing instances using population cues
abstract
Reliably measuring the similarity of two shapes or images (instances) is an important problem for various computer vision applications such as classification, recognition, and retrieval. While pairwise measures take advantage of the geometric differences between two instances to quantify their similarity, recent advances use relationships among the population of instances when quantifying pairwise measures. In this paper, we propose a novel method which refines pairwise similarity measures using population cues by examining the most similar instances shared by the compared shapes or images. We then use this refined measure to organize instances into disjoint components that consist of similar instances. Connectivity is then established between components to avoid hard constraints on what instances can be retrieved, improving retrieval performance. To evaluate the proposed method we conduct experiments on the well-known MPEG-7 and Swedish Leaf shape datasets as well as the Nister and Stewenius image dataset. We show that the proposed method is versatile, performing very well on its own or in concert with existing methods.
Andrew Temlyakov, Pahal Dalal, Jarrell W. Waggoner, Dhaval Salvi, Song Wang 0002
WACV4
2012 Video In Sentences Out
Andrei Barbu, Alexander Bridge, Zachary Burchill, Dan Coroian, Sven J. Dickinson, Sanja Fidler, Aaron Michaux, Sam Mussman, N. Siddharth 0001, Dhaval Salvi, Lara Schmidt, Jiangnan Shangguan, Jeffrey Mark Siskind, Jarrell W. Waggoner, Song Wang 0002, Jinlian Wei
UAI10
2010 Free-shape subwindow search for object localization
abstract
Object localization in an image is usually handled by searching for an optimal subwindow that tightly covers the object of interest. However, the subwindows considered in previous work are limited to rectangles or other specified, simple shapes. With such specified shapes, no subwindow can cover the object of interest tightly. As a result, the desired subwindow around the object of interest may not be optimal in terms of the localization objective function, and cannot be detected by a subwindow search algorithm. In this paper, we propose a new graph-theoretic approach for object localization by searching for an optimal subwindow without pre-specifying its shape. Instead, we require the resulting subwindow to be well aligned with edge pixels that are detected from the image. This requirement is quantified and integrated into the localization objective function based on the widely-used bag of visual words technique. We show that the ratio-contour graph algorithm can be adapted to find the optimal free-shape subwindow in terms of the new localization objective function. In the experiment, we test the proposed approach on the PASCAL VOC 2006 and VOC 2007 databases for localizing several categories of animals. We find that its performance is better than the previous efficient subwindow search algorithm.
Yu Cao 0003, Dhaval Salvi, Kenton Oliver, Jarrell W. Waggoner, Song Wang 0002
CVPR3