Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ankur Datta

dblp:90/7 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-authorArtificial intelligence and machine learning · 10 · 4 first-authorDatabases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Image recognition and object detection · 44% 3D vision · 41% Video understanding and tracking · 16%
Computer graphics and multimedia
3 papers
Multimedia analysis and retrieval · 73% Rendering · 27%

Topics — the 13 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › motion estimation › non-rigid motion estimation
articulated motion estimation
0.222011
Linearized Motion Estimation for Articulated Planes · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Linear motion estimation for systems of articulated planes · CVPR 2008
Computer vision › 3D vision
motion estimation
0.222011
Linearized Motion Estimation for Articulated Planes · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Linear motion estimation for systems of articulated planes · CVPR 2008
Computer vision › Image recognition and object detection › efficient visual recognition
efficient detection
0.212013
Efficient Maximum Appearance Search for Large-Scale Object Detection · CVPR 2013
Computer vision › Image recognition and object detection › object detection
large-scale object detection
0.212013
Efficient Maximum Appearance Search for Large-Scale Object Detection · CVPR 2013
Computer vision › Image recognition and object detection
object detection
0.212013
Efficient Maximum Appearance Search for Large-Scale Object Detection · CVPR 2013
Multimedia analysis and retrieval
video indexing
0.112012
Large-Scale Vehicle Detection, Indexing, and Search in Urban Surveillance Videos · IEEE Trans. Multim. 2012
Multimedia analysis and retrieval
video surveillance
0.112012
Large-Scale Vehicle Detection, Indexing, and Search in Urban Surveillance Videos · IEEE Trans. Multim. 2012
Rendering
novel view synthesis
0.112009
Dynamic seethroughs: Synthesizing hidden views of moving objects · ISMAR 2009
Computer vision › 3D vision › motion estimation › rigid motion estimation
planar motion estimation
0.112008
Linear motion estimation for systems of articulated planes · CVPR 2008
Computer vision › Image recognition and object detection › object detection › category-specific object detection
vehicle detection
0.012012
Large-Scale Vehicle Detection, Indexing, and Search in Urban Surveillance Videos · IEEE Trans. Multim. 2012
Information retrieval › search engines › semantic search › entity retrieval
attribute-based retrieval
0.012012
Large-Scale Vehicle Detection, Indexing, and Search in Urban Surveillance Videos · IEEE Trans. Multim. 2012
Rendering
image-based rendering
0.012009
Dynamic seethroughs: Synthesizing hidden views of moving objects · ISMAR 2009
Computer vision › Video understanding and tracking › object tracking › human motion tracking
human body tracking
0.012008
Linear motion estimation for systems of articulated planes · CVPR 2008

Methods — techniques the papers use, named apart from their topics

fisher vector encoding · 0.5occlusion modeling · 0.4feature selection · 0.4linear SVM · 0.3gaussian mixture model · 0.3MoSIFT · 0.3discriminative scoring · 0.2karush-kuhn-tucker system · 0.1gradient-based estimation · 0.1feature-based estimation · 0.1piece-wise planar model · 0.12d projective invariant · 0.1
YearPublicationVenuePosition
2018 E-commerce Product Query Classification Using Implicit User's Feedback from Clicks
abstract
Query classification (QC) has been widely studied to understand users' search intent. For e-commerce search queries, users typically search for either a specific product or a category of products. In both cases, a query can be associated with a category label that belongs to a taxonomy tree describing the items in the catalog. However, product-related search queries are typically short, ambiguous, and continuously changing depending on seasonal trends and the introduction of new products over time. Traditional supervised approaches to e-commerce QC are not feasible due to the high cost of manual annotation and the high volume of traffic on e-commerce search engines. In this work, we introduce an unsupervised method to collect large amounts of query classification data using user's implicit click feedback. We obtain a large multi-label dataset containing 403,349 unique queries from 2,085 categories. We compare and contrast different state-of-the-art text classifiers and demonstrate that an ensemble of linear SVMs models achieves a micro-F1 score of 0.60 and 0.82 at leaf and top level, respectively.
Yiu-Chang Lin, Ankur Datta, Giuseppe Di Fabbrizio
IEEE BigData2
2018 Dense Bynet: Residual Dense Network for Image Super Resolution
abstract
This paper proposes a method, Dense ByNet, for single image super-resolution based on a convolutional neural network (CNN). The main innovation is a new architecture that combines several CNN design choices. Using a residual network as a basis, it introduces dense connections inside residual blocks, significantly reducing the number of parameters. Second, we apply dilation convolutions to increase the spatial context. Lastly, we propose modifications to the activation and cost functions. We evaluate the method on benchmark datasets and show that it achieves state-of-the-art results over multiple upscaling factors in terms of peak SNR and structural similarity (SSIM).
Jiu Xu, Yeongnam Chae, Björn Stenger, Ankur Datta
ICIP4
2017 Web-Scale Language-Independent Cataloging of Noisy Product Listings for E-Commerce
abstract
Pradipto Das, Yandi Xia, Aaron Levine, Giuseppe Di Fabbrizio, Ankur Datta. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Pradipto Das, Yandi Xia, Aaron Levine, Giuseppe Di Fabbrizio, Ankur Datta
EACL (1)5
2016 Large-scale taxonomy categorization for noisy product listings
abstract
E-commerce catalogs include a continuously growing number of products that are constantly updated. Each item in a catalog is characterized by several attributes and identified by a taxonomy label. Categorizing products with their taxonomy labels is fundamental to effectively search and organize listings in a catalog. However, manual and/or rule based approaches to categorization are not scalable. In this paper, we compare several classifiers to product taxonomy categorization of top-level categories. We first investigate a number of feature sets and observe that a combination of word unigrams from product names and navigational breadcrumbs work best for categorization. Secondly, we apply correspondence topic models to detect noisy data and introduce a lightweight manual process to improve dataset quality. Finally, we evaluate linear models, gradient boosted trees (GBTs) and convolutional neural networks (CNNs) with pre-trained word embeddings demonstrating that, compared to other baselines, GBTs and CNNs yield the highest gains in error reduction.
Pradipto Das, Yandi Xia, Aaron Levine, Giuseppe Di Fabbrizio, Ankur Datta
IEEE BigData5
2013 Tree-based vehicle color classification using spatial features on publicly available continuous data
abstract
Several recent investigations attempt to classify vehicles into a small number (5-7) of colors. A significant complication arises, however; a large proportion of vehicles (>50%) are various shades of gray: white, black, silver, gray, and variations such as gun metal and pearly white. Distinguishing such shades of gray in vehicle body color from lighting changes is an unsolved problem. Furthermore, previous studies have evaluated their performance on private datasets precluding a comparison of methodologies. In this paper, we release a public dataset with ground truth color classification for future evaluations and comparisons based on the publicly available i-LIDS data [9]. We describe a method to perform vehicle color classification into 7 frequently occurring colors including dark red, dark blue and light silver, using pose dependent vehicle detection, vehicle alignment, and vehicle body part masks. We introduce new features for tree-based vehicle color classification based on the reliability of color information and the relative color of various vehicle parts.
Lisa M. Brown, Ankur Datta, Sharath Pankanti
AVSS2
2013 Efficient Maximum Appearance Search for Large-Scale Object Detection
abstract
In recent years, efficiency of large-scale object detection has arisen as an important topic due to the exponential growth in the size of benchmark object detection datasets. Most current object detection methods focus on improving accuracy of large-scale object detection with efficiency being an afterthought. In this paper, we present the Efficient Maximum Appearance Search (EMAS) model which is an order of magnitude faster than the existing state-of-the-art large-scale object detection approaches, while maintaining comparable accuracy. Our EMAS model consists of representing an image as an ensemble of densely sampled feature points with the proposed Point wise Fisher Vector encoding method, so that the learnt discriminative scoring function can be applied locally. Consequently, the object detection problem is transformed into searching an image sub-area for maximum local appearance probability, thereby making EMAS an order of magnitude faster than the traditional detection methods. In addition, the proposed model is also suitable for incorporating global context at a negligible extra computational cost. EMAS can also incorporate fusion of multiple features, which greatly improves its performance in detecting multiple object categories. Our experiments show that the proposed algorithm can perform detection of 1000 object classes in less than one minute per image on the Image Net ILSVRC2012 dataset and for 107 object classes in less than 5 seconds per image for the SUN09 dataset using a single CPU.
Qiang Chen 0007, Rogério Feris, Ankur Datta, Liangliang Cao, ZhongYang Huang, Shuicheng Yan
CVPR4
2013 Spatio-temporal fisher vector coding for surveillance event detection
abstract
We present a generic event detection system evaluated in the Surveillance Event Detection (SED) task of TRECVID 2012. We investigate a statistical approach with spatio-temporal features applied to seven event classes, which were defined by the SED task. This approach is based on local spatio-temporal descriptors, called MoSIFT and generated by pair-wise video frames. A Gaussian Mixture Model(GMM) is learned to model the distribution of the low level features. Then for each sliding window, the Fisher vector encoding [improvedFV] is used to generate the sample representation. The model is learnt using a Linear SVM for each event. The main novelty of our system is the introduction of Fisher vector encoding into video event detection. Fisher vector encoding has demonstrated great success in image classification. The key idea is to model the low level visual features as a Gaussian Mixture Model and to generate an intermediate vector representation for bag of features. FV encoding uses higher order statistics in place of histograms in the standard BoW. FV has several good properties: (a) it can naturally separate the video specific information from the noisy local features and (b) we can use a linear model for this representation. We build an efficient implementation for FV encoding which can attain a 10 times speed-up over real-time. We also take advantage of non-trivial object localization techniques to feed into the video event detection, e.g. multi-scale detection and non-maximum suppression. This approach outperformed the results of all other teams submissions in TRECVID SED 2012 on four of the seven event types.
Qiang Chen 0007, Yang Cai 0002, Lisa M. Brown, Ankur Datta, Quanfu Fan, Rogério Feris, Shuicheng Yan, Alex Hauptmann 0001, Sharath Pankanti
ACM Multimedia4
2013 Boosting object detection performance in crowded surveillance videos
abstract
We present a novel approach to automatically create efficient and accurate object detectors tailored to work well on specific video surveillance cameras (specific-domain detectors), using samples acquired with the help of a more expensive, general-domain detector (trained using images from multiple cameras). Our method requires no manual labels from the target domain. We automatically collect training data using tracking over short periods of time from high-confidence samples selected by the general-domain detector. In this context, a novel confidence measure is proposed for detectors based on a cascade of classifiers, which are frequently adopted for computer vision applications that require real-time processing. We demonstrate our proposed approach on the problem of vehicle detection in crowded surveillance videos, showing that an automatically generated detector significantly outperforms the original general-domain detector with much less feature computations.
Rogério Feris, Ankur Datta, Sharath Pankanti, Ming-Ting Sun
WACV2
2012 Appearance modeling for person re-identification using Weighted Brightness Transfer Functions
Ankur Datta, Lisa M. Brown, Rogério Feris, Sharath Pankanti
ICPR1
2012 Unsupervised model selection for view-invariant object detection in surveillance environments
Behjat Siddiquie, Rogério Feris, Ankur Datta, Larry Davis 0001
ICPR3
2012 Exploiting Color Strength to Improve Color Correction
abstract
Color information is an important feature for many vision algorithms including color correction, image retrieval and tracking. In this paper, we study the limitations of color measurement accuracy and explore how this information can be used to improve the performance of color correction. In particular, we show that a strong correlation exists between the error in hue measurements on one hand and saturation and intensity on the other hand. We introduce the notion of color strength, which is a combination of saturation and intensity information to determine when hue information in a scene is reliable. We verify the predictive capability of this model on two different datasets with ground truth color information. Further, we show how color strength information can be used to significantly improve color correction accuracy for the 11K real-world SFU gray ball dataset.
Lisa M. Brown, Ankur Datta, Sharath Pankanti
ISM2
2012 Large-Scale Vehicle Detection, Indexing, and Search in Urban Surveillance Videos
abstract
We present a novel approach for visual detection and attribute-based search of vehicles in crowded surveillance scenes. Large-scale processing is addressed along two dimensions: 1) large-scale indexing, where hundreds of billions of events need to be archived per month to enable effective search and 2) learning vehicle detectors with large-scale feature selection, using a feature pool containing millions of feature descriptors. Our method for vehicle detection also explicitly models occlusions and multiple vehicle types (e.g., buses, trucks, SUVs, cars), while requiring very few manual labeling. It runs quite efficiently at an average of 66 Hz on a conventional laptop computer. Once a vehicle is detected and tracked over the video, fine-grained attributes are extracted and ingested into a database to allow future search queries such as “Show me all blue trucks larger than 7 ft. length traveling at high speed northbound last Saturday, from 2 pm to 5 pm”. We perform a comprehensive quantitative analysis to validate our approach, showing its usefulness in realistic urban surveillance settings.
Rogério Feris, Behjat Siddiquie, James Petterson, Yun Zhai, Ankur Datta, Lisa M. Brown, Sharath Pankanti
IEEE Trans. Multim.5
2011 Hierarchical ranking of facial attributes
abstract
We propose a novel hierarchical structured prediction approach for ranking images of faces based on attributes. We view ranking as a bipartite graph matching problem; learning to rank under this setting can be achieved through structured prediction techniques that directly optimize the matching measures. Our key contribution is a novel model that combines structured predictors for different feature descriptors in a hierarchical fashion, enabling accurate ranking. We demonstrate our method on an important application which consists of searching for people over short intervals of time based on facial attributes. Given queries containing physical traits of a person (e.g., red hat, beard, and sunglasses), and an input database of face images, our system ranks the images in the database according to the query. Experiments show that our proposed hierarchical ranking approach poses significant enhancements in terms of accuracy over the non-hierarchical baseline.
Ankur Datta, Rogério Feris, Daniel A. Vaquero
FG1
2011 Linearized Motion Estimation for Articulated Planes
abstract
In this paper, we describe the explicit application of articulation constraints for estimating the motion of a system of articulated planes. We relate articulations to the relative homography between planes and show that these articulations translate into linearized equality constraints on a linear least-squares system, which can be solved efficiently using a Karush-Kuhn-Tucker system. The articulation constraints can be applied for both gradient-based and feature-based motion estimation algorithms and to illustrate this, we describe a gradient-based motion estimation algorithm for an affine camera and a feature-based motion estimation algorithm for a projective camera that explicitly enforces articulation constraints. We show that explicit application of articulation constraints leads to numerically stable estimates of motion. The simultaneous computation of motion estimates for all of the articulated planes in a scene allows us to handle scene areas where there is limited texture information and areas that leave the field of view. Our results demonstrate the wide applicability of the algorithm in a variety of challenging real-world cases such as human body tracking, motion estimation of rigid, piecewise planar scenes, and motion estimation of triangulated meshes.
Ankur Datta, Yaser Sheikh, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 Dynamic seethroughs: Synthesizing hidden views of moving objects
abstract
This paper presents a method to create an illusion of seeing moving objects through occluding surfaces in a video. This illusion is achieved by transferring information from a camera viewing the occluded area. In typical view interpolation approaches for 3D scenes, some form of correspondence across views is required. For occluded areas, establishing direct correspondence is impossible as information is missing in one of the views. Instead, we use a 2D projective invariant to capture information about occluded objects (which may be moving). Since invariants are quantities that do not change across views, a visually compelling rendering of hidden areas is achieved without the need for explicit correspondences. A piece-wise planar model of the scene allows the entire rendering process to take place without any 3D reconstruction, while still producing visual parallax. Because of the simplicity and robustness of the 2D invariant, we are able to transfer both static backgrounds and moving objects in real time. A complete working system has been implemented that runs live at 5Hz. Applications for this technology include the ability to look through corners at tight intersections for automobile safety, concurrent visualization of a surveillance camera network, and monitoring systems for patients/elderly/children.
Peter C. Barnum, Yaser Sheikh, Ankur Datta, Takeo Kanade
ISMAR3
2008 Linear motion estimation for systems of articulated planes
abstract
In this paper, we describe the explicit application of articulation constraints for estimating the motion of a system of planes. We relate articulations to the relative homography between planes and show that for affine cameras, these articulations translate into linear equality constraints on a linear least squares system, yielding accurate and numerically stable estimates of motion. The global nature of motion estimation allows us to handle areas where there is limited texture information and areas that leave the field of view. Our results demonstrate the accuracy of the algorithm in a variety of cases such as human body tracking, motion estimation of rigid, piecewise planar scenes and motion estimation of triangulated meshes.
Ankur Datta, Yaser Sheikh, Takeo Kanade
CVPR1
2008 On the sustained tracking of human motion
abstract
In this paper, we propose an algorithm for sustained tracking of humans, where we combine frame-to-frame articulated motion estimation with a per-frame body detection algorithm. The proposed approach can automatically recover from tracking error and drift. The frame-to-frame motion estimation algorithm replaces traditional dynamic models within a filtering framework. Stable and accurate per-frame motion is estimated via an image-gradient based algorithm that solves a linear constrained least squares system. The per-frame detector learns appearance of different body parts and dasiasketchespsila expected gradient maps to detect discriminant pose configurations in images. The resulting online algorithm is computationally efficient and has been widely tested on a large dataset of sequences of drivers in vehicles. It shows stability and sustained accuracy over thousands of frames.
Yaser Sheikh, Ankur Datta, Takeo Kanade
FG2
2003 Novel feature vector for image authentication
abstract
As the dissemination of images grows across the Internet, authentication of images is becoming an important research issue. Image watermarking techniques are the most used forms to authenticate an image, however they introduce distortions in the original image and are vulnerable to a multitude of attacks. In this paper we present a novel method to authenticate images. We propose that instead of watermarking images, we can authenticate an image by calculating an invariant feature vector which models the image based on its shape and illumination characteristics. We propose an algorithm to calculate such a robust feature vector which withstands a variety of attacks and hence can be a good alternative to watermarking techniques.
Ankur Datta, Niels da Vitoria Lobo, John J. Leeson
ICME1