VLDB 2026 Research / reviewers in the wild / expert
Tony X. Han
dblp:11/5694
· DBLP profile ↗
42ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 4 first-authorArtificial intelligence and machine learning · 20 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Image recognition and object detection · 29% Video understanding and tracking · 19% Efficient and distributed learning · 19% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 89% Indexing and storage engines · 11% | |
| Computer graphics and multimedia
3 papers |
Image and video processing · 76% Visualization and visual analytics · 24% |
Topics — the 30 heaviest of 39, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
object detection |
0.7 | 4 | 2017 | Learning Efficient Object Detection Models with Knowledge Distillation · NIPS 2017 Detection Evolution with Multi-order Contextual Co-occurrence · CVPR 2013 Detection by detections: Non-parametric detector adaptation for a video · CVPR 2012 |
Information retrieval
image retrieval |
0.3 | 2 | 2015 | Sketch-Based Image Retrieval Through Hypothesis-Driven Object Boundary Selection With HLR Descriptor · IEEE Trans. Multim. 2015 Contextual weighting for vocabulary tree based image retrieval · ICCV 2011 |
Information retrieval
retrieval models |
0.3 | 2 | 2015 | Sketch-Based Image Retrieval Through Hypothesis-Driven Object Boundary Selection With HLR Descriptor · IEEE Trans. Multim. 2015 Contextual weighting for vocabulary tree based image retrieval · ICCV 2011 |
Computer vision › Video understanding and tracking
object tracking |
0.3 | 2 | 2015 | Max-Confidence Boosting With Uncertainty for Visual Tracking · IEEE Trans. Image Process. 2015 Discriminative Tracking by Metric Learning · ECCV (3) 2010 |
Computer vision › Image recognition and object detection › object detection
efficient object detection |
0.3 | 1 | 2017 | Learning Efficient Object Detection Models with Knowledge Distillation · NIPS 2017 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.3 | 1 | 2017 | Learning Efficient Object Detection Models with Knowledge Distillation · NIPS 2017 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2017 | Learning Efficient Object Detection Models with Knowledge Distillation · NIPS 2017 |
Information retrieval › image retrieval
sketch-based image retrieval |
0.2 | 1 | 2015 | Sketch-Based Image Retrieval Through Hypothesis-Driven Object Boundary Selection With HLR Descriptor · IEEE Trans. Multim. 2015 |
Natural language and speech › Information extraction and text analysis › document understanding › document image analysis
font recognition |
0.2 | 1 | 2014 | Large-Scale Visual Font Recognition · CVPR 2014 |
Computer vision › Segmentation and scene understanding › image segmentation › binary segmentation
foreground-background segmentation |
0.2 | 1 | 2013 | Ensemble Video Object Cut in Highly Dynamic Scenes · CVPR 2013 |
Computer vision › Video understanding and tracking
video object segmentation |
0.2 | 1 | 2013 | Ensemble Video Object Cut in Highly Dynamic Scenes · CVPR 2013 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › visual domain adaptation
detector adaptation |
0.1 | 1 | 2012 | Detection by detections: Non-parametric detector adaptation for a video · CVPR 2012 |
Computer vision › Video understanding and tracking
video object detection |
0.1 | 1 | 2012 | Detection by detections: Non-parametric detector adaptation for a video · CVPR 2012 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.1 | 1 | 2011 | Adapting an object detector by considering the worst case: A conservative approach · CVPR 2011 |
Information retrieval › image retrieval
large-scale image retrieval |
0.1 | 1 | 2011 | Contextual weighting for vocabulary tree based image retrieval · ICCV 2011 |
Indexing and storage engines › vector index
vocabulary tree |
0.1 | 1 | 2011 | Contextual weighting for vocabulary tree based image retrieval · ICCV 2011 |
Machine learning › Representation and self-supervised learning › representation learning › feature extraction
bag-of-features |
0.1 | 1 | 2010 | Randomized Locality Sensitive Vocabularies for Bag-of-Features Model · ECCV (3) 2010 |
Computer vision › Video understanding and tracking › object tracking
discriminative tracking |
0.1 | 1 | 2010 | Discriminative Tracking by Metric Learning · ECCV (3) 2010 |
Computer vision › 3D vision
3d scene understanding |
0.1 | 1 | 2009 | Building recognition using sketch-based representations and spectral graph matching · ICCV 2009 |
Computer vision › Image recognition and object detection
pedestrian detection |
0.1 | 1 | 2009 | An HOG-LBP human detector with partial occlusion handling · ICCV 2009 |
Computer vision › 3D vision
structure from motion |
0.1 | 1 | 2009 | Building recognition using sketch-based representations and spectral graph matching · ICCV 2009 |
Computer vision › 3D vision › feature matching › local feature matching
wide-baseline matching |
0.1 | 1 | 2009 | Building recognition using sketch-based representations and spectral graph matching · ICCV 2009 |
Machine learning › Deep learning architectures and training
teacher-student framework |
0.1 | 1 | 2017 | Learning Efficient Object Detection Models with Knowledge Distillation · NIPS 2017 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.1 | 1 | 2015 | Max-Confidence Boosting With Uncertainty for Visual Tracking · IEEE Trans. Image Process. 2015 |
Image and video processing › image segmentation › contour detection
object boundary detection |
0.1 | 1 | 2015 | Sketch-Based Image Retrieval Through Hypothesis-Driven Object Boundary Selection With HLR Descriptor · IEEE Trans. Multim. 2015 |
Computer vision › 3D vision › motion capture
articulated body tracking |
0.1 | 1 | 2006 | Efficient Nonparametric Belief Propagation with Application to Articulated Body Tracking · CVPR (1) 2006 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
belief propagation |
0.1 | 1 | 2006 | Efficient Nonparametric Belief Propagation with Application to Articulated Body Tracking · CVPR (1) 2006 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › belief propagation
nonparametric belief propagation |
0.1 | 1 | 2006 | Efficient Nonparametric Belief Propagation with Application to Articulated Body Tracking · CVPR (1) 2006 |
Visualization and visual analytics › flow visualization
feature tracking |
0.1 | 1 | 2005 | On Optimizing Template Matching via Performance Characterization · ICCV 2005 |
Image and video processing › image matching
template matching |
0.1 | 1 | 2005 | On Optimizing Template Matching via Performance Characterization · ICCV 2005 |
Methods — techniques the papers use, named apart from their topics
hypothesis-driven selection · 0.4histogram of line relationship · 0.4weighted cross-entropy loss · 0.3teacher bounded loss · 0.3hint learning · 0.3online learning · 0.2boosting · 0.2nearest class mean classifier · 0.2max-margin template selection · 0.2local feature metric learning · 0.2local feature embedding · 0.2deformable part model · 0.2spatial verification · 0.1contextual weighting · 0.1randomized vocabularies · 0.1locality-sensitive hashing · 0.1spectral graph matching · 0.1sketch-based representation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Competing ratio loss for discriminative multi-class image classification
Ke Zhang 0005, Yurong Guo 0001, Dongliang Chang, Zhenbing Zhao, Zhanyu Ma, Tony X. Han |
Neurocomputing | 7 |
| 2019 | Object instance detection with pruned Alexnet and extended training data
Rui Wang 0039, Tony X. Han |
Signal Process. Image Commun. | 3 |
| 2018 | Localized Deep Norm-CNN Structure for Face VerificationabstractFace verification is still a big challenging problem due to the different image conditions such as expression, pose, and illumination. To address these challenges, we propose a new Deep Leaning structure called Localized Deep-Norm CNN. Our model focuses on finding the correlations of features inside the sub region of each learning face by adding a localized feature normalization layer. The model can recover all the important correlated features of face images. Intuitively, the Localized Deep-Norm CNN model mimics the primary visual context of the learned face image by combining the localized extracted features representations. The local relational face features are extracted and normalized by assigning each sub-block to a local CNN model. Then, the global features are constructed by combining these localized high-level features to one fully connected layer to produce the final feature space of 4608d dimensions. In our model, two different optimization techniques are proposed to optimize the loss functions. The first optimization modifies the SoftMax loss function by using a cosine similarity metric instead of Euclidean inner-product layer. The second optimization is done by combining different loss functions with different metric learning. Our model achieves robustness accuracy of 99.19% in characterizing the similarity of multiple faces which is 0.16% improving on the LFW performance results. Adil Al-Azzawi, Hasanain Al-Sadr, Jianlin Cheng, Tony X. Han |
ICMLA | 4 |
| 2018 | Surprisingly Easy Network Compression and Data Extension for Object Instance DetectionabstractTo detect instances in unstructured environment with mobile system, we develop a light weight but accurate learning model denoted as B-PA(BING Pruned Alexnet). Our method first utilizes BING(Binarized Normed Gradient) to compute bounding boxes, then builds a compressed network for recognition by pruning neurons and cutting fully connected layers on the original noted Alexnet. Addressing the problem that the training samples for instance detection are limited and of small variation, we extend the training data by combining data augmentation with synthetic generation. Our B-PA model takes only 5.3MB, which is 50 times smaller but with equivalent or even higher accuracy than the original Alexnet. Experiment results demonstrate that our method outperforms the state-of-art instance detection algorithms on WRGB-D Dataset and GMU Kitchen Dataset. Rui Wang 0039, Tony X. Han |
VCIP | 3 |
| 2018 | Residual Networks of Residual Networks: Multilevel Residual NetworksabstractA residual networks family with hundreds or even thousands of layers dominates major image recognition tasks, but building a network by simply stacking residual blocks inevitably limits its optimization ability. This paper proposes a novel residual network architecture, residual networks of residual networks (RoR), to dig the optimization ability of residual networks. RoR substitutes optimizing residual mapping of residual mapping for optimizing original residual mapping. In particular, RoR adds levelwise shortcut connections upon original residual networks to promote the learning capability of residual networks. More importantly, RoR can be applied to various kinds of residual networks (ResNets, Pre-ResNets, and WRN) and significantly boost their performance. Our experiments demonstrate the effectiveness and versatility of RoR, where it achieves the best performance in all residual-network-like structures. Our RoR-3-WRN58-4 + SD models achieve new state-of-the-art results on CIFAR-10, CIFAR-100, and SVHN, with the test errors of 3.77%, 19.73%, and 1.59%, respectively. RoR-3 models also achieve state-of-the-art results compared with ResNets on the ImageNet data set. Ke Zhang 0005, Tony X. Han, Xingfang Yuan, Liru Guo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Learning Efficient Object Detection Models with Knowledge DistillationabstractDespite significant accuracy improvement in convolutional neural networks (CNN) based object detectors, they often require prohibitive runtimes to process an image for real-time applications. State-of-the-art models often use very deep networks with a large number of floating point operations. Efforts such as model compression learn compact models with fewer number of parameters, but with much reduced accuracy. In this work, we propose a new framework to learn compact and fast ob- ject detection networks with improved accuracy using knowledge distillation [20] and hint learning [34]. Although knowledge distillation has demonstrated excellent improvements for simpler classification setups, the complexity of detection poses new challenges in the form of regression, region proposals and less voluminous la- bels. We address this through several innovations such as a weighted cross-entropy loss to address class imbalance, a teacher bounded loss to handle the regression component and adaptation layers to better learn from intermediate teacher distribu- tions. We conduct comprehensive empirical evaluation with different distillation configurations over multiple datasets including PASCAL, KITTI, ILSVRC and MS-COCO. Our results show consistent improvement in accuracy-speed trade-offs for modern multi-class detection models. Guobin Chen, Wongun Choi, Xiang Yu 0002, Tony X. Han, Manmohan Krishna Chandraker |
NIPS | 4 |
| 2016 | Multiple Instance Learning Convolutional Neural Networks for object recognitionabstractConvolutional Neural Networks (CNN) have demonstrated its successful applications in computer vision, speech recognition, and natural language processing. For object recognition, CNNs might be limited by its strict label requirement and an implicit assumption that images are supposed to be target-object-dominated for optimal solutions. However, the labeling procedure, necessitating laying out the locations of target objects, is very tedious, making high-quality large-scale dataset prohibitively expensive. Data augmentation schemes are widely used when deep networks suffer the insufficient training data problem. All the images produced through data augmentation share the same label, which may be problematic since not all data augmentation methods are label-preserving. In this paper, we propose a weakly supervised CNN framework named Multiple Instance Learning Convolutional Neural Networks (MILCNN) to solve this problem. We apply MILCNN framework to object recognition and report state-of-the-art performance on three benchmark datasets: CIFAR10, CIFAR100 and ILSVRC2015 classification dataset. Tony X. Han, Ming-Chang Liu, Ahmad Khodayari-Rostamabad |
ICPR | 2 |
| 2016 | Constellational contour parsing for deformable object detection
Tony X. Han, Zhihai He, Wenming Cao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Robust tracking via monocular active vision for an intelligent teaching system
Rui Wang 0039, Hao Dong 0005, Tony X. Han, Lei Mei |
Vis. Comput. | 3 |
| 2015 | Scene text detection based on component-level fusion and region-level verificationabstractIn this paper, we present a novel scene text detection method that combines the advantages of component-based methods and region-based methods, while overcoming their inherent limitations. We first extract text regions as candidates, and then aggregate these text components in these regions into words and text lines. To separate non-text components in the background from text components, we perform both character-level filtering and word-level classification with a trained linear SVM (support vector machine) classifier. Our extensive experiments on ICDAR2003 and ICDAR2011 datasets have shown that our method outperforms the state-of-the-art methods in text detection. Guanghan Ning, Tony X. Han, Zhihai He |
ICIP | 2 |
| 2015 | Coupled ensemble graph cuts and object verification for animal segmentation from highly cluttered videosabstractIn this paper, we consider animal object segmentation from wildlife monitoring videos captured by motion-triggered cameras, called camera-traps. This is a very challenging task because the wildlife monitoring scenes are often highly cluttered and dynamic. To address this issue, we propose to explore the ideas of coupled ensemble graph cuts and object verification. We consider video object cut as an ensemble of frame-level background-foreground object classifiers which fuse information across frames and refine their segmentation results in a collaborative and iterative manner. To significantly reduce false positives in foreground animal detection and segmentation, we learn an object verification model to further classify if the segmented image patch belongs to the background or the animal. Our extensive experimental results and performance comparisons over a diverse set of challenging camera-trap data, as well as the new Change Detection 2014 benchmark dataset, demonstrate that the proposed framework outperforms various state-of-the-art algorithms and has the capability to handle even the most challenging objects in a wide variety of video sequences. Zhi Zhang 0005, Tony X. Han, Zhihai He |
ICIP | 2 |
| 2015 | Selective Pooling Vector for Fine-Grained RecognitionabstractWe propose a new framework for image recognition by selectively pooling local visual descriptors, and show its superior discriminative power on fine-grained image classification tasks. The representation is based on selecting the most confident local descriptors for nonlinear function learning using a linear approximation in an embedded higher dimensional space. The advantage of our Selective Pooling Vector over the previous state-of-the-art Super Vector and Fisher Vector representations, is that it ensures a more accurate learning function, which proves to be important for classifying details in fine-grained image recognition. Our experimental results corroborate this claim: with a simple linear SVM as the classifier, the selective pooling vector achieves significant performance gains on standard benchmark datasets for various fine-grained tasks such as the CMU Multi-PIE dataset for face recognition, the Caltech-UCSD Bird dataset and the Stanford Dogs dataset for fine-grained object categorization. On all datasets we outperform the state of the arts and boost the recognition rates to 96.4%, 48.9%, 52.0% respectively. Jianchao Yang, Hailin Jin, Eli Shechtman, Jonathan Brandt, Tony X. Han |
WACV | 6 |
| 2015 | Max-Confidence Boosting With Uncertainty for Visual TrackingabstractThe challenges in visual tracking call for a method which can reliably recognize the subject of interests in an environment, where the appearance of both the background and the foreground change with time. Many existing studies model this problem as tracking by classification with online updating of the classification models, however, most of them overlook the ambiguity in visual modeling and do not consider the prior information in the tracking process. In this paper, we present a novel visual tracking method called max-confidence boosting (MCB), which explores a new way of online updating ambiguous visual phenomenon. The MCB framework models uncertainty in prior knowledge utilizing the indeterministic labels, which are used in updating models from previous frames and the new frame. Our proposed MCB tracker allows ambiguity in the tracking process and can effectively alleviate the drift problem. Many experimental results in challenging video sequences verify the success of our method, and our MCB tracker outperforms a number of the state-of-the-art tracking by classification methods. Wen Guo 0003, Liangliang Cao, Tony X. Han, Shuicheng Yan, Changsheng Xu |
IEEE Trans. Image Process. | 3 |
| 2015 | Sketch-Based Image Retrieval Through Hypothesis-Driven Object Boundary Selection With HLR DescriptorabstractThe appearance gap between sketches and photo- realistic images is a fundamental challenge in sketch-based image retrieval (SBIR) systems. The existence of noisy edges on photo- realistic images is a key factor in the enlargement of the appearance gap and significantly degrades retrieval performance . To bridge the gap, we propose a framework consisting of a new line segment -based descriptor named histogram of line relationship (HLR) and a new noise impact reduction algorithm known as object boundary selection . HLR treats sketches and extracted edges of photo- realistic images as a series of piece-wise line segments and captures the relationship between them. Based on the HLR, the object boundary selection algorithm aims to reduce the impact of noisy edges by selecting the shaping edges that best correspond to the object boundaries. Multiple hypotheses are generated for descriptors by hypothetical edge selection. The selection algorithm is formulated to find the best combination of hypotheses to maximize the retrieval score; a fast method is also proposed. To reduce the distraction of false matches in the scoring process, two constraints on spatial and coherent aspects are introduced . We tested the HLR descriptor and the proposed framework on public datasets and a new image dataset of three million images, which we recently collected for SBIR evaluation purposes. We compared the proposed HLR with state-of-the-art descriptors (SHoG, GF-HOG). The experimental results show that our HLR descriptor outperforms them. Combined with the object boundary selection algorithm, our framework significantly improves SBIR performance. Jian Zhang 0002, Tony X. Han, Zhenjiang Miao |
IEEE Trans. Multim. | 3 |
| 2014 | Randomized Support Vector Forest
Xutao Lv, Tony X. Han, Zicheng Liu 0001, Zhihai He |
BMVC | 2 |
| 2014 | Large-Scale Visual Font RecognitionabstractThis paper addresses the large-scale visual font recognition (VFR) problem, which aims at automatic identification of the typeface, weight, and slope of the text in an image or photo without any knowledge of content. Although visual font recognition has many practical applications, it has largely been neglected by the vision community. To address the VFR problem, we construct a large-scale dataset containing 2,420 font classes, which easily exceeds the scale of most image categorization datasets in computer vision. As font recognition is inherently dynamic and open-ended, i.e., new classes and data for existing categories are constantly added to the database over time, we propose a scalable solution based on the nearest class mean classifier (NCM). The core algorithm is built on local feature embedding, local feature metric learning and max-margin template selection, which is naturally amenable to NCM and thus to such open-ended classification problems. The new algorithm can generalize to new classes and new data at little added cost. Extensive experiments demonstrate that our approach is very effective on our synthetic test images, and achieves promising results on real world test images. Jianchao Yang, Hailin Jin, Jonathan Brandt, Eli Shechtman, Aseem Agarwala, Tony X. Han |
CVPR | 7 |
| 2014 | Deep convolutional neural network based species recognition for wild animal monitoringabstractWe proposed a novel deep convolutional neural network based species recognition algorithm for wild animal classification on very challenging camera-trap imagery data. The imagery data were captured with motion triggered camera trap and were segmented automatically using the state of the art graph-cut algorithm. The moving foreground is selected as the region of interests and is fed to the proposed species recognition algorithm. For the comparison purpose, we use the traditional bag of visual words model as the baseline species recognition algorithm. It is clear that the proposed deep convolutional neural network based species recognition achieves superior performance. To our best knowledge, this is the first attempt to the fully automatic computer vision based species recognition on the real camera-trap images. We also collected and annotated a standard camera-trap dataset of 20 species common in North America, which contains 14, 346 training images and 9, 530 testing images, and is available to public for evaluation and benchmark purpose. Guobin Chen, Tony X. Han, Zhihai He, Roland Kays, Tavis Forrester |
ICIP | 2 |
| 2014 | Multi-scale embedded descriptor for shape classification
Tony X. Han, Zhihai He |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Detection Evolution with Multi-order Contextual Co-occurrenceabstractContext has been playing an increasingly important role to improve the object detection performance. In this paper we propose an effective representation, Multi-Order Contextual co-Occurrence (MOCO), to implicitly model the high level context using solely detection responses from a baseline object detector. The so-called (1st-order) context feature is computed as a set of randomized binary comparisons on the response map of the baseline object detector. The statistics of the 1st-order binary context features are further calculated to construct a high order co-occurrence descriptor. Combining the MOCO feature with the original image feature, we can evolve the baseline object detector to a stronger context aware detector. With the updated detector, we can continue the evolution till the contextual improvements saturate. Using the successful deformable-part-model detector [13] as the baseline detector, we test the proposed MOCO evolution framework on the PASCAL VOC 2007 dataset [8] and Caltech pedestrian dataset [7]: The proposed MOCO detector outperforms all known state-of-the-art approaches, contextually boosting deformable part models (ver. 5) [13] by 3.3% in mean average precision on the PASCAL 2007 dataset. For the Caltech pedestrian dataset, our method further reduces the log-average miss rate from 48% to 46% and the miss rate at 1 FPPI from 25% to 23%, compared with the best prior art [6]. Yuanyuan Ding, Jing Xiao 0006, Tony X. Han |
CVPR | 4 |
| 2013 | Ensemble Video Object Cut in Highly Dynamic ScenesabstractWe consider video object cut as an ensemble of frame-level background-foreground object classifiers which fuses information across frames and refine their segmentation results in a collaborative and iterative manner. Our approach addresses the challenging issues of modeling of background with dynamic textures and segmentation of foreground objects from cluttered scenes. We construct patch-level bag-of-words background models to effectively capture the background motion and texture dynamics. We propose a foreground salience graph (FSG) to characterize the similarity of an image patch to the bag-of-words background models in the temporal domain and to neighboring image patches in the spatial domain. We incorporate this similarity information into a graph-cut energy minimization framework for foreground object segmentation. The background-foreground classification results at neighboring frames are fused together to construct a foreground probability map to update the graph weights. The resulting object shapes at neighboring frames are also used as constraints to guide the energy minimization process during graph cut. Our extensive experimental results and performance comparisons over a diverse set of challenging videos with dynamic scenes, including the new Change Detection Challenge Dataset, demonstrate that the proposed ensemble video object cut method outperforms various state-of-the-art algorithms. Xiaobo Ren, Tony X. Han, Zhihai He |
CVPR | 2 |
| 2012 | Histogram of Oriented Normal Vectors for Object Recognition with a Depth Sensor
Xiaoyu Wang 0002, Xutao Lv, Tony X. Han, James Keller 0001, Zhihai He, Marjorie Skubic, Shihong Lao |
ACCV (2) | 4 |
| 2012 | Detection by detections: Non-parametric detector adaptation for a videoabstractWe propose an approach to improving the detection results of a generic offline trained detector on a specific video. Our method does not leverage visual tracking as most detection by tracking methods do. Instead, the proposed detection by detections approach can serve as a more confident initialization for detection by tracking methods. Different from other supervised detector adaptation methods, we constrain the task to videos and no supervised labels for the target video are required for the adaptation; we intend to fill the gap between detection by tracking and pure detection by frames. As a non-parametric detector adaptation method, confident detections are collected to re-rank and to group other detections. We focus on methods with high precision detection results since it is necessitated in real application. Extensive experiments with two state-of-the-art detectors demonstrate the efficacy of our approach. Xiaoyu Wang 0002, Gang Hua 0001, Tony X. Han |
CVPR | 3 |
| 2012 | Semi-supervised learning for robust car windshield tracking and monitoring in live traffic videosabstractThis paper deals with the problem of car-windshield tracking in live traffic video. To avoid a comprehensive labeled dataset that covers most appearance variations, we aim to appropriately involve unlabeled examples and efficiently update the discriminative model in an online semi-supervised setting. Our approach follows the state-of-the- art “learning by detection” approach, yet different from it in the following aspects. First, instead of assigning hard labels to new added examples, we leave them unlabeled. Second, we focus on exploring the intrinsic manifold structure of data marginal distribution and studying its role in kernel function optimization. The proposed online semi-supervised learning framework involves a 3D mean-shift optimization for windshield localization and is followed by a block-based decision for co-driver detection. The experimental results demonstrate the effectiveness of the proposed method. Zhongna Zhou, Tony X. Han, Zhihai He |
ICIP | 2 |
| 2012 | Recognizing Emotions From an Ensemble of FeaturesabstractThis paper details the authors' efforts to push the baseline of emotion recognition performance on the Geneva Multimodal Emotion Portrayals (GEMEP) Facial Expression Recognition and Analysis database. Both subject-dependent and subject-independent emotion recognition scenarios are addressed in this paper. The approach toward solving this problem involves face detection, followed by key-point identification, then feature generation, and then, finally, classification. An ensemble of features consisting of hierarchical Gaussianization, scale-invariant feature transform, and some coarse motion features have been used. In the classification stage, we used support vector machines. The classification task has been divided into person-specific and person-independent emotion recognitions using face recognition with either manual labels or automatic algorithms. We achieve 100% performance for the person-specific one, 66% performance for the person-independent one, and 80% performance for overall results, in terms of classification rate, for emotion recognition with manual identification of subjects. Usman Tariq, Kai-Hsiang Lin, Zhen Li 0028, Vuong Le, Thomas S. Huang, Xutao Lv, Tony X. Han |
IEEE Trans. Syst. Man Cybern. Part B | 9 |
| 2012 | Detection of Sudden Pedestrian Crossings for Driving Assistance SystemsabstractIn this paper, we study the problem of detecting sudden pedestrian crossings to assist drivers in avoiding accidents. This application has two major requirements: to detect crossing pedestrians as early as possible just as they enter the view of the car-mounted camera and to maintain a false alarm rate as low as possible for practical purposes. Although many current sliding-window-based approaches using various features and classification algorithms have been proposed for image-/video-based pedestrian detection, their performance in terms of accuracy and processing speed falls far short of practical application requirements. To address this problem, we propose a three-level coarse-to-fine video-based framework that detects partially visible pedestrians just as they enter the camera view, with low false alarm rate and high speed. The framework is tested on a new collection of high-resolution videos captured from a moving vehicle and yields a performance better than that of state-of-the-art pedestrian detection while running at a frame rate of 55 fps. Yanwu Xu 0001, Dong Xu 0001, Stephen Lin 0001, Tony X. Han, Xianbin Cao 0001, Xuelong Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2011 | Adapting an object detector by considering the worst case: A conservative approachabstractThe performance of an offline-trained classifier can be improved on-site by adapting the classifier towards newly acquired data. However, the adaptation rate is a tuning parameter affecting the performance gain substantially. Poor selection of the adaptation rate may worsen the performance of the original classifier. To solve this problem, we propose a conservative model adaptation method by considering the worst case during the adaptation process. We first construct a random cover of the set of the adaptation data from its partition. For each element in the cover (i.e. a portion of the whole adaptation data set), we define the cross-entropy error function in the form of logistic regression. The element in the cover with the maximum cross-entropy error corresponds to the worst case in the adaptation. Therefore we can convert the conservative model adaptation into the classic min-max optimization problem: finding the adaptation parameters that minimize the maximum of the cross-entropy errors of the cover. Taking the object detection as a testbed, we implement an adapted object detector based on binary classification. Under different adaptation scenarios and different datasets including PASCAL, ImageNet, INRIA, and TUD-Pedestrian, the proposed adaption method achieves significant performance gain and is compared favorably with the state-of-the-art adaptation method with the fine tuned adaptation rate. Without the need of tuning the adaptation rates, the proposed conservative model adaptation method can be extended to other adaptive classification tasks. Tony X. Han, Shihong Lao |
CVPR | 2 |
| 2011 | Emotion recognition from an ensemble of featuresabstractThis work details the authors' efforts to push the baseline of expression recognition performance on a realistic database. Both subject-dependent and subject-independent emotion recognition scenarios are addressed in this work. These two happen frequently in real life settings. The approach towards solving this problem involves face detection, followed by key point identification, then feature generation and then finally classification. An ensemble of features comprising of Hierarchial Gaussianization (HG), Scale Invariant Feature Transform (SIFT) and Optic Flow have been incorporated. In the classification stage we used SVMs. The classification task has been divided into person specific and person independent emotion recognition. Both manual labels and automatic algorithms for person verification have been attempted. They both give similar performance. Usman Tariq, Kai-Hsiang Lin, Zhen Li 0028, Vuong Le, Thomas S. Huang, Xutao Lv, Tony X. Han |
FG | 9 |
| 2011 | Contextual weighting for vocabulary tree based image retrievalabstractIn this paper we address the problem of image retrieval from millions of database images. We improve the vocabulary tree based approach by introducing contextual weighting of local features in both descriptor and spatial domains. Specifically, we propose to incorporate efficient statistics of neighbor descriptors both on the vocabulary tree and in the image spatial domain into the retrieval. These contextual cues substantially enhance the discriminative power of individual local features with very small computational overhead. We have conducted extensive experiments on benchmark datasets, i.e., the UKbench, Holidays, and our new Mobile dataset, which show that our method reaches state-of-the-art performance with much less computation. Furthermore, the proposed method demonstrates excellent scalability in terms of both retrieval accuracy and efficiency on large-scale experiments using 1.26 million images from the ImageNet database as distractors. Xiaoyu Wang 0002, Ming Yang 0007, Timothée Cour, Shenghuo Zhu, Kai Yu 0001, Tony X. Han |
ICCV | 6 |
| 2011 | Generative Group Activity Analysis with Quaternion Descriptor
Guangyu Zhu 0002, Shuicheng Yan, Tony X. Han, Changsheng Xu |
MMM (2) | 3 |
| 2010 | Randomized Locality Sensitive Vocabularies for Bag-of-Features Model
Yadong Mu, Ju Sun, Tony X. Han, Loong Fah Cheong, Shuicheng Yan |
ECCV (3) | 3 |
| 2010 | Discriminative Tracking by Metric Learning
Xiaoyu Wang 0002, Gang Hua 0001, Tony X. Han |
ECCV (3) | 3 |
| 2010 | Efficient Facial Attribute Recognition with a Spatial CodebookabstractThere is a large number of possible facial attributes such as hairstyle, with/without glasses, with/without mustache, etc. Considering large number of facial attributes and their combinations, it is difficult to build attributes classifiers for all possible combinations needed in various applications, especially at the designing stage. To tackle this important and challenging problem, we propose a novel efficient facial attributes recognition algorithm using a learned spatial codebook. The Maximum Entropy and Maximum Orthogonality (MEMO) criterion is followed to learn the spatial codebook. With a spatial codebook constructed at the designing stage, attribute classifiers can be trained on demand with a small number of exemplars with high accuracy on the testing data. Meanwhile, up to 600 times speedup is achieved in the on-demand training process, compared to current state-of-the-art method. The effectiveness of the proposed method is supported by convincing experimental results. Yoshihisa Ijiri, Shihong Lao, Tony X. Han, Hiroshi Murase |
ICPR | 3 |
| 2009 | Building recognition using sketch-based representations and spectral graph matchingabstractIn this work, we address the problem of building recognition across two camera views with large changes in scales and viewpoints. The main idea is to construct a semantically rich sketch-based representation for buildings which is invariant under large scale and perspective changes. After multi-scale maximally stable extremal regions (MSER) detection, the proposed approach finds repeated structural components of buildings, such as window, doors, and facades, and extracts semantically rich features, which are organized into a sketch-based representation of buildings. These descriptors are then clustered in association with different planes of the building and matched across video frames using spectral graph analysis. Our experiments demonstrate that the proposed approach outperforms SIFT-based matching schemes, especially for images with large viewpoint changes. Yu-Chia Chung, Tony X. Han, Zhihai He |
ICCV | 2 |
| 2009 | An HOG-LBP human detector with partial occlusion handlingabstractBy combining Histograms of Oriented Gradients (HOG) and Local Binary Pattern (LBP) as the feature set, we propose a novel human detection approach capable of handling partial occlusion. Two kinds of detectors, i.e., global detector for whole scanning windows and part detectors for local regions, are learned from the training data using linear SVM. For each ambiguous scanning window, we construct an occlusion likelihood map by using the response of each block of the HOG feature to the global detector. The occlusion likelihood map is then segmented by Mean-shift approach. The segmented portion of the window with a majority of negative response is inferred as an occluded region. If partial occlusion is indicated with high likelihood in a certain scanning window, part detectors are applied on the unoccluded regions to achieve the final classification on the current scanning window. With the help of the augmented HOG-LBP feature and the global-part occlusion handling method, we achieve a detection rate of 91.3% with FPPW= 10−6, 94.7% with FPPW= 10−5, and 97.9% with FPPW= 10−4on the INRIA dataset, which, to our best knowledge, is the best human detection performance on the INRIA dataset. The global-part occlusion handling method is further validated using synthesized occlusion data constructed from the INRIA and Pascal dataset. Xiaoyu Wang 0002, Tony X. Han, Shuicheng Yan |
ICCV | 2 |
| 2009 | Video-based Activity Monitoring for Indoor EnvironmentsabstractIn this work, we study how continuous video monitoring and intelligent video processing can be used in eldercare to assist the independent living of elders and to improve the efficiency of eldercare practice. More specifically, we construct an advanced silhouette extraction and tracking algorithm for indoor environments. An adaptive learning method was developed to estimate the physical location and moving speed of a person from a single camera view without calibration. Then hierarchical decision tree and dimension reduction methods were used for human action recognition. We extract important ADL (activities of daily living) statistics for automated functional assessment. Our extensive tests over these massive video datasets demonstrate that the proposed automated activity analysis system is very efficient. Zhongna Zhou, Yu-Chia Chung, Zhihai He, Tony X. Han, James Keller 0001 |
ISCAS | 5 |
| 2009 | Hierarchical Space-Time Model Enabling Efficient Search for Human ActionsabstractWe propose a five-layer hierarchical space-time model (HSTM) for representing and searching human actions in videos. From a features point of view, both invariance and selectivity are desirable characteristics, which seem to contradict each other. To make these characteristics coexist, we introduce a coarse-to-fine search and verification scheme for action searching, based on the HSTM model. Because going through layers of the hierarchy corresponds to progressively turning the knob between invariance and selectivity, this strategy enables search for human actions ranging from rapid movements of sports to subtle motions of facial expressions. The introduction of the Histogram of Gabor Orientations feature makes the searching for actions go smoothly across the hierarchical layers of the HSTM model. The efficient matching is achieved by applying integral histograms to compute the features in the top two layers. The HSTM model was tested on three selected challenging video sequences and on the KTH human action database. And it achieved improvement over other state-of-the-art algorithms. These promising results validate that the HSTM model is both selective and robust for searching human actions. Huazhong Ning, Tony X. Han, Dirk Bernhardt-Walther, Ming Liu 0009, Thomas S. Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Activity Analysis, Summarization, and Visualization for Indoor Human Activity MonitoringabstractIn this work, we study how continuous video monitoring and intelligent video processing can be used in eldercare to assist the independent living of elders and to improve the efficiency of eldercare practice. More specifically, we develop an automated activity analysis and summarization for eldercare video monitoring. At the object level, we construct an advanced silhouette extraction, human detection and tracking algorithm for indoor environments. At the feature level, we develop an adaptive learning method to estimate the physical location and moving speed of a person from a single camera view without calibration. At the action level, we explore hierarchical decision tree and dimension reduction methods for human action recognition. We extract important ADL (activities of daily living) statistics for automated functional assessment. To test and evaluate the proposed algorithms and methods, we deploy the camera system in a real living environment for about a month and have collected more than 200 hours (in excess of 600 G bytes) of activity monitoring videos. Our extensive tests over these massive video datasets demonstrate that the proposed automated activity analysis system is very efficient. Zhongna Zhou, Yu-Chia Chung, Zhihai He, Tony X. Han, James Keller 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2007 | A Drifting-proof Framework for Tracking and Online Appearance LearningabstractIn order to avoid the notorious drifting problem for tracking system, a new integrated appearance learning framework is proposed in this paper. Previous tracking frameworks with appearance learning ability (Black, et al., 1998) either require supervised offline training or will fail inevitably if the tracker locks on the background. While in our framework, no offline training is required. Given the location of the object in the first frame of the video sequence, we model the foreground (the image patch containing the object)/background difference as the transition cost in our tracking objective function. A tracker based on dynamic programming (DP) and template prediction (Toyama, et al., 2001) is carried out on the pixels with high foreground-likelihood. The typical views (i.e. appearance model) proposed by the tracker are used to initialize the states of a hidden Markov model (HMM). With the learned HMM, the tracking results and the appearance model can be further refined until the video sequence and all of these estimated parameters/hidden variables can be well explained by the HMM. Through this iterative procedure, typical views of the object, transition probabilities between the typical views, and location of the object are simultaneously estimated with strong confidence. The experiments show that the proposed framework achieves fairly satisfied results for several challenging video sequences and therefore has many potential applications for video analysis Tony X. Han, Ming Liu 0009, Thomas S. Huang |
WACV | 1 |
| 2006 | Efficient Nonparametric Belief Propagation with Application to Articulated Body TrackingabstractAn efficient Nonparametric Belief Propagation (NBP) algorithm is developed in this paper. While the recently proposed nonparametric belief propagation algorithm has wide applications such as articulated tracking [22, 19], superresolution [6], stereo vision and sensor calibration [10], the hardcore of the algorithm requires repeatedly sampling from products of mixture of Gaussians, which makes the algorithm computationally very expensive. To avoid the slow sampling process, we applied mixture Gaussian density approximation by mode propagation and kernel fitting [2, 7]. The products of mixture of Gaussians are approximated accurately by just a few mode propagation and kernel fitting steps, while the sampling method (e.g. Gibbs sampler) needs many samples to achieve similar approximation results. The proposed algorithm is then applied to articulated body tracking for several scenarios. The experimental results show the robustness and the efficiency of the proposed algorithm. The proposed efficient NBP algorithm also has potentials in other applications mentioned above. Tony X. Han, Huazhong Ning, Thomas S. Huang |
CVPR (1) | 1 |
| 2005 | Online appearance learning by template predictionabstractA new object tracking framework with online appearance learning ability is proposed in this paper. The object appearances are modeled as a set of probability mass functions (PMF), defined as "object-model-set". The averaged object appearance in the video is also modeled as a PMF, named as "universal model". Given an initial template of the target object, which is the only element in the initial object-model-set, the framework tries to track the object by looking into the whole input video sequence. The dynamic programming (DP) is applied to achieve a best spatial-scale matching between the observations and the current model set, across the whole input video. The object-model-set is iteratively updated if the prediction of the matched image patch using current object-model-set is less than its prior computed using universal model. The PMFs of such matched image patches are added into the object-model-set. Thus the object appearance, which is modeled as the set of PMFs of typical views, is learned online. The tracking results can be further refined given the updated object-model-set. This make the proposed tracking framework robust to the appearance variation caused by 3D motion, partial occlusion and illumination. Also, the learned typical views facilitate other vision tasks such as recognition or 3D reconstruction. Tracking results and the learned typical view on the challenging video sequence experimentally show the robustness and strong online learning ability of the proposed frame work. Ming Liu 0009, Tony X. Han, Thomas S. Huang |
AVSS | 2 |
| 2005 | On Optimizing Template Matching via Performance CharacterizationabstractTemplate matching is a fundamental operator in computer vision and is widely used in feature tracking, motion estimation, image alignment, and mosaicing. Under a certain parameterized warping model, the traditional template matching algorithm estimates the geometric warp parameters that minimize the SSD between the target and a warped template. The performance of the template matching can be characterized by deriving the distribution of warp parameter estimate as a function of the ideal template, the ideal warp parameters, and a given noise or perturbation model. In this paper, we assume a discretization of the warp parameter space and derive the theoretical expression for the probability mass function (PMF) of the parameter estimate. As the PMF is also a function of the template size, we can optimize the choice of the template or block size by determining the template/block size that gives the estimate with minimum entropy. Experimental results illustrate the correctness of the theory. An experiment involving feature point tracking in face video is shown to illustrate the robustness of the algorithm in a real-world problem. Tony X. Han, Visvanathan Ramesh, Ying Zhu 0006, Thomas S. Huang |
ICCV | 1 |
| 2004 | Optimal segmentation of signals and its application to image denoising and boundary feature extraction
Tony X. Han, Steven M. Kay, Thomas S. Huang |
ICIP | 1 |