VLDB 2026 Research / reviewers in the wild / expert
Ying Wu 0001
dblp:64/5840-1
· DBLP profile ↗
203ranked-venue papers
20as first author
26since 2021 · last 2025
0000-0002-3523-7054ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 158 · 16 first-author · 18 since 2021Artificial intelligence and machine learning · 139 · 19 first-author · 21 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GPVK-VL: Geometry-Preserving Virtual Keyframes for Visual Localization under Large Viewpoint ChangesabstractVisual localization, the task of determining the position and orientation of a camera, typically involves three core components: offline construction of a keyframe database, efficient online keyframes retrieval, and robust local feature matching. However, significant challenges arise when there are large viewpoint disparities between the query view and the database, such as attempting localization in a corridor previously build from an opposing direction. Intuitively, this issue can be addressed by synthesizing a set of virtual keyframes that cover all viewpoints. However, existing methods for synthesizing novel views to assist localization often fail to ensure geometric accuracy under large viewpoint changes. In this paper, we introduce a confidence-aware geometric prior into 2D Gaussian splatting to ensure the geometric accuracy of the scene. Then we can render novel views through the mesh with clear structures and accurate geometry, even under significant viewpoint changes, enabling the synthesis of a comprehensive set of virtual keyframes. Incorporating this geometry-preserving virtual keyframe database into the localization pipeline significantly enhances the robustness of visual localization. Yunxuan Li, Lei Fan 0005, Xiaoying Xing, Jianxiong Zhou, Ying Wu 0001 |
CVPR | 5 |
| 2025 | Incremental Object Keypoint LearningabstractExisting progress in object keypoint estimation primarily benefits from the conventional supervised learning paradigm based on numerous data labeled with pre-defined keypoints. However, these well-trained models can hardly detect the undefined new keypoints in test time, which largely hinders their feasibility for diverse downstream tasks. To handle this, various solutions are explored but still suffer from either limited generalizability or transferability. Therefore, in this paper, we explore a novel keypoint learning paradigm in that we only annotate new keypoints in the new data and incrementally train the model, without retaining any old data, called Incremental object Keypoint Learning (IKL). A two-stage learning scheme as a novel baseline tailored to IKL is developed. In the first Knowledge Association stage, given the data labeled with only new keypoints, an auxiliary KA-Net is trained to automatically associate the old keypoints to these new ones based on their spatial and intrinsic anatomical relations. In the second Mutual Promotion stage, based on a keypoint-oriented spatial distillation loss, we jointly leverage the auxiliary KA-Net and the old model for knowledge consolidation to mutually promote the estimation of all old and new keypoints. Owing to the investigation of the correlations between new and old keypoints, our proposed method can not just effectively mitigate the catastrophic forgetting of old keypoints, but may even further improve the estimation of the old ones and achieve a positive transfer beyond anti-forgetting. Such an observation has been solidly verified by extensive experiments on different keypoint datasets, where our method exhibits superiority in alleviating the forgetting issue and boosting performance while enjoying labeling efficiency even under the low-shot data regime. Mingfu Liang, Jiahuan Zhou, Xu Zou 0002, Ying Wu 0001 |
CVPR | 4 |
| 2025 | Autoscape: Geometry-Consistent Long-Horizon Scene GenerationabstractThis paper proposes AutoScape, a long-horizon driving scene generation framework. At its core is a novel RGB-D diffusion model that iteratively generates sparse, geometrically consistent keyframes, serving as reliable anchors for the scene's appearance and geometry. To maintain long-range geometric consistency, the model 1) jointly handles image and depth in a shared latent space, 2) explicitly conditions on the existing scene geometry (i.e., rendered point clouds) from previously generated keyframes, and 3) steers the sampling process with a warp-consistent guidance. Given high-quality RGB-D keyframes, a video diffusion model then interpolates between them to produce dense and coherent video frames. AutoScape generates realistic and geometrically consistent driving videos of over 20 seconds, improving the long-horizon FID and FVD scores over the prior state-of-the-art by 48.6\% and 43.0\%, respectively. Ziyu Jiang, Mingfu Liang, Bingbing Zhuang, Jong-Chyi Su, Sparsh Garg, Ying Wu 0001, Manmohan Krishna Chandraker |
ICCV | 7 |
| 2024 | Evidential Active Recognition: Intelligent and Prudent Open-World Embodied PerceptionabstractActive recognition enables robots to intelligently explore novel observations, thereby acquiring more information while circumventing undesired viewing conditions. Recent approaches favor learning policies from simulated or collected data, wherein appropriate actions are more frequently selected when the recognition is accurate. However, most recognition modules are developed under the closed-world assumption, which makes them ill-equipped to handle unexpected inputs, such as the absence of the target object in the current observation. To address this issue, we propose treating active recognition as a sequential evidence-gathering process, providing by-step uncertainty quantification and reliable prediction under the evidence combination theory. Additionally, the reward function developed in this paper effectively characterizes the merit of actions when operating in open-world environments. To evaluate the performance, we collect a dataset from an indoor simulator, encompassing various recognition challenges such as distance, occlusion levels, and visibility. Through a series of experiments on recognition and robustness analysis, we demonstrate the necessity of introducing uncertainties to active recognition and the superior performance of the proposed method. Lei Fan 0005, Mingfu Liang, Yunxuan Li, Gang Hua 0001, Ying Wu 0001 |
CVPR | 5 |
| 2024 | Active Open-Vocabulary Recognition: Let Intelligent Moving Mitigate CLIP LimitationsabstractActive recognition, which allows intelligent agents to explore observations for better recognition performance, serves as a prerequisite for various embodied AI tasks, such as grasping, navigation and room arrangements. Given the evolving environment and the multitude of object classes, it is impractical to include all possible classes during the training stage. In this paper, we aim at advancing active open-vocabulary recognition, empowering embodied agents to actively perceive and classify arbitrary objects. However, directly adopting recent open-vocabulary classification models, like Contrastive Language Image Pretraining (CLIP), poses its unique challenges. Specifically, we observe that CLIP's performance is heavily affected by the viewpoint and occlusions, compromising its reliability in unconstrained embod-ied perception scenarios. Further, the sequential nature of observations in agent-environment interactions necessitates an effective method for integrating features that maintains discriminative strength for open-vocabulary classification. To address these issues, we introduce a novel agent for active open-vocabulary recognition. The proposed method leverages inter-frame and inter-concept similarities to navigate agent movements and to fuse features, without relying on class-specific knowledge. Compared to baseline CLIP model with 29.6% accuracy on ShapeNet dataset, the proposed agent could achieve 53.3% accuracy for open-vocabulary recognition, without any fine-tuning to the equipped CLIP model. Additional experiments conducted with the Habitat simulator further affirm the efficacy of our method. Lei Fan 0005, Jianxiong Zhou, Xiaoying Xing, Ying Wu 0001 |
CVPR | 4 |
| 2024 | AIDE: An Automatic Data Engine for Object Detection in Autonomous DrivingabstractAutonomous vehicle (AV) systems rely on robust perception models as a cornerstone of safety assurance. However, objects encountered on the road exhibit a long-tailed distri-bution, with rare or unseen categories posing challenges to a deployed perception model. This necessitates an expen-sive process of continuously curating and annotating data with significant human effort. We propose to leverage recent advances in vision-language and large language models to design an Automatic Data Engine (AIDE) that automati-cally identifies issues, efficiently curates data, improves the model through auto-labeling, and verifies the model through generation of diverse scenarios. This process operates it-eratively, allowing for continuous self-improvement of the model. We further establish a benchmark for open-world detection on AV datasets to comprehensively evaluate vari-ous learning paradigms, demonstrating our method's supe-rior performance at a reduced cost. Mingfu Liang, Jong-Chyi Su, Samuel Schulter, Sparsh Garg, Shiyu Zhao 0001, Ying Wu 0001, Manmohan Krishna Chandraker |
CVPR | 6 |
| 2024 | Micro-expression spotting with a novel wavelet convolution magnification network in long videos
Jianxiong Zhou, Ying Wu 0001 |
Pattern Recognit. Lett. | 2 |
| 2024 | Outlier-Probability-Based Feature Adaptation for Robust Unsupervised Anomaly Detection on Contaminated Training DataabstractIn the realm of large-scale industrial manufacturing, the precise detection of defective parts stands as a critical imperative. While current unsupervised anomaly detection algorithms exhibit commendable accuracy when applied to clean training datasets, their susceptibility to contaminated training data limits their real-world efficacy. In response to this challenge, this paper proposes a novel Outlier-Probability-Based Feature Adaptation (OPFA) network to realize robust unsupervised anomaly detection on contaminated training data. This method distinguishes itself by maintaining both high accuracy and robustness in the face of contaminated training data, enabling effective learning of discriminative features for anomaly detection. Specifically, the model enhances feature representations through the contraction of normal features and the contrast between normal and outlier features. Our methodology employs an iterative mechanism, featuring three core designs. First, outlier detection evaluates the outlier probabilities of current feature embeddings, providing a basis for subsequent improvements. Second, Gaussian Mixture Model (GMM) is leveraged to model the distributions of normal feature embeddings. Third, the adaptive network refines feature representations based on the GMM models and outlier scores of feature embeddings. Ablation experiments underscore the effectiveness of each component within our model. Furthermore, our approach outperforms other state-of-the-art methods on three benchmark datasets, demonstrating a notable advantage especially in scenarios with contaminated training data. Jianxiong Zhou, Ying Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Flexible Visual Recognition by Evidential Modeling of Confusion and IgnoranceabstractIn real-world scenarios, typical visual recognition systems could fail under two major causes, i.e., the misclassification between known classes and the excusable misbehavior on unknown-class images. To tackle these deficiencies, flexible visual recognition should dynamically predict multiple classes when they are unconfident between choices and reject making predictions when the input is entirely out of the training distribution. Two challenges emerge along with this novel task. First, prediction uncertainty should be separately quantified as confusion depicting inter-class uncertainties and ignorance identifying out-of-distribution samples. Second, both confusion and ignorance should be comparable between samples to enable effective decision-making. In this paper, we propose to model these two sources of uncertainty explicitly with the theory of Subjective Logic. Regarding recognition as an evidence-collecting process, confusion is then defined as conflicting evidence, while ignorance is the absence of evidence. By predicting Dirichlet concentration parameters for singletons, comprehensive subjective opinions, including confusion and ignorance, could be achieved via further evidence combinations. Through a series of experiments on synthetic data analysis, visual recognition, and open-set detection, we demonstrate the effectiveness of our methods in quantifying two sources of uncertainties and dealing with flexible recognition. Lei Fan 0005, Bo Liu 0043, Ying Wu 0001, Gang Hua 0001 |
ICCV | 4 |
| 2023 | Emotional Voice Conversion with Semi-Supervised Generative Modeling
Huayi Zhan, Hong Cheng 0002, Ying Wu 0001 |
INTERSPEECH | 4 |
| 2023 | TOA: Task-oriented Active VQAabstractKnowledge-based visual question answering (VQA) requires external knowledge to answer the question about an image. Early methods explicitly retrieve knowledge from external knowledge bases, which often introduce noisy information. Recently large language models like GPT-3 have shown encouraging performance as implicit knowledge source and revealed planning abilities. However, current large language models can not effectively understand image inputs, thus it remains an open problem to extract the image information and input to large language models. Prior works have used image captioning and object descriptions to represent the image. However, they may either drop the essential visual information to answer the question correctly or involve irrelevant objects to the task-of-interest. To address this problem, we propose to let large language models make an initial hypothesis according to their knowledge, then actively collect the visual evidence required to verify the hypothesis. In this way, the model can attend to the essential visual information in a task-oriented manner. We leverage several vision modules from the perspectives of spatial attention (i.e., Where to look) and attribute attention (i.e., What to look), which is similar to human cognition. The experiments show that our proposed method outperforms the baselines on open-ended knowledge-based VQA datasets and presents clear reasoning procedure with better interpretability. Xiaoying Xing, Mingfu Liang, Ying Wu 0001 |
NeurIPS | 3 |
| 2023 | Avoiding Lingering in Learning Active Recognition by Adversarial DisturbanceabstractThis paper considers the active recognition scenario, where the agent is empowered to intelligently acquire observations for better recognition. The agents usually compose two modules, i.e., the policy and the recognizer, to select actions and predict the category. While using ground-truth class labels to supervise the recognizer, the policy is typically updated with rewards determined by the current in-training recognizer, like whether achieving correct predictions. However, this joint learning process could lead to unintended solutions, like a collapsed policy that only visits views that the recognizer is already sufficiently trained to obtain rewards, which harms the generalization ability. We call this phenomenon lingering to depict the agent being reluctant to explore challenging views during training. Existing approaches to tackle the exploration-exploitation trade-off could be ineffective as they usually assume reliable feedback during exploration to update the estimate of rarely-visited states. This assumption is invalid here as the reward from the recognizer could be insufficiently trained.To this end, our approach integrates another adversarial policy to constantly disturb the recognition agent during training, forming a competing game to promote active explorations and avoid lingering. The reinforced adversary, rewarded when the recognition fails, contests the recognition agent by turning the camera to challenging observations. Extensive experiments across two datasets validate the effectiveness of the proposed approach regarding its recognition performances, learning efficiencies, and especially robustness in managing environmental noises. Lei Fan 0005, Ying Wu 0001 |
WACV | 2 |
| 2023 | Temporal Feature Enhancement Dilated Convolution Network for Weakly-supervised Temporal Action LocalizationabstractWeakly-supervised Temporal Action Localization (WTAL) aims to classify and localize action instances in untrimmed videos with only video-level labels. Existing methods typically use snippet-level RGB and optical flow features extracted from pre-trained extractors directly. Because of two limitations: the short temporal span of snippets and the inappropriate initial features, these WTAL methods suffer from the lack of effective use of temporal information and have limited performance. In this paper, we propose the Temporal Feature Enhancement Dilated Convolution Network (TFE-DCN) to address these two limitations. The proposed TFE-DCN has an enlarged receptive field that covers a long temporal span to observe the full dynamics of action instances, which makes it powerful to capture temporal dependencies between snippets. Furthermore, we propose the Modality Enhancement Module that can enhance RGB features with the help of enhanced optical flow features, making the overall features appropriate for the WTAL task. Experiments conducted on THUMOS’14 and ActivityNet v1.3 datasets show that our proposed approach far outperforms state-of-the-art WTAL methods. Jianxiong Zhou, Ying Wu 0001 |
WACV | 2 |
| 2023 | Discriminative Self-Paced Group-Metric Adaptation for Online Visual IdentificationabstractExisting solutions to instance-level visual identification usually aim to learn faithful and discriminative feature extractors from offline training data and directly use them for the unseen online testing data. However, their performance is largely limited due to the severe distribution shifting issue between training and testing samples. Therefore, we propose a novel online group-metric adaptation model to adapt the offline learned identification models for the online data by learning a series of metrics for all sharing-subsets. Each sharing-subset is obtained from the proposed novel frequent sharing-subset mining module and contains a group of testing samples that share strong visual similarity relationships to each other. Furthermore, to handle potentially large-scale testing samples, we introduce self-paced learning (SPL) to gradually include samples into adaptation from easy to difficult which elaborately simulates the learning principle of humans. Unlike existing online visual identification methods, our model simultaneously takes both the sample-specific discriminant and the set-based visual similarity among testing samples into consideration. Our method is generally suitable to any off-the-shelf offline learned visual identification baselines for online performance improvement which can be verified by extensive experiments on several widely-used visual identification benchmarks. Jiahuan Zhou, Bing Su 0001, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Balancing Between Forgetting and Acquisition in Incremental Subpopulation Learning
Mingfu Liang, Jiahuan Zhou, Ying Wu 0001 |
ECCV (26) | 4 |
| 2022 | Unsupervised Depth Completion and Denoising for RGB-D SensorsabstractDepth information is considered valuable as it describes geometric structures, which benefits various robotic tasks. However, the depth acquired by RGB-D sensors still suffers from two deficiencies, i.e., incompletion and noises. Previous methods complete depth by exploring hand-tuned models or raising surface assumptions, while nowadays, deep approaches intend to solve this problem with rendered image pairs. For depth denoising, as a consequence of different sensor mechanisms, most methods can only work under specific devices. With existing methods, three challenges emerge: the onerous training set collecting process, the mismatch between existing models and present RGB-D sensors, and the non-real-time computation. In this paper, we first state depth completion and denoising are inherently different and without the need to collect or render complete and noiseless ground truths. We address all mentioned challenges with two separate un-supervised learning procedures. The completion network takes color and incomplete depth as input and predicts values to the unobserved area, which combines prior knowledge and color-depth correlations. The denoising step exploits image sequences to construct noise models in a self-supervised manner with the ability to cater to different sensors. Experimental comparisons and ablation studies demonstrate that even without human-labeled ground truths, the proposed method could produce better completion results and also reduce noises in real-time. Lei Fan 0005, Yunxuan Li, Ying Wu 0001 |
ICRA | 4 |
| 2022 | A Novel Phoneme-based Modeling for Text-independent Speaker Identification
Xin Wang 0064, Chuan Xie, Huayi Zhan, Ying Wu 0001 |
INTERSPEECH | 5 |
| 2022 | A Sequential Decision-theoretic Method for Detecting Mobile Robots Localization FailuresabstractMany methods in mobile robotics usually utilize current sensor measurement to evaluate the localization performance of robots, for example in scan matching and particle filter methods. This immediately detecting methodology tend to cause a problem that a well-localization robot obtains a poor sensor measurement, the robot may mistake momentary observation noise for a localization failure. In this paper, we propose a new robot localization fault detection method for resolving this problem. We model robot localization fault detection as a sequential decision-making problem, where the decision of detecting a localization failure is based on a long-term sensor measurements. We employ two parameters of false-positive and false-negative observation error probabilities, which can eliminate the influence of noisy observations. Further, the proposed method derives Bayesian update equations for the integration of a long-term observations and presents an analytic formula representing the belief function of the reliability of localization results. Experimental studies validate the effectiveness of the proposed method. Menghong Liu, Huayi Zhan, Ying Wu 0001 |
IV | 4 |
| 2022 | Introduction to the Special Section of CVPR 2017abstractThe papers in this special section were presented at the Computer Vision and Pattern Recognition conference. Yanxi Liu 0001, James M. Rehg, Camillo J. Taylor, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Learning Meta-Distance for Sequences by Learning a Ground Metric via Virtual Sequence RegressionabstractDistance between sequences is structural by nature because it needs to establish the temporal alignments among the temporally correlated vectors in sequences with varying lengths. Generally, distances for sequences heavily depend on the ground metric between the vectors in sequences to infer the alignments and hence can be viewed as meta-distances upon the ground metric. Learning such meta-distance from multi-dimensional sequences is appealing but challenging. We propose to learn the meta-distance through learning a ground metric for the vectors in sequences. The learning samples are sequences of vectors for which how the ground metric between vectors induces the meta-distance is given. The objective is that the meta-distance induced by the learned ground metric produces large values for sequences from different classes and small values for those from the same class. We formulate the ground metric as a parameter of the meta-distance and regress each sequence to an associated pre-generated virtual sequence w.r.t. the meta-distance, where the virtual sequences for sequences of different classes are well-separated. We develop general iterative solutions to learn both the Mahalanobis metric and the deep metric induced by a neural network for any ground-metric-based sequence distance. Experiments on several sequence datasets demonstrate the effectiveness and efficiency of the proposed methods. Bing Su 0001, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Linear and Deep Order-Preserving Wasserstein Discriminant AnalysisabstractSupervised dimensionality reduction for sequence data learns a transformation that maps the observations in sequences onto a low-dimensional subspace by maximizing the separability of sequences in different classes. It is typically more challenging than conventional dimensionality reduction for static data, because measuring the separability of sequences involves non-linear procedures to manipulate the temporal structures. In this paper, we propose a linear method, called order-preserving Wasserstein discriminant analysis (OWDA), and its deep extension, namely DeepOWDA, to learn linear and non-linear discriminative subspace for sequence data, respectively. We construct novel separability measures between sequence classes based on the order-preserving Wasserstein (OPW) distance to capture the essential differences among their temporal structures. Specifically, for each class, we extract the OPW barycenter and construct the intra-class scatter as the dispersion of the training sequences around the barycenter. The inter-class distance is measured as the OPW distance between the corresponding barycenters. We learn the linear and non-linear transformations by maximizing the inter-class distance and minimizing the intra-class scatter. In this way, the proposed OWDA and DeepOWDA are able to concentrate on the distinctive differences among classes by lifting the geometric relations with temporal constraints. Experiments on four 3D action recognition datasets show the effectiveness of OWDA and DeepOWDA. Bing Su 0001, Jiahuan Zhou, Ji-Rong Wen, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | FLAR: A Unified Prototype Framework for Few-sample Lifelong Active RecognitionabstractIntelligent agents with visual sensors are allowed to actively explore their observations for better recognition performance. This task is referred to as Active Recognition (AR). Currently, most methods toward AR are implemented under a fixed-category setting, which constrains their applicability in realistic scenarios that need to incrementally learn new classes without retraining from scratch. Further, collecting massive data for novel categories is expensive. To address this demand, in this paper, we propose a unified framework towards Few-sample Lifelong Active Recognition (FLAR), which aims at performing active recognition on progressively arising novel categories that only have few training samples. Three difficulties emerge with FLAR: the lifelong recognition policy learning, the knowledge preservation of old categories, and the lack of training samples. To this end, our approach integrates prototypes, a robust representation for limited training samples, into a reinforcement learning solution, which motivates the agent to move towards views resulting in more discriminative features. Catastrophic forgetting during lifelong learning is then alleviated with knowledge distillation. Extensive experiments across two datasets, respectively for object and scene recognition, demonstrate that even without large training samples, the proposed approach could learn to actively recognize novel categories in a class-incremental behavior. Lei Fan 0005, Peixi Xiong, Ying Wu 0001 |
ICCV | 4 |
| 2021 | Contrastive Learning for Label Efficient Semantic SegmentationabstractCollecting labeled data for the task of semantic segmentation is expensive and time-consuming, as it requires dense pixel-level annotations. While recent Convolutional Neural Network (CNN) based semantic segmentation approaches have achieved impressive results by using large amounts of labeled training data, their performance drops significantly as the amount of labeled data decreases. This happens because deep CNNs trained with the de facto cross-entropy loss can easily overfit to small amounts of labeled data. To address this issue, we propose a simple and effective contrastive learning-based training strategy in which we first pretrain the network using a pixel-wise, label-based contrastive loss, and then fine-tune it using the cross-entropy loss. This approach increases intra-class compactness and inter-class separability, thereby resulting in a better pixel classifier. We demonstrate the effectiveness of the proposed training strategy using the Cityscapes and PASCAL VOC 2012 segmentation datasets. Our results show that pretraining with the proposed contrastive loss results in large performance gains (more than 20% absolute improvement in some settings) when the amount of labeled data is limited. In many settings, the proposed contrastive pretraining strategy, which does not use any additional data, is able to match or outperform the widely-used ImageNet pretraining strategy that uses more than a million additional labeled images. Xiangyun Zhao, Raviteja Vemulapalli, Philip Andrew Mansfield, Boqing Gong, Bradley Green, Lior Shapira, Ying Wu 0001 |
ICCV | 7 |
| 2021 | Morphable Detector for Object Detection on DemandabstractMany emerging applications of intelligent robots need to explore and understand new environments, where it is desirable to detect objects of novel classes on the fly with minimum online efforts. This is an object detection on demand (ODOD) task. It is challenging, because it is impossible to annotate a large number of data on the fly, and the embedded systems are usually unable to perform back-propagation which is essential for training. Most existing few-shot detection methods are confronted here as they need extra training. We propose a novel morphable detector (MD), that simply "morphs" some of its changeable parameters online estimated from the few samples, so as to detect novel classes without any extra training. The MD has two sets of parameters, one for the feature embedding and the other for class representation (called "prototypes"). Each class is associated with a hidden prototype to be learned by integrating the visual and semantic embeddings. The learning of the MD is based on the alternate learning of the feature embedding and the prototypes in an EM-like approach which allows the recovery of an unknown prototype from a few samples of a novel class. Once an MD is learned, it is able to use a few samples of a novel class to directly compute its prototype to fulfill the online morphing process. We have shown the superiority of the MD in Pascal [12], COCO [27] and FSOD [13] datasets. Xiangyun Zhao, Xu Zou 0002, Ying Wu 0001 |
ICCV | 3 |
| 2021 | Towards Unconstrained Facial Landmark Detection Robust to Diverse Cropping MannersabstractFacial landmark detection is one crucial step for face-based image/video analysis. Despite the fact that recently many facial landmark detection models have achieved remarkable performance, most state-of-the-art heatmap regression-based methods heavily rely on initialization of the face detector. However, there inevitably exists semantic gaps among different annotators or face detectors. An improper facial bounding box will tremendously drop off the performance of the facial landmark detection model. Facial landmark detection would be more practical if robust to face images cropped by diverse manners (see Figure 1, the col.1 shows face images cropped by a proper bounding box, col.2 and col.3 show face images cropped by an oversize and a small bounding boxes respectively). To this end, we present a “Unconstrained Facial Landmark Detection(UFLD)” mechanism, that aims at enhancing the robustness of facial landmark detection, to deal with the inconsistent cropping manner issue. UFLD consists of two aspects: a Transformation-Invariant Landmark Detector(TILD) and an Availability-Guided Solver(AGS). TILD gives the ability to detect consistent landmarks for face images cropped by diverse manners. And AGS can alleviate the by-effect of “landmarks outside the image” caused by improper cropping results or TILD, and further promote the performance. The proposed mechanism achieved above 6.5% improvement in standard normalized landmarks mean error reduction on face images cropped by diverse manners compared to baselines. Xu Zou 0002, Luxin Yan, Sheng Zhong 0001, Ying Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Multi-Scale Low-Discriminative Feature Reactivation for Weakly Supervised Object LocalizationabstractFor weakly supervised object localization (WSOL), how to avoid the network focusing only on some small discriminative parts is a main challenge needed to solve. The widely-used Class Activation Mapping (CAM) based paradigm usually employs Adversarial Learning (AL) strategy to search more object parts by constantly hiding discovered object features, but the adversarial process is difficult to control. In this paper, we propose a novel CAM-based framework with Multi-scale Low-Discriminative Feature Reactivation (mLDFR) for WSOL. The mLDFR framework reactivates the low-discriminative object parts via bottom-up continuous feature maps recalibration and multi-scale object category mapping. Compared with the AL-based methods, our method fully improves the localization power of the network without damaging the classification power and can perform multi-instance localization, which are hard to achieve under the AL-based framework. Moreover, the mLDFR framework is flexible, and can be built on the top of various classical CNN backbones. Experimental results demonstrate the superiority of our method. With VGG16 as backbone, we achieve 46.96% Cls-Loc top1 err and 66.12% CorLoc on ILSVRC2014, 38.07% Cls-Loc top1 err and 75.04% CorLoc on CUB200-2011, surpassing the state-of-the-arts by a large margin. Bo Wang 0147, Chunfeng Yuan, Bing Li 0001, Xinmiao Ding, Zeya Li, Ying Wu 0001, Weiming Hu 0004 |
IEEE Trans. Image Process. | 6 |
| 2020 | Uncertainty-Aware Score Distribution Learning for Action Quality AssessmentabstractAssessing action quality from videos has attracted growing attention in recent years. Most existing approaches usually tackle this problem based on regression algorithms, which ignore the intrinsic ambiguity in the score labels caused by multiple judges or their subjective appraisals. To address this issue, we propose an uncertainty-aware score distribution learning (USDL) approach for action quality assessment (AQA). Specifically, we regard an action as an instance associated with a score distribution, which describes the probability of different evaluated scores. Moreover, under the circumstance where finer-grained score labels are available (e.g., difficulty degree of an action or multiple scores from different judges), we further devise a multi-path uncertainty-aware score distribution learning (MUSDL) method to explore the disentangled components of a score. In order to demonstrate the effectiveness of our proposed methods, We conduct experiments on two AQA datasets containing various Olympic actions. Our approaches set new state-of-the-arts under the Spearman's Rank Correlation (i.e., 0.8102 on AQA-7 and 0.9273 on MTL-AQA). Yansong Tang, Zanlin Ni, Jiahuan Zhou, Jiwen Lu, Ying Wu 0001, Jie Zhou 0001 |
CVPR | 6 |
| 2020 | TA-Student VQA: Multi-Agents Training by Self-QuestioningabstractThere are two main challenges in Visual Question Answering (VQA). The first one is that each model obtains its strengths and shortcomings when applied to several questions; what is more, the “ceiling effect” for specific questions is difficult to overcome with simple consecutive training. The second challenge is that even the state-of-the-art dataset is of large scale, questions targeted at a single image are off in format and lack diversity in content. We introduce our self-questioning model with multi-agent training: TA-student VQA. This framework differs from standard VQA algorithms by involving question-generating mechanisms and collaborative learning questions between question-answering agents. Thus, TA-student VQA overcomes the limitation of the content diversity and format variation of questions and improves the overall performance of multiple question-answering agents. We evaluate our model on VQA-v2, which outperforms algorithms without such mechanisms. In addition, TA-student VQA achieves a greater model capacity, allowing it to answer more generated questions in addition to those in the annotated datasets. Peixi Xiong, Ying Wu 0001 |
CVPR | 2 |
| 2020 | Online Joint Multi-Metric Adaptation From Frequent Sharing-Subset Mining for Person Re-IdentificationabstractPerson Re-IDentification (P-RID), as an instance-level recognition problem, still remains challenging in computer vision community. Many P-RID works aim to learn faithful and discriminative features/metrics from offline training data and directly use them for the unseen online testing data. However, their performance is largely limited due to the severe data shifting issue between training and testing data. Therefore, we propose an online joint multi-metric adaptation model to adapt the offline learned P-RID models for the online data by learning a series of metrics for all the sharing-subsets. Each sharing-subset is obtained from the proposed novel frequent sharing-subset mining module and contains a group of testing samples which share strong visual similarity relationships to each other. Unlike existing online P-RID methods, our model simultaneously takes both the sample-specific discriminant and the set-based visual similarity among testing samples into consideration so that the adapted multiple metrics can refine the discriminant of all the given testing samples jointly via a multi-kernel late fusion framework. Our proposed model is generally suitable to any offline learned P-RID baselines for online boosting, the performance improvement by our model is not only verified by extensive experiments on several widely-used P-RID benchmarks (CUHK03, Market1501, DukeMTMC-reID and MSMT17) and state-of-the-art P-RID baselines but also guaranteed by the provided in-depth theoretical analyses. Jiahuan Zhou, Bing Su 0001, Ying Wu 0001 |
CVPR | 3 |
| 2020 | Object Detection with a Unified Label Space from Multiple Datasets
Xiangyun Zhao, Samuel Schulter, Yi-Hsuan Tsai, Manmohan Krishna Chandraker, Ying Wu 0001 |
ECCV (14) | 6 |
| 2020 | Baselines Extraction from Curved Document Images via Slope Fields RecoveryabstractBaselines estimation is a critical preprocessing step for many tasks of document image processing and analysis. The problem is very challenging due to arbitrarily complicated page layouts and various types of image quality degradations. This paper proposes a method based on slope fields recovery for curved baseline extraction from a distorted document image captured by a hand-held camera. Our method treats the curved baselines as the solution curves of an ordinary differential equation defined on a slope field. By assuming the page shape is a smooth and developable surface, we investigate a type of intrinsic geometric constraints of baselines to estimate the latent slope field. The curved baselines are finally obtained by solving an ordinary differential equation through the Euler method. Unlike the traditional text-lines based methods, our method is free from text-lines detection and segmentation. It can exploit multiple visual cues other than horizontal text-lines available in images for baselines extraction and is quite robust to document scripts, various types of image quality degradation (e.g., image distortion, blur and non-uniform illumination), large areas of non-textual objects and complex page layouts. Extensive experiments on synthetic and real-captured document images are implemented to evaluate the performance of the proposed method. Gaofeng Meng, Chunhong Pan, Shiming Xiang, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Learning Low-Dimensional Temporal Representations with Latent AlignmentsabstractLow-dimensional discriminative representations enhance machine learning methods in both performance and complexity. This has motivated supervised dimensionality reduction (DR), which transforms high-dimensional data into a discriminative subspace. Most DR methods require data to be i.i.d. However, in some domains, data naturally appear in sequences, where the observations are temporally correlated. We propose a DR method, namely, latent temporal linear discriminant analysis (LT-LDA), to learn low-dimensional temporal representations. We construct the separability among sequence classes by lifting the holistic temporal structures, which are established based on temporal alignments and may change in different subspaces. We jointly learn the subspace and the associated latent alignments by optimizing an objective that favors easily separable temporal structures. We show that this objective is connected to the inference of alignments and thus allows for an iterative solution. We provide both theoretical insight and empirical evaluations on several real-world sequence datasets to show the applicability of our method. Bing Su 0001, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Learning Visual Instance Retrieval from Failure: Efficient Online Local Metric Adaptation from Negative SamplesabstractExisting visual instance retrieval (VIR) approaches attempt to learn a faithful global matching metric or discriminative feature embedding offline to cover enormous visual appearance variations, so as to directly use it online on various unseen probes for retrieval. However, their requirement for a huge set of positive training pairs is very demanding in practice and the performance is largely constrained for the unseen testing samples due to the severe data shifting issue. In contrast, this paper advocates a different paradigm: part of the learning can be performed online but with nominal costs, so as to achieve online metric adaptation for different query probes. By exploiting easily-available negative samples, we propose a novel solution to achieve the optimal local metric adaptation effectively and efficiently. The insight of our method is the local hard negative samples can actually provide tight constraints to fine tune the metric locally. Our local metric adaptation method is generally applicable to be used on top of any offline-learned baselines. In addition, this paper gives in-depth theoretical analyses of the proposed method to guarantee the reduction of the classification error both asymptotically and practically. Extensive experiments on various VIR tasks have confirmed our effectiveness and superiority. Jiahuan Zhou, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Does Learning Specific Features for Related Parts Help Human Pose Estimation?abstractHuman pose estimation (HPE) is inherently a homogeneous multi-task learning problem, with the localization of each body part as a different task. Recent HPE approaches universally learn a shared representation for all parts, from which their locations are linearly regressed. However, our statistical analysis indicates not all parts are related to each other. As a result, such a sharing mechanism can lead to negative transfer and deteriorate the performance. This potential issue drives us to raise an interesting question. Can we identify related parts and learn specific features for them to improve pose estimation? Since unrelated tasks no longer share a high-level representation, we expect to avoid the adverse effect of negative transfer. In addition, more explicit structural knowledge, e.g., ankles and knees are highly related, is incorporated into the model, which helps resolve ambiguities in HPE. To answer this question, we first propose a data-driven approach to group related parts based on how much information they share. Then a part-based branching network (PBN) is introduced to learn representations specific to each part group. We further present a multi-stage version of this network to repeatedly refine intermediate features and pose estimates. Ablation experiments indicate learning specific features significantly improves the localization of occluded parts and thus benefits HPE. Our approach also outperforms all state-of-the-art methods on two benchmark datasets, with an outstanding advantage when occlusion occurs. Wei Tang 0016, Ying Wu 0001 |
CVPR | 2 |
| 2019 | Visual Query Answering by Entity-Attribute Graph Matching and ReasoningabstractVisual Query Answering (VQA) is of great significance in offering people convenience: one can raise a question for details of objects, or high-level understanding about the scene, over an image. This paper proposes a novel method to address the VQA problem. In contrast to prior works, our method that targets single scene VQA, replies on graph-based techniques and involves reasoning. In a nutshell, our approach is centered on three graphs. The first graph, referred to as inference graph G_I, is constructed via learning over labeled data. The other two graphs, referred to as query graph Q and entity-attribute graph EAG, are generated from natural language query NLQ and image Img, that are issued from users, respectively. As EAG often does not take sufficient information to answer Q, we develop techniques to infer missing information of EAG with G_I. Based on EAG and Q, we provide techniques to find matches of Q in EAG, as the answer of NLQ in Img. Unlike commonly used VQA methods that are based on end-to-end neural networks, our graph-based method shows well-designed reasoning capability, and thus is highly interpretable. We also create a dataset on soccer match (Soccer-VQA) with rich annotations. The experimental results show that our approach outperforms the state-of-the-art method and has high potential for future investigation. Peixi Xiong, Huayi Zhan, Xin Wang 0064, Baivab Sinha, Ying Wu 0001 |
CVPR | 5 |
| 2019 | Order-Preserving Wasserstein Discriminant AnalysisabstractSupervised dimensionality reduction for sequence data projects the observations in sequences onto a low-dimensional subspace to better separate different sequence classes. It is typically more challenging than conventional dimensionality reduction for static data, because measuring the separability of sequences involves non-linear procedures to manipulate the temporal structures. This paper presents a linear method, namely Order-preserving Wasserstein Discriminant Analysis (OWDA), which learns the projection by maximizing the inter-class distance and minimizing the intra-class scatter. For each class, OWDA extracts the order-preserving Wasserstein barycenter and constructs the intra-class scatter as the dispersion of the training sequences around the barycenter. The inter-class distance is measured as the order-preserving Wasserstein distance between the corresponding barycenters. OWDA is able to concentrate on the distinctive differences among classes by lifting the geometric relations with temporal constraints. Experiments show that OWDA achieves competitive results on three 3D action recognition datasets. Bing Su 0001, Jiahuan Zhou, Ying Wu 0001 |
ICCV | 3 |
| 2019 | Recognizing Part Attributes With Insufficient DataabstractRecognizing the attributes of objects and their parts is central to many computer vision applications. Although great progress has been made to apply object-level recognition, recognizing the attributes of parts remains less applicable since the training data for part attributes recognition is usually scarce especially for internet-scale applications. Furthermore, most existing part attribute recognition methods rely on the part annotations which are more expensive to obtain. In order to solve the data insufficiency problem and get rid of dependence on the part annotation, we introduce a novel Concept Sharing Network (CSN) for part attribute recognition. A great advantage of CSN is its capability of recognizing the part attribute (a combination of part location and appearance pattern) that has insufficient or zero training data, by learning the part location and appearance pattern respectively from the training data that usually mix them in a single label. Extensive experiments on CUB, Celeb A, and a newly proposed human attribute dataset demonstrate the effectiveness of CSN and its advantages over other methods, especially for the attributes with few training samples. Further experiments show that CSN can also perform zero-shot part attribute recognition. Xiangyun Zhao, Yi Yang 0007, Feng Zhou 0002, Xiao Tan 0001, Yuchen Yuan, Sid Ying-Ze Bao, Ying Wu 0001 |
ICCV | 7 |
| 2019 | Learning Robust Facial Landmark Detection via Hierarchical Structured EnsembleabstractHeatmap regression-based models have significantly advanced the progress of facial landmark detection. However, the lack of structural constraints always generates inaccurate heatmaps resulting in poor landmark detection performance. While hierarchical structure modeling methods have been proposed to tackle this issue, they all heavily rely on manually designed tree structures. The designed hierarchical structure is likely to be completely corrupted due to the missing or inaccurate prediction of landmarks. To the best of our knowledge, in the context of deep learning, no work before has investigated how to automatically model proper structures for facial landmarks, by discovering their inherent relations. In this paper, we propose a novel Hierarchical Structured Landmark Ensemble (HSLE) model for learning robust facial landmark detection, by using it as the structural constraints. Different from existing approaches of manually designing structures, our proposed HSLE model is constructed automatically via discovering the most robust patterns so HSLE has the ability to robustly depict both local and holistic landmark structures simultaneously. Our proposed HSLE can be readily plugged into any existing facial landmark detection baselines for further performance improvement. Extensive experimental results demonstrate our approach significantly outperforms the baseline by a large margin to achieve a state-of-the-art performance. Xu Zou 0002, Sheng Zhong 0001, Luxin Yan, Xiangyun Zhao, Jiahuan Zhou, Ying Wu 0001 |
ICCV | 6 |
| 2019 | VIASEG: Visual Information Assisted Lightweight Point Cloud SegmentationabstractRapid and precise point cloud segmentation is one of the prerequisites for real-time and robust autonomous perception and environmental understanding, which requires a balance between speed and accuracy in architecture design. However, recent lightweight architectures, though fast enough, rely on domain adaptation from time-consuming-constructed synthetic dataset and sophisticated post-processing procedure to improve their performance, neglecting the rich visual information acquired by cameras aside from LiDAR sensors. In this paper, such color information is embedded at data-level to boost the performance of real-time point cloud segmentation. Furthermore, a multiscale lightweight fully convolutional network, VIASeg, is proposed based on the newly designed Super Squeeze Residual module and Semantic Connection from higher convolutional layers to lower layers, which improves the performance by feature denoising with high level semantic information. The superiority of the proposed method is validated and demonstrated in the comparative and ablative experimental analysis, while maintaining the real-time characteristic. Zhibin Zhong, Chi Zhang 0020, Yuehu Liu, Ying Wu 0001 |
ICIP | 4 |
| 2019 | Learning Distance for Sequences by Learning a Ground MetricabstractLearning distances that operate directly on multi-dimensional sequences is challenging because such distances are structural by nature and the vectors in sequences are not independent. Generally, distances for sequences heavily depend on the ground metric between the vectors in sequences. We propose to learn the distance for sequences through learning a ground Mahalanobis metric for the vectors in sequences. The learning samples are sequences of vectors for which how the ground metric between vectors induces the overall distance is given, and the objective is that the distance induced by the learned ground metric produces large values for sequences from different classes and small values for those from the same class. We formulate the metric as a parameter of the distance, bring closer each sequence to an associated virtual sequence w.r.t. the distance to reduce the number of constraints, and develop a general iterative solution for any ground-metric-based sequence distance. Experiments on several sequence datasets demonstrate the effectiveness and efficiency of our method. Bing Su 0001, Ying Wu 0001 |
ICML | 2 |
| 2018 | Easy Identification From Better Constraints: Multi-Shot Person Re-Identification From Reference ConstraintsabstractMulti-shot person re-identification (MsP-RID) utilizes multiple images from the same person to facilitate identification. Considering the fact that motion information may not be discriminative nor reliable enough for MsP-RID, this paper is focused on handling the large variations in the visual appearances through learning discriminative visual metrics for identification. Existing metric learning-based methods usually exploit pair-wise or triple-wise similarity constraints, that generally demands intensive optimization in metric learning, or leads to degraded performances by using sub-optimal solutions. In addition, as the training data are significantly imbalanced, the learning can be largely dominated by the negative pairs and thus produces unstable and non-discriminative results. In this paper, we propose a novel type of similarity constraint. It assigns the sample points to a set of reference points to produce a linear number of reference constraints. Several optimal transport-based schemes for reference constraint generation are proposed and studied. Based on those constraints, by utilizing a typical regressive metric learning model, the closed-form solution of the learned metric can be easily obtained. Extensive experiments and comparative studies on several public MsP-RID benchmarks have validated the effectiveness of our method and its significant superiority over the state-of-the-art MsP-RID methods in terms of both identification accuracy and running speed. Jiahuan Zhou, Bing Su 0001, Ying Wu 0001 |
CVPR | 3 |
| 2018 | Fictitious GAN: Training GANs with Historical Models
Yin Xia, Randall Berry, Ying Wu 0001 |
ECCV (1) | 5 |
| 2018 | Exploiting Vector Fields for Geometric Rectification of Distorted Document Images
Gaofeng Meng, Yuanqi Su, Ying Wu 0001, Shiming Xiang, Chunhong Pan |
ECCV (16) | 3 |
| 2018 | Deeply Learned Compositional Models for Human Pose Estimation
Wei Tang 0016, Pei Yu, Ying Wu 0001 |
ECCV (3) | 3 |
| 2018 | A Modulation Module for Multi-task Learning with Applications in Image Retrieval
Xiangyun Zhao, Xiaohui Shen, Xiaodan Liang, Ying Wu 0001 |
ECCV (1) | 5 |
| 2018 | Fused Discriminative Metric Learning for Low Resolution Pedestrian DetectionabstractLow resolution (LR) is one of the most challenging factor in pedestrian detection. In this paper, we propose a fused discriminative metric learning (F-DML) approach for low resolution pedestrian detection without explicit super resolution. We firstly learn a discriminative high resolution (HR) feature space as target space. Then, an optimal Mahanalobis metric is learned to transform the LR feature space into a new LR classification space, which largely preserves the discriminative structure of the HR feature space. Finally, a weighted K-nearest neighbors classifier is applied in the LR classification space which inherits good discrimination from HR feature space. A new training strategy is proposed to find the fewest and most representative LR-HR exemplars. In addition, we build a new dataset for the evaluation of low resolution pedestrian detection methods. Extensive experimental results demonstrate that the proposed approach performs favorably against the state-of-the-art methods. Xinzhao Li, Yuehu Liu, Zeqi Chen, Jiahuan Zhou, Ying Wu 0001 |
ICIP | 5 |
| 2018 | Learning Low-Dimensional Temporal RepresentationsabstractLow-dimensional discriminative representations enhance machine learning methods in both performance and complexity, motivating supervised dimensionality reduction (DR) that transforms high-dimensional data to a discriminative subspace. Most DR methods require data to be i.i.d., however, in some domains, data naturally come in sequences, where the observations are temporally correlated. We propose a DR method called LT-LDA to learn low-dimensional temporal representations. We construct the separability among sequence classes by lifting the holistic temporal structures, which are established based on temporal alignments and may change in different subspaces. We jointly learn the subspace and the associated alignments by optimizing an objective which favors easily-separable temporal structures, and show that this objective is connected to the inference of alignments, thus allows an iterative solution. We provide both theoretical insight and empirical evaluation on real-world sequence datasets to show the interest of our method. Bing Su 0001, Ying Wu 0001 |
ICML | 2 |
| 2018 | Active target tracking: A simplified view aligning method for binocular camera model
Xinzhao Li, Yuanqi Su, Yuehu Liu, Shaozhuo Zhai, Ying Wu 0001 |
Comput. Vis. Image Underst. | 5 |
| 2018 | Guest Editors' Introduction to the Special Section on Learning with Shared Information for Computer Vision and Multimedia AnalysisabstractThe twelve papers in this special section focus on learning systems with shared information for computer vision and multimedia communication analysis. In the real world, a realistic setting for computer vision or multimedia recognition problems is that we have some classes containing lots of training data and many classes containing a small amount of training data. Therefore, how to use frequent classes to help learning rare classes for which it is harder to collect the training data is an open question. Learning with shared information is an emerging topic in machine learning, computer vision and multimedia analysis. There are different levels of components that can be shared during concept modeling and machine learning stages, such as sharing generic object parts, sharing attributes, sharing transformations, sharing regularization parameters and sharing training examples, etc. Regarding the specific methods, multi-task learning, transfer learning and deep learning can be seen as using different strategies to share information. These learning with shared information methods are very effective in solving real-world large-scale problems. Trevor Darrell, Christoph H. Lampert, Nicu Sebe, Ying Wu 0001, Yan Yan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Discriminative Dimensionality Reduction for Multi-Dimensional SequencesabstractSince the observables at particular time instants in a temporal sequence exhibit dependencies, they are not independent samples. Thus, it is not plausible to apply i.i.d. assumption-based dimensionality reduction methods to sequence data. This paper presents a novel supervised dimensionality reduction approach for sequence data, called Linear Sequence Discriminant Analysis (LSDA). It learns a linear discriminative projection of the feature vectors in sequences to a lower-dimensional subspace by maximizing the separability of the sequence classes such that the entire sequences are holistically discriminated. The sequence class separability is constructed based on the sequence statistics, and the use of different statistics produces different LSDA methods. This paper presents and compares two novel LSDA methods, namely M-LSDA and D-LSDA. M-LSDA extracts model-based statistics by exploiting the dynamical structure of the sequence classes, and D-LSDA extracts the distance-based statistics by computing the pairwise similarity of samples from the same sequence class. Extensive experiments on several different tasks have demonstrated the effectiveness and the general applicability of the proposed methods. Bing Su 0001, Xiaoqing Ding, Hao Wang 0005, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Motion Correlation Discovery for Visual TrackingabstractMotion information plays an important role in identifying moving objects, which has not been well utilized in state-of-the-art tracking algorithms. In this letter, we propose a unified framework integrating two tracking problems, i.e., pixel-level foreground probabilistic inference and motion parameter estimation. Our model employs motion fields to propagate probability forward, and discovers motion patterns in the spatial domain to distinguish targets from the background. It takes advantage of continuity and inertia of both target and camera motion, and provides reliable evidence to resolve confusion caused by appearance similarity between targets and the background. Target localization is effectively achieved from the pixel-level foreground probabilistic map. Experimental results demonstrate that the proposed method significantly improves our baseline method and achieves performance comparable to state-of-the-art tracking methods with more complex features. Sheng Zhong 0001, Ying Wu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2018 | Joint Video Object Discovery and Segmentation by Coupled Dynamic Markov NetworksabstractIt is a challenging task to extract segmentation mask of a target from a single noisy video, which involves object discovery coupled with segmentation. To solve this challenge, we present a method to jointly discover and segment an object from a noisy video, where the target disappears intermittently throughout the video. Previous methods either only fulfill video object discovery, or video object segmentation presuming the existence of the object in each frame. We argue that jointly conducting the two tasks in a unified way will be beneficial. In other words, video object discovery and video object segmentation tasks can facilitate each other. To validate this hypothesis, we propose a principled probabilistic model, where two dynamic Markov networks are coupled-one for discovery and the other for segmentation. When conducting the Bayesian inference on this model using belief propagation, the bi-directional message passing reveals a clear collaboration between these two inference tasks. We validated our proposed method in five data sets. The first three video data sets, i.e., the SegTrack data set, the YouTube-objects data set, and the Davis data set, are not noisy, where all video frames contain the objects. The two noisy data sets, i.e., the XJTU-Stevens data set, and the Noisy-ViDiSeg data set, newly introduced in this paper, both have many frames that do not contain the objects. When compared with state of the art, it is shown that although our method produces inferior results on video data sets without noisy frames, we are able to obtain better results on video data sets with noisy frames. Ziyi Liu 0001, Le Wang 0003, Gang Hua 0001, Qilin Zhang 0004, Zhenxing Niu, Ying Wu 0001, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 6 |
| 2018 | Heteroscedastic Max-Min Distance Analysis for Dimensionality ReductionabstractMax-min distance analysis (MMDA) performs dimensionality reduction by maximizing the minimum pairwise distance between classes in the latent subspace under the homoscedastic assumption, which can address the class separation problem caused by the Fisher criterion, but is incapable of tackling heteroscedastic data properly. In this paper, we propose two heteroscedastic MMDA (HMMDA) methods to employ the differences of class covariances. Whitened HMMDA (WHMMDA) extends MMDA by utilizing the Chernoff distance as the separability measure between classes in the whitened space. Orthogonal HMMDA (OHMMDA) incorporates the maximization of the minimal pairwise Chernoff distance and the minimization of class compactness into a trace quotient formulation with an orthogonal constraint of the transformation, which can be solved by bisection search. Two variants of OHMMDA further encode the margin information by using only neighboring samples to construct the intra-class and inter-class scatters. Experiments on several UCI datasets and two face databases demonstrate the effectiveness of the HMMDA methods. Bing Su 0001, Xiaoqing Ding, Changsong Liu, Ying Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Human Action Segmentation using 3D Fully Convolutional Network
Pei Yu, Jiang Wang 0001, Ying Wu 0001 |
BMVC | 3 |
| 2017 | Towards a Unified Compositional Model for Visual Pattern ModelingabstractCompositional models represent visual patterns as hierarchies of meaningful and reusable parts. They are attractive to vision modeling due to their ability to decompose complex patterns into simpler ones and resolve the lowlevel ambiguities in high-level image interpretations. However, current compositional models separate structure and part discovery from parameter estimation, which generally leads to suboptimal learning and fitting of the model. Moreover, the commonly adopted latent structural learning is not scalable for deep architectures. To address these difficult issues for compositional models, this paper quests for a unified framework for compositional pattern modeling, inference and learning. Represented by And-or graphs (AOGs), it jointly models the compositional structure, parts, features, and composition/sub-configuration relationships. We show that the inference algorithm of the proposed framework is equivalent to a feed-forward network. Thus, all the parameters can be learned efficiently via the highly-scalable back-propagation (BP) in an end-to-end fashion. We validate the model via the task of handwritten digit recognition. By visualizing the processes of bottom-up composition and top-down parsing, we show that our model is fully interpretable, being able to learn the hierarchical compositions from visual primitives to visual patterns at increasingly higher levels. We apply this new compositional model to natural scene character recognition and generic object detection. Experimental results have demonstrated its effectiveness. Wei Tang 0016, Pei Yu, Jiahuan Zhou, Ying Wu 0001 |
ICCV | 4 |
| 2017 | Efficient Online Local Metric Adaptation via Negative Samples for Person Re-identificationabstractMany existing person re-identification (PRID) methods typically attempt to train a faithful global metric offline to cover the enormous visual appearance variations, so as to directly use it online on various probes for identity matching. However, their need for a huge set of positive training pairs is very demanding in practice. In contrast to these methods, this paper advocates a different paradigm: part of the learning can be performed online but with nominal costs, so as to achieve online metric adaptation for different input probes. A major challenge here is that no positive training pairs are available for the probe anymore. By only exploiting easily-available negative samples, we propose a novel solution to achieve local metric adaptation effectively and efficiently. For each probe at the test time, it learns a strictly positive semi-definite dedicated local metric. Comparing to offline global metric learning, its computational cost is negligible. The insight of this new method is that the local hard negative samples can actually provide tight constraints to fine tune the metric locally. This new local metric adaptation method is generally applicable, as it can be used on top of any global metric to enhance its performance. In addition, this paper gives in-depth theoretical analysis and justification of the new method. We prove that our new method guarantees the reduction of the classification error asymptotically, and prove that it actually learns the optimal local metric to best approximate the asymptotic case by a finite number of training data. Extensive experiments and comparative studies on almost all major benchmarks (VIPeR, QMUL GRID, CUHK Campus, CUHK03 and Market-1501) have confirmed the effectiveness and superiority of our method. Jiahuan Zhou, Pei Yu, Wei Tang 0016, Ying Wu 0001 |
ICCV | 4 |
| 2017 | Deep Networks for Degraded Document Image Binarization through Pyramid ReconstructionabstractBinarization of document images is an important processing step for document images analysis and recognition. However, this problem is quite challenging in some cases because of the quality degradation of document images, such as varying illumination, complicated backgrounds, image noises due to ink spots, water stains or document creases. In this paper, we propose a framework based on deep convolutional neural-network (DCNN) for adaptive binarization of degraded document images. The basic idea of our method is to decompose a degraded document image into a spatial pyramid structure by using DCNN, with each layer at different scale. Then the foreground image is sequentially reconstructed from these layers in a coarse-to-fine manner by using deconvolutional network. Such kind of decomposition is quite beneficial, since multi-resolution supervision information can be directly introduced into network learning. We also define several loss functions about label consistency and foregrounds smoothing to further regularize the training of the network. Experimental results demonstrate the effectiveness of the proposed method. Gaofeng Meng, Kun Yuan 0003, Ying Wu 0001, Shiming Xiang, Chunhong Pan |
ICDAR | 3 |
| 2017 | Discriminative Transformation for Multi-Dimensional Temporal SequencesabstractFeature space transformation techniques have been widely studied for dimensionality reduction in vector-based feature space. However, these techniques are inapplicable to sequence data because the features in the same sequence are not independent. In this paper, we propose a method called max-min inter-sequence distance analysis (MMSDA) to transform features in sequences into a low-dimensional subspace such that different sequence classes are holistically separated. To utilize the temporal dependencies, MMSDA first aligns features in sequences from the same class to an adapted number of temporal states, and then, constructs the sequence class separability based on the statistics of these ordered states. To learn the transformation, MMSDA formulates the objective of maximizing the minimal pairwise separability in the latent subspace as a semi-definite programming problem and provides a new tractable and effective solution with theoretical proofs by constraints unfolding and pruning, convex relaxation, and within-class scatter compression. Extensive experiments on different tasks have demonstrated the effectiveness of MMSDA. Bing Su 0001, Xiaoqing Ding, Changsong Liu, Hao Wang 0005, Ying Wu 0001 |
IEEE Trans. Image Process. | 5 |
| 2017 | Unsupervised Hierarchical Dynamic Parsing and Encoding for Action RecognitionabstractGenerally, the evolution of an action is not uniform across the video, but exhibits quite complex rhythms and non-stationary dynamics. To model such non-uniform temporal dynamics, in this paper, we describe a novel hierarchical dynamic parsing and encoding method to capture both the locally smooth dynamics and globally drastic dynamic changes. It parses the dynamics of an action into different layers and encodes such multi-layer temporal information into a joint representation for action recognition. At the first layer, the action sequence is parsed in an unsupervised manner into several smooth-changing stages corresponding to different key poses or temporal structures by temporal clustering. The dynamics within each stage are encoded by mean-pooling or rank-pooling. At the second layer, the temporal information of the ordered dynamics extracted from the previous layer is encoded again by rank-pooling to form the overall representation. Extensive experiments on a gesture action data set (Chalearn Gesture) and three generic action data sets (Olympic Sports, Hollywood2, and UCF101) have demonstrated the effectiveness of the proposed method. Bing Su 0001, Jiahuan Zhou, Xiaoqing Ding, Ying Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Online Variable Coding Length Product Quantization for Fast Nearest Neighbor Search in Mobile RetrievalabstractQuantization methods are crucial for efficient nearest neighbor search in many applications such as image, music, or product search. As mobile devices are becoming increasingly more popular, the quantization methods on mobile devices are more important, because a large portion of the search queries are becoming performed on mobile devices. One important characteristic of the communication on mobile devices is the inherent unreliability of their communication channels. In order to adapt the quality changes of the communication channels, we need to change the coding length of the quantization accordingly. The existing quantization methods use fixed-length codebooks, and it is expensive to retrain another codebook with different coding length. In this paper, we propose a novel variable length product quantization framework that consists of a set of fast universal scalar quantizers. The framework is capable of producing variable length quantization without retraining the codebook. Each data vector is transformed into a new space to reduce the correlation across dimensions. A proper number of bits is allocated to represent the scalar component in each dimension according to the given coding length. For each component, we estimate its probability density function (PDF) and design an efficient universal scalar quantizer based on the PDF and the allocated bits. To reduce distortion, we learn a Gaussian mixture model for the data. The experimental results show that, compared to state-of-the-art product quantization methods, our approach can construct the codebooks online for variable coding lengths and achieve the comparable performance. Jin Li 0011, Xuguang Lan, Xiangwei Li, Jiang Wang 0001, Nanning Zheng 0001, Ying Wu 0001 |
IEEE Trans. Multim. | 6 |
| 2016 | Learning Reconstruction-Based Remote Gaze EstimationabstractIt is a challenging problem to accurately estimate gazes from low-resolution eye images that do not provide fine and detailed features for eyes. Existing methods attempt to establish the mapping between the visual appearance space to the gaze space. Different from the direct regression approach, the reconstruction-based approach represents appearance and gaze via local linear reconstruction in their own spaces. A common treatment is to use the same local reconstruction in the two spaces, i.e., the reconstruction weights in the appearance space are transferred to the gaze space for gaze reconstruction. However, this questionable treatment is taken for granted but has never been justified, leading to significant errors in gaze estimation. This paper is focused on the study of this fundamental issue. It shows that the distance metric in the appearance space needs to be adjusted, before the same reconstruction can be used. A novel method is proposed to learn the metric, such that the affinity structure of the appearance space under this new metric is as close as possible to the affinity structure of the gaze space under the normal Euclidean metric. Furthermore, the local affinity structure invariance is utilized to further regularize the solution to the reconstruction weights, so as to obtain a more robust and accurate solution. Effectiveness of the proposed method is validated and demonstrated through extensive experiments on different subjects. Pei Yu, Jiahuan Zhou, Ying Wu 0001 |
CVPR | 3 |
| 2016 | Hierarchical Dynamic Parsing and Encoding for Action Recognition
Bing Su 0001, Jiahuan Zhou, Xiaoqing Ding, Hao Wang 0005, Ying Wu 0001 |
ECCV (4) | 5 |
| 2016 | Finding the right exemplars for reconstructing single image super-resolutionabstractExemplar-based methods have shown their potential in synthesizing novel but visually plausible contents for image super-resolution (SR), by using the implicit knowledge conveyed by the exemplar database. In practice, however, it is common that unwanted artifacts and low quality results are produced due to the using of inappropriate exemplars. How are the “right” exemplars defined and identified? This fundamental issue has not be well addressed in these methods. This paper proposes a novel solution to this issue by learning a new distance metric in the LR space, such that affinity structure of the LR space under the new metric is as close to that of the HR space. Based on this learned best metric, appropriate exemplars can be identified. In addition, the proposed method is able to automatically determine the appropriate number of exemplars to use. Extensive experiments have shown that our method is able to handle regions with different properties and to obtain visually appealing super-resolution results with sharp details and smooth edges. Jiahuan Zhou, Ying Wu 0001 |
ICIP | 2 |
| 2015 | Heteroscedastic max-min distance analysisabstractMany discriminant analysis methods such as LDA and HLDA actually maximize the average pairwise distances between classes, which often causes the class separation problem. Max-min distance analysis (MMDA) addresses this problem by maximizing the minimum pairwise distance in the latent subspace, but it is developed under the homoscedastic assumption. This paper proposes Heteroscedastic MMDA (HMMDA) methods that explore the discriminative information in the difference of intra-class scatters for dimensionality reduction. WHMMDA maximizes the minimal pairwise Chenoff distance in the whitened space. OHMMDA incorporates this objective and the minimization of class compactness into a trace quotient formulation and imposes an orthogonal constraint to the final transformation, which can be solved by a bisection search algorithm. Two variants of OHMMDA are further proposed to encode the margin information. Experiments on several UCI Machine Learning datasets and the Yale Face database demonstrate the effectiveness of the proposed HMMDA methods. Bing Su 0001, Xiaoqing Ding, Changsong Liu, Ying Wu 0001 |
CVPR | 4 |
| 2015 | A novel unsupervised approach to discovering regions of interest in traffic images
Zhenyu An, Zhenwei Shi 0001, Ying Wu 0001, Changshui Zhang |
Pattern Recognit. | 3 |
| 2015 | Sparse Unmixing of Hyperspectral Data Using Spectral A Priori InformationabstractGiven a spectral library, sparse unmixing aims at finding the optimal subset of endmembers from it to model each pixel in the hyperspectral scene. However, sparse unmixing still remains a challenging task due to the usually high mutual coherence of the spectral library. In this paper, we exploit the spectral a priori information in the hyperspectral image to alleviate this difficulty. It assumes that some materials in the spectral library are known to exist in the scene. Such information can be obtained via field investigation or hyperspectral data analysis. Then, we propose a novel model to incorporate the spectral a priori information into sparse unmixing. Based on the alternating direction method of multipliers, we present a new algorithm, which is termed sparse unmixing using spectral a priori information (SUnSPI), to solve the model. Experimental results on both synthetic and real data demonstrate that the spectral a priori information is beneficial to sparse unmixing and that SUnSPI can exploit this information effectively to improve the abundance estimation. Wei Tang 0016, Zhenwei Shi 0001, Ying Wu 0001, Changshui Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2014 | Unifying Spatial and Attribute Selection for Distracter-Resilient TrackingabstractVisual distracters are detrimental and generally very difficult to handle in target tracking, because they generate false positive candidates for target matching. The resilience of region-based matching to the distracters depends not only on the matching metric, but also on the characteristics of the target region to be matched. The two tasks, i.e., learning the best metric and selecting the distracter-resilient target regions, actually correspond to the attribute selection and spatial selection processes in the human visual perception. This paper presents an initial attempt to unify the modeling of these two tasks for an effective solution, based on the introduction of a new quantity called Soft Visual Margin. As a function of both matching metric and spatial location, it measures the discrimination between the target and its spatial distracters, and characterizes the reliability of matching. Different from other formulations of margin, this new quantity is analytical and is insensitive to noisy data. This paper presents a novel method to jointly determine the best spatial location and the optimal metric. Based on that, a solid distracter-resilient region tracker is designed, and its effectiveness is validated and demonstrated through extensive experiments. Nan Jiang 0016, Ying Wu 0001 |
CVPR | 2 |
| 2014 | Cross-View Action Modeling, Learning, and RecognitionabstractExisting methods on video-based action recognition are generally view-dependent, i.e., performing recognition from the same views seen in the training data. We present a novel multiview spatio-temporal and-or graph (MST-AOG) representation for cross-view action recognition, i.e., the recognition is performed on the video from an unknown and unseen view. As a compositional model, MST-AOG compactly represents the hierarchical combinatorial structures of cross-view actions by explicitly modeling the geometry, appearance and motion variations. This paper proposes effective methods to learn the structure and parameters of MST-AOG. The inference based on MST-AOG enables action recognition from novel views. The training of MST-AOG takes advantage of the 3D human skeleton data obtained from Kinect cameras to avoid annotating enormous multi-view video frames, which is error-prone and time-consuming, but the recognition does not need 3D information and is based on 2D video input. A new Multiview Action3D dataset has been created and will be released. Extensive experiments have demonstrated that this new action representation significantly improves the accuracy and robustness for cross-view action recognition on 2D videos. Jiang Wang 0001, Xiaohan Nie, Yin Xia, Ying Wu 0001, Song-Chun Zhu |
CVPR | 4 |
| 2014 | Learning Fine-Grained Image Similarity with Deep RankingabstractLearning fine-grained image similarity is a challenging task. It needs to capture between-class and within-class image differences. This paper proposes a deep ranking model that employs deep learning techniques to learn similarity metric directly from images. It has higher learning capability than models based on hand-crafted features. A novel multiscale network structure has been developed to describe the images effectively. An efficient triplet sampling algorithm is also proposed to learn the model with distributed asynchronized stochastic gradient. Extensive experiments show that the proposed algorithm outperforms models based on hand-crafted visual features and deep classification models. Jiang Wang 0001, Yang Song 0009, Thomas K. Leung, Charles Rosenberg 0001, Jingbin Wang, James Philbin, Bo Chen 0019, Ying Wu 0001 |
CVPR | 8 |
| 2014 | Mining discriminative 3D Poselet for cross-view action recognitionabstractThis paper presents a novel approach to cross-view action recognition. Traditional cross-view action recognition methods typically rely on local appearance/motion features. In this paper, we take advantage of the recent developments of depth cameras to build a more discriminative cross-view action representation. In this representation, an action is characterized by the spatio-temporal configuration of 3D Poselets, which are discriminatively discovered with a novel Poselet mining algorithm and can be detected with view-invariant 3D Poselet detectors. The Kinect skeleton is employed to facilitate the 3D Poselet mining and 3D Poselet detectors learning, but the recognition is solely based on 2D video input. Extensive experiments have demonstrated that this new action representation significantly improves the accuracy and robustness for cross-view action recognition. Jiang Wang 0001, Xiaohan Nie, Yin Xia, Ying Wu 0001 |
WACV | 4 |
| 2014 | Automatic, fast, online calibration between depth and color cameras
Ilya Mikhelson, Philip Greggory Lee, Alan V. Sahakian, Ying Wu 0001, Aggelos K. Katsaggelos |
J. Vis. Commun. Image Represent. | 4 |
| 2014 | Spatially-Constrained Similarity Measurefor Large-Scale Object RetrievalabstractOne fundamental problem in object retrieval with the bag-of-words model is its lack of spatial information. Although various approaches are proposed to incorporate spatial constraints into the model, most of them are either too strict or too loose so that they are only effective in limited cases. In this paper, a new spatially-constrained similarity measure (SCSM) is proposed to handle object rotation, scaling, view point change and appearance deformation. The similarity measure can be efficiently calculated by a voting-based method using inverted files. During the retrieval process, object localization in the database images can also be simultaneously achieved using SCSM without post-processing. Furthermore, based on the retrieval and localization results of SCSM, we introduce a novel and robust re-ranking method with the k-nearest neighbors of the query for automatically refining the initial search results. Extensive performance evaluations on six public data sets show that SCSM significantly outperforms other spatial models including RANSAC-based spatial verification, while k-NN re-ranking outperforms most state-of-the-art approaches using query expansion. We also adapted SCSM for mobile product image search with an iterative algorithm to simultaneously extract the product instance from the mobile query image, identify the instance, and retrieve visually similar product images. Experiments on two product image search data sets show that our approach can robustly localize and extract the product in the query image, and hence drastically improve the retrieval accuracy over baseline methods. Xiaohui Shen, Zhe Lin 0001, Jonathan Brandt, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | Learning Actionlet Ensemble for 3D Human Action RecognitionabstractHuman action recognition is an important yet challenging task. Human actions usually involve human-object interactions, highly articulated motions, high intra-class variations, and complicated temporal structures. The recently developed commodity depth sensors open up new possibilities of dealing with this problem by providing 3D depth data of the scene. This information not only facilitates a rather powerful human motion capturing technique, but also makes it possible to efficiently model human-object interactions and intra-class variations. In this paper, we propose to characterize the human actions with a novel actionlet ensemble model, which represents the interaction of a subset of human joints. The proposed model is robust to noise, invariant to translational and temporal misalignment, and capable of characterizing both the human motion and the human-object interactions. We evaluate the proposed approach on three challenging action recognition datasets captured by Kinect devices, a multiview action recognition dataset captured with Kinect device, and a dataset captured by a motion capture system. The experimental evaluations show that the proposed approach achieves superior performance to the state-of-the-art algorithms. Jiang Wang 0001, Zicheng Liu 0001, Ying Wu 0001, Junsong Yuan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Regularized Simultaneous Forward-Backward Greedy Algorithm for Sparse Unmixing of Hyperspectral DataabstractSparse unmixing assumes that each observed signature of a hyperspectral image is a linear combination of only a few spectra (endmembers) in an available spectral library. It then estimates the fractional abundances of these endmembers in the scene. The sparse unmixing problem still remains a great difficulty due to the usually high correlation of the spectral library. Under such circumstances, this paper presents a novel algorithm termed as the regularized simultaneous forward-backward greedy algorithm (RSFoBa) for sparse unmixing of hyperspectral data. The RSFoBa has low computational complexity of getting an approximate solution for the l0problem directly and can exploit the joint sparsity among all the pixels in the hyperspectral data. In addition, the combination of the forward greedy step and the backward greedy step makes the RSFoBa more stable and less likely to be trapped into the local optimum than the conventional greedy algorithms. Furthermore, when updating the solution in each iteration, a regularizer that enforces the spatial-contextual coherence within the hyperspectral image is considered to make the algorithm more effective. We also show that the sublibrary obtained by the RSFoBa can serve as input for any other sparse unmixing algorithms to make them more accurate and time efficient. Experimental results on both synthetic and real data demonstrate the effectiveness of the proposed algorithm. Wei Tang 0016, Zhenwei Shi 0001, Ying Wu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2014 | Context-Aware Discovery of Visual Co-Occurrence PatternsabstractOnce an image is decomposed into a number of visual primitives, e.g., local interest points or regions, it is of great interests to discover meaningful visual patterns from them. Conventional clustering of visual primitives, however, usually ignores the spatial and feature structure among them, thus cannot discover high-level visual patterns of complex structure. To overcome this problem, we propose to consider spatial and feature contexts among visual primitives for pattern discovery. By discovering spatial co-occurrence patterns among visual primitives and feature co-occurrence patterns among different types of features, our method can better address the ambiguities of clustering visual primitives. We formulate the pattern discovery problem as a regularized k-means clustering where spatial and feature contexts are served as constraints to improve the pattern discovery results. A novel self-learning procedure is proposed to utilize the discovered spatial or feature patterns to gradually refine the clustering result. Our self-learning procedure is guaranteed to converge and experiments on real images validate the effectiveness of our method. Hongxing Wang 0001, Junsong Yuan 0001, Ying Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2013 | Large Displacement Optical Flow from Nearest Neighbor FieldsabstractWe present an optical flow algorithm for large displacement motions. Most existing optical flow methods use the standard coarse-to-fine framework to deal with large displacement motions which has intrinsic limitations. Instead, we formulate the motion estimation problem as a motion segmentation problem. We use approximate nearest neighbor fields to compute an initial motion field and use a robust algorithm to compute a set of similarity transformations as the motion candidates for segmentation. To account for deviations from similarity transformations, we add local deformations in the segmentation process. We also observe that small objects can be better recovered using translations as the motion candidates. We fuse the motion results obtained under similarity transformations and under translations together before a final refinement. Experimental validation shows that our method can successfully handle large displacement motions. Although we particularly focus on large displacement motions in this work, we make no sacrifice in terms of overall performance. In particular, our method ranks at the top of the Middlebury benchmark. Zhuoyuan Chen, Hailin Jin, Zhe Lin 0001, Scott Cohen, Ying Wu 0001 |
CVPR | 5 |
| 2013 | Detecting and Aligning Faces by Image RetrievalabstractDetecting faces in uncontrolled environments continues to be a challenge to traditional face detection methods due to the large variation in facial appearances, as well as occlusion and clutter. In order to overcome these challenges, we present a novel and robust exemplar-based face detector that integrates image retrieval and discriminative learning. A large database of faces with bounding rectangles and facial landmark locations is collected, and simple discriminative classifiers are learned from each of them. A voting-based method is then proposed to let these classifiers cast votes on the test image through an efficient image retrieval technique. As a result, faces can be very efficiently detected by selecting the modes from the voting maps, without resorting to exhaustive sliding window-style scanning. Moreover, due to the exemplar-based framework, our approach can detect faces under challenging conditions without explicitly modeling their variations. Evaluation on two public benchmark datasets shows that our new face detection approach is accurate and efficient, and achieves the state-of-the-art performance. We further propose to use image retrieval for face validation (in order to remove false positives) and for face alignment/landmark localization. The same methodology can also be easily generalized to other face-related tasks, such as attribute recognition, as well as general object detection. Xiaohui Shen, Zhe Lin 0001, Jonathan Brandt, Ying Wu 0001 |
CVPR | 4 |
| 2013 | Robust Dictionary Learning by Error Source DecompositionabstractSparsity models have recently shown great promise in many vision tasks. Using a learned dictionary in sparsity models can in general outperform predefined bases in clean data. In practice, both training and testing data may be corrupted and contain noises and outliers. Although recent studies attempted to cope with corrupted data and achieved encouraging results in testing phase, how to handle corruption in training phase still remains a very difficult problem. In contrast to most existing methods that learn the dictionary from clean data, this paper is targeted at handling corruptions and outliers in training data for dictionary learning. We propose a general method to decompose the reconstructive residual into two components: a non-sparse component for small universal noises and a sparse component for large outliers, respectively. In addition, further analysis reveals the connection between our approach and the ``partial'' dictionary learning approach, updating only part of the prototypes (or informative code words) with remaining (or noisy code words) fixed. Experiments on synthetic data as well as real applications have shown satisfactory performance of this new robust dictionary learning approach. Zhuoyuan Chen, Ying Wu 0001 |
ICCV | 2 |
| 2013 | Learning Maximum Margin Temporal Warping for Action RecognitionabstractTemporal misalignment and duration variation in video actions largely influence the performance of action recognition, but it is very difficult to specify effective temporal alignment on action sequences. To address this challenge, this paper proposes a novel discriminative learning-based temporal alignment method, called maximum margin temporal warping (MMTW), to align two action sequences and measure their matching score. Based on the latent structure SVM formulation, the proposed MMTW method is able to learn a phantom action template to represent an action class for maximum discrimination against other classes. The recognition of this action class is based on the associated learned alignment of the input action. Extensive experiments on five benchmark datasets have demonstrated that this MMTW model is able to significantly promote the accuracy and robustness of action recognition under temporal misalignment and variations. Jiang Wang 0001, Ying Wu 0001 |
ICCV | 2 |
| 2013 | Single image super-resolution based on space structure learning
Heng Su, Nan Jiang 0016, Ying Wu 0001, Jie Zhou 0001 |
Pattern Recognit. Lett. | 3 |
| 2013 | What Are We Tracking: A Unified Approach of Tracking and RecognitionabstractTracking is essentially a matching problem. While traditional tracking methods mostly focus on low-level image correspondences between frames, we argue that high-level semantic correspondences are indispensable to make tracking more reliable. Based on that, a unified approach of low-level object tracking and high-level recognition is proposed for single object tracking, in which the target category is actively recognized during tracking. High-level offline models corresponding to the recognized category are then adaptively selected and combined with low-level online tracking models so as to achieve better tracking performance. Extensive experimental results show that our approach outperforms state-of-the-art online models in many challenging tracking scenarios such as drastic view change, scale change, background clutter, and morphable objects. Jialue Fan, Xiaohui Shen, Ying Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | Decomposing and regularizing sparse/non-sparse components for motion field estimationabstractRegularizing motion field is critical to achieve accurate estimation of the motion field. As the motion field may include discontinuity (e.g., at the motion boundaries), traditional smoothness regularization may not work well. Among many approaches to handling motion discontinuity, recent attempts pursued a sparse representation of the motion field for regularization, and achieved quite encouraging results. However, statistics show that these methods tend to over-sparsify the motion field, and thus confronted by the non-sparse noise in practice. In this paper, we propose to decompose the motion field into sparse and non-sparse components for the motion boundaries and small universal noises, respectively. This separation approach regularizes these two sources differently. We propose a novel and efficient optimization algorithm to solve this problem. In addition, our study reveals the in-depth connection between this noise separation approach and the influence function approach in robust statistics. We validate and evaluate our new approach on the Middlebury benchmark, and have achieved outstanding testing performance. Zhuoyuan Chen, Jiang Wang 0001, Ying Wu 0001 |
CVPR | 3 |
| 2012 | Order determination and sparsity-regularized metric learning adaptive visual trackingabstractRecent attempts of integrating metric learning in visual tracking have produced encouraging results. Instead of using fixed and pre-specified metric in visual appearance matching, these methods are able to learn and adjust the metric adaptively by finding the best projection of the feature space. Such learned metric is by design the best to discriminate the target of interest and its distracters from the background. However, an important issue remained unaddressed is how we can determine the optimal dimensionality of the projection to achieve best discrimination. Using inappropriate dimensions for the projection is likely to result in larger classification error, or higher computational costs and over-fitting. This paper presents a novel solution to this structural order determination problem, by introducing sparsity regularization for metric learning (or SRML). This regularization leads to the lowest possible dimensionality of the projection and thus determining the best order. This can actually be viewed as the minimum description length regularization in metric learning. The experiments validate this new approach on standard benchmark datasets, and demonstrate its effectiveness in visual tracking applications. Nan Jiang 0016, Wenyu Liu 0001, Ying Wu 0001 |
CVPR | 3 |
| 2012 | Object retrieval and localization with spatially-constrained similarity measure and k-NN re-rankingabstractOne fundamental problem in object retrieval with the bag-of-visual words (BoW) model is its lack of spatial information. Although various approaches are proposed to incorporate spatial constraints into the BoW model, most of them are either too strict or too loose so that they are only effective in limited cases. We propose a new spatially-constrained similarity measure (SCSM) to handle object rotation, scaling, view point change and appearance deformation. The similarity measure can be efficiently calculated by a voting-based method using inverted files. Object retrieval and localization are then simultaneously achieved without post-processing. Furthermore, we introduce a novel and robust re-ranking method with the k-nearest neighbors of the query for automatically refining the initial search results. Extensive performance evaluations on six public datasets show that SCSM significantly outperforms other spatial models, while k-NN re-ranking outperforms most state-of-the-art approaches using query expansion. Xiaohui Shen, Zhe Lin 0001, Jonathan Brandt, Shai Avidan, Ying Wu 0001 |
CVPR | 5 |
| 2012 | A unified approach to salient object detection via low rank matrix recoveryabstractSalient object detection is not a pure low-level, bottom-up process. Higher-level knowledge is important even for task-independent image saliency. We propose a unified model to incorporate traditional low-level features with higher-level guidance to detect salient objects. In our model, an image is represented as a low-rank matrix plus sparse noises in a certain feature space, where the non-salient regions (or background) can be explained by the low-rank matrix, and the salient regions are indicated by the sparse noises. To ensure the validity of this model, a linear transform for the feature space is introduced and needs to be learned. Given an image, its low-level saliency is then extracted by identifying those sparse noises when recovering the low-rank matrix. Furthermore, higher-level knowledge is fused to compose a prior map, and is treated as a prior term in the objective function to improve the performance. Extensive experiments show that our model can comfortably achieves comparable performance to the existing methods even without the help from high-level knowledge. The integration of top-down priors further improves the performance and achieves the state-of-the-art. Moreover, the proposed model can be considered as a prototype framework not only for general salient object detection, but also for potential task-dependent saliency applications. Xiaohui Shen, Ying Wu 0001 |
CVPR | 2 |
| 2012 | Mining actionlet ensemble for action recognition with depth camerasabstractHuman action recognition is an important yet challenging task. The recently developed commodity depth sensors open up new possibilities of dealing with this problem but also present some unique challenges. The depth maps captured by the depth cameras are very noisy and the 3D positions of the tracked joints may be completely wrong if serious occlusions occur, which increases the intra-class variations in the actions. In this paper, an actionlet ensemble model is learnt to represent each action and to capture the intra-class variance. In addition, novel features that are suitable for depth data are proposed. They are robust to noise, invariant to translational and temporal misalignments, and capable of characterizing both the human motion and the human-object interactions. The proposed approach is evaluated on two challenging action recognition datasets captured by commodity depth cameras, and another dataset captured by a MoCap system. The experimental evaluations show that the proposed approach achieves superior performance to the state of the art algorithms. Jiang Wang 0001, Zicheng Liu 0001, Ying Wu 0001, Junsong Yuan 0001 |
CVPR | 3 |
| 2012 | Mobile Product Image Search by Automatic Query Object Extraction
Xiaohui Shen, Zhe Lin 0001, Jonathan Brandt, Ying Wu 0001 |
ECCV (4) | 4 |
| 2012 | Robust 3D Action Recognition with Random Occupancy Patterns
Jiang Wang 0001, Zicheng Liu 0001, Jan Chorowski, Zhuoyuan Chen, Ying Wu 0001 |
ECCV (2) | 5 |
| 2012 | Dynamic hand gesture recognition: An exemplar-based approach from motion divergence fields
Xiaohui Shen, Gang Hua 0001, Lance Williams, Ying Wu 0001 |
Image Vis. Comput. | 4 |
| 2012 | Scribble Tracker: A Matting-Based Approach for Robust TrackingabstractModel updating is a critical problem in tracking. Inaccurate extraction of the foreground and background information in model adaptation would cause the model to drift and degrade the tracking performance. The most direct yet difficult solution to the drift problem is to obtain accurate boundaries of the target. We approach such a solution by proposing a novel model adaptation framework based on the combination of matting and tracking. In our framework, coarse tracking results automatically provide sufficient and accurate scribbles for matting, which makes matting applicable in a tracking system. Meanwhile, accurate boundaries of the target can be obtained from matting results even when the target has large deformation. An effective model combining short-term features and long-term appearances is further constructed and successfully updated based on such accurate boundaries. The model can successfully handle occlusion by explicit inference. Extensive experiments show that our adaptation scheme largely avoids model drift and significantly outperforms other discriminative tracking models. Jialue Fan, Xiaohui Shen, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Discriminative Metric Preservation for Tracking Low-Resolution TargetsabstractTracking low-resolution (LR) targets is a practical yet quite challenging problem in real video analysis applications. Lack of discriminative details in the visual appearance of the LR target leads to the matching ambiguity, which confronts most existing tracking methods. Although artificially enhancing the video resolution by superresolution (SR) techniques before analyzing might be an option, the high demand of computational cost can hardly meet the requirements of the tracking scenario. This paper presents a novel solution to track LR targets without explicitly performing SR. This new approach is based on discriminative metric preservation that preserves the data affinity structure in the high-resolution (HR) feature space for effective and efficient matching of LR images. In addition, we substantialize this new approach in a solid case study of differential tracking under metric preservation and derive a closed-form solution to motion estimation for LR video. In addition, this paper extends the basic linear metric preservation method to a more powerful nonlinear kernel metric preservation method. Such a solution to LR target tracking is discriminative, robust, and efficient. Extensive experiments validate the entrustments and effectiveness of the proposed approach and demonstrate the improved performance of the proposed method in tracking LR targets. Nan Jiang 0016, Heng Su, Wenyu Liu 0001, Ying Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2012 | Spatially Adaptive Block-Based Super-ResolutionabstractSuper-resolution technology provides an effective way to increase image resolution by incorporating additional information from successive input images or training samples. Various super-resolution algorithms have been proposed based on different assumptions, and their relative performances can differ in regions of different characteristics within a single image. Based on this observation, an adaptive algorithm is proposed in this paper to integrate a higher level image classification task and a lower level super-resolution process, in which we incorporate reconstruction-based super-resolution algorithms, single-image enhancement, and image/video classification into a single comprehensive framework. The target high-resolution image plane is divided into adaptive-sized blocks, and different suitable super-resolution algorithms are automatically selected for the blocks. Then, a deblocking process is applied to reduce block edge artifacts. A new benchmark is also utilized to measure the performance of super-resolution algorithms. Experimental results with real-life videos indicate encouraging improvements with our method. Heng Su, Ying Wu 0001, Daniel Tretter, Jie Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | Super-Resolution Without Dense FlowabstractSuper-resolution is a widely applied technique that improves the resolution of input images by software methods. Most conventional reconstruction-based super-resolution algorithms assume accurate dense optical flow fields between the input frames, and their performance degrades rapidly when the motion estimation result is not accurate enough. However, optical flow estimation is usually difficult, particularly when complicated motion is presented in real-world videos. In this paper, we explore a new way to solve this problem by using sparse feature point correspondences between the input images. The feature point correspondences, which are obtained by matching a set of feature points, are usually precise and much more robust than dense optical flow fields. This is because the feature points represent well-selected significant locations in the image, and performing matching on the feature point set is usually very accurate. In order to utilize the sparse correspondences in conventional super-resolution, we extract an adaptive support region with a reliable local flow field from each corresponding feature point pair. The normalized prior is also proposed to increase the visual consistency of the reconstructed result. Extensive experiments on real data were carried out, and results show that the proposed algorithm produces high-resolution images with better quality, particularly in the presence of large-scale or complicated motion fields. Heng Su, Ying Wu 0001, Jie Zhou 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | Discovering Thematic Objects in Image Collections and VideosabstractGiven a collection of images or a short video sequence, we define a thematic object as the key object that frequently appears and is the representative of the visual contents. Successful discovery of the thematic object is helpful for object search and tagging, video summarization and understanding, etc. However, this task is challenging because 1) there lacks a priori knowledge of the thematic objects, such as their shapes, scales, locations, and times of re-occurrences, and 2) the thematic object of interest can be under severe variations in appearances due to viewpoint and lighting condition changes, scale variations, etc. Instead of using a top-down generative model to discover thematic visual patterns, we propose a novel bottom-up approach to gradually prune uncommon local visual primitives and recover the thematic objects. A multilayer candidate pruning procedure is designed to accelerate the image data mining process. Our solution can efficiently locate thematic objects of various sizes and can tolerate large appearance variations of the same thematic object. Experiments on challenging image and video data sets and comparisons with existing methods validate the effectiveness of our method. Junsong Yuan 0001, Gangqiang Zhao, Yun Fu 0001, Zhu Li 0001, Aggelos K. Katsaggelos, Ying Wu 0001 |
IEEE Trans. Image Process. | 6 |
| 2012 | Mining Visual Collocation Patterns via Self-Supervised Subspace LearningabstractTraditional text data mining techniques are not directly applicable to image data which contain spatial information and are characterized by high-dimensional visual features. It is not a trivial task to discover meaningful visual patterns from images because the content variations and spatial dependence in visual data greatly challenge most existing data mining methods. This paper presents a novel approach to coping with these difficulties for mining visual collocation patterns. Specifically, the novelty of this work lies in the following new contributions: 1) a principled solution to the discovery of visual collocation patterns based on frequent itemset mining and 2) a self-supervised subspace learning method to refine the visual codebook by feeding back discovered patterns via subspace learning. The experimental results show that our method can discover semantically meaningful patterns efficiently and effectively. Junsong Yuan 0001, Ying Wu 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2011 | Tracking low resolution objects by metric preservationabstractTracking low resolution (LR) targets is a practical yet quite challenging problem in real applications. The loss of discriminative details in the visual appearance of the L-R targets confronts most existing visual tracking methods. Although the resolution of the LR video inputs may be enhanced by super resolution (SR) techniques, the large computational cost for high-quality SR does not make it an attractive option. This paper presents a novel solution to track LR targets without performing explicit SR. This new approach is based on discriminative metric preservation that preserves the structure in the high resolution feature space for LR matching. In addition, we integrate metric preservation with differential tracking to derive a closed-form solution to motion estimation for LR video. Extensive experiments have demonstrated the effectiveness and efficiency of the proposed approach. Nan Jiang 0016, Wenyu Liu 0001, Heng Su, Ying Wu 0001 |
CVPR | 4 |
| 2011 | Adaptive and discriminative metric differential trackingabstractMatching the visual appearances of the target over consecutive image frames is the most critical issue in video-based object tracking. Choosing an appropriate distance metric for matching determines its accuracy and robustness, and significantly influences the tracking performance. This paper presents a new tracking approach that incorporates adaptive metric into differential tracking method. This new approach automatically learns an optimal distance metric for more accurate matching, and obtains a closed-form analytical solution to motion estimation and differential tracking. Extensive experiments validate the effectiveness of adaptive metric, and demonstrate the improved performance of the proposed new tracking method. Nan Jiang 0016, Wenyu Liu 0001, Ying Wu 0001 |
CVPR | 3 |
| 2011 | Nonlocal mattingabstractThis work attempts to considerably reduce the amount of user effort in the natural image matting problem. The key observation is that the nonlocal principle, introduced to denoise images, can be successfully applied to the alpha matte to obtain sparsity in matte representation, and therefore dramatically reduce the number of pixels a user needs to manually label. We show how to avoid making the user provide redundant and unnecessary input, develop a method for clustering the image pixels for the user to label, and a method to perform high-quality matte extraction. We show that this algorithm is therefore faster, easier, and higher quality than state of the art methods. Philip Greggory Lee, Ying Wu 0001 |
CVPR | 2 |
| 2011 | Action recognition with multiscale spatio-temporal contextsabstractThe popular bag of words approach for action recognition is based on the classifying quantized local features density. This approach focuses excessively on the local features but discards all information about the interactions among them. Local features themselves may not be discriminative enough, but combined with their contexts, they can be very useful for the recognition of some actions. In this paper, we present a novel representation that captures contextual interactions between interest points, based on the density of all features observed in each interest point's mutliscale spatio-temporal contextual domain. We demonstrate that augmenting local features with our contextual feature significantly improves the recognition performance. Jiang Wang 0001, Zhuoyuan Chen, Ying Wu 0001 |
CVPR | 3 |
| 2011 | Mining discriminative co-occurrence patterns for visual recognitionabstractThe co-occurrence pattern, a combination of binary or local features, is more discriminative than individual features and has shown its advantages in object, scene, and action recognition. We discuss two types of co-occurrence patterns that are complementary to each other, the conjunction (AND) and disjunction (OR) of binary features. The necessary condition of identifying discriminative co-occurrence patterns is firstly provided. Then we propose a novel data mining method to efficiently discover the optimal co-occurrence pattern with minimum empirical error, despite the noisy training dataset. This mining procedure of AND and OR patterns is readily integrated to boosting, which improves the generalization ability over the conventional boosting decision trees and boosting decision stumps. Our versatile experiments on object, scene, and action categorization validate the advantages of the discovered discriminative co-occurrence patterns. Junsong Yuan 0001, Ming Yang 0007, Ying Wu 0001 |
CVPR | 3 |
| 2011 | Motion divergence fields for dynamic hand gesture recognitionabstractAlthough it is in general difficult to track articulated hand motion, exemplar-based approaches provide a robust solution for hand gesture recognition. Presumably, a rich set of dynamic hand gestures are needed for a meaningful recognition system. How to build the visual representation for the motion patterns is the key for scalable recognition. We propose a novel representation based on the divergence map of the gestural motion field, which transforms motion patterns into spatial patterns. Given the motion divergence maps, we leverage modern image feature detectors to extract salient spatial patterns, such as Maximum Stable Extremal Regions (MSER). A local descriptor is extracted from each region to capture the local motion pattern. The descriptors from gesture exemplars are subsequently indexed using a pre-trained vocabulary tree. New gestures are then matched efficiently with the database gestures with a TF-IDF scheme. Our extensive experiments on a large hand gesture database with 10 categories and 1050 video samples validate the efficacy of the extracted motion patterns for gesture recognition. The proposed approach achieves an overall recognition rate of 97.62%, while the average recognition time is only 34.53 ms. Xiaohui Shen, Gang Hua 0001, Lance Williams, Ying Wu 0001 |
FG | 4 |
| 2011 | Adaptive incremental video super-resolution with temporal consistencyabstractVideo super-resolution can be generally divided into two categories: incremental video super-resolution and simultaneous video super-resolution. Incremental video super-resolution algorithms are usually faster, but their results cannot be guaranteed to be visually consistent to the human vision system. An adaptive incremental video super-resolution framework with the temporal consistency constraint is proposed in this paper. The temporal consistency among the video frames is enforced by imposing the similarity between the adjacent reconstructed HR frames. The variances of the potential functions, which affect the weights of the different terms in the utility function, are adaptively determined so that the algorithm is robust to various motion and image content situations. Some considerations, such as the incremental motion estimation, are also incorporated to improve the efficiency of the algorithm, which makes the proposed algorithm near-realtime. The experimental results show that the proposed algorithm can generate HR video with high quality while saving the computational time as well. Heng Su, Ying Wu 0001, Jie Zhou 0001 |
ICIP | 2 |
| 2011 | Contextual saliencyabstractMatching local salient points is limited in some computer vision problems (e.g., wide baseline point matching and its applications), since local features vary dramatically under large view changes. In this paper, we propose a new definition on salient points which includes context information around the given point. The proposed contextual salient points are extracted based on the geometry structure of the manifold embedded in the image feature space. Experimental results show the benefit of contextual salient points over local salient points. Jialue Fan, Ying Wu 0001 |
VCIP | 2 |
| 2011 | Learning spatio-temporal dependency of local patches for complex motion segmentation
Jiang Xu 0002, Junsong Yuan 0001, Ying Wu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2011 | Discriminative Video Pattern Search for Efficient Action DetectionabstractActions are spatiotemporal patterns. Similar to the sliding window-based object detection, action detection finds the reoccurrences of such spatiotemporal patterns through pattern matching, by handling cluttered and dynamic backgrounds and other types of action variations. We address two critical issues in pattern matching-based action detection: 1) the intrapattern variations in actions, and 2) the computational efficiency in performing action pattern search in cluttered scenes. First, we propose a discriminative pattern matching criterion for action classification, called naive Bayes mutual information maximization (NBMIM). Each action is characterized by a collection of spatiotemporal invariant features and we match it with an action class by measuring the mutual information between them. Based on this matching criterion, action detection is to localize a subvolume in the volumetric video space that has the maximum mutual information toward a specific action class. A novel spatiotemporal branch-and-bound (STBB) search algorithm is designed to efficiently find the optimal solution. Our proposed action detection method does not rely on the results of human detection, tracking, or background subtraction. It can handle action variations such as performing speed and style variations as well as scale changes well. It is also insensitive to dynamic and cluttered backgrounds and even to partial occlusions. The cross-data set experiments on action detection, including KTH, CMU action data sets, and another new MSR action data set, demonstrate the effectiveness and efficiency of the proposed multiclass multiple-instance action detection method. Junsong Yuan 0001, Zicheng Liu 0001, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Special Issue on Video Analysis on Resource-Limited SystemsabstractThe 17 papers in this special issue focus on resource-limited systems. Rama Chellappa, Andrea Cavallaro, Ying Wu 0001, Caifeng Shan, Yun Fu 0001, Kari Pulli |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Learning Adaptive Metric for Robust Visual TrackingabstractMatching the visual appearances of the target over consecutive image frames is the most critical issue in video-based object tracking. Choosing an appropriate distance metric for matching determines its accuracy and robustness, and thus significantly influences the tracking performance. Most existing tracking methods employ fixed pre-specified distance metrics. However, this simple treatment is problematic and limited in practice, because a pre-specified metric does not likely to guarantee the closest match to be the true target of interest. This paper presents a new tracking approach that incorporates adaptive metric learning into the framework of visual object tracking. Collecting a set of supervised training samples on-the-fly in the observed video, this new approach automatically learns the optimal distance metric for more accurate matching. The design of the learned metric ensures that the closest match is very likely to be the true target of interest based on the supervised training. Such a learned metric is discriminative and adaptive. This paper substantializes this new approach in a solid case study of adaptive-metric differential tracking, and obtains a closed-form analytical solution to motion estimation and visual tracking. Moreover, this paper extends the basic linear distance metric learning method to a more powerful nonlinear kernel metric learning method. Extensive experiments validate the effectiveness of the proposed approach, and demonstrate the improved performance of the proposed new tracking method. Nan Jiang 0016, Wenyu Liu 0001, Ying Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2010 | Sparsity model for robust optical flow estimation at motion discontinuitiesabstractThis paper introduces a new sparsity prior to the estimation of dense flow fields. Based on this new prior, a complex flow field with motion discontinuities can be accurately estimated by finding the sparsest representation of the flow field in certain domains. In addition, a stronger additional sparsity constraint on the flow gradients is incorporated into the model to cope with the measurement noises. Robust estimation techniques are also employed to identify the outliers and to refine the results. This new sparsity model can accurately and reliably estimate the entire dense flow field from a small portion of measurements when other measurements are corrupted by noise. Experiments show that our method significantly outperforms traditional methods that are based on global or piecewise smoothness priors. Xiaohui Shen, Ying Wu 0001 |
CVPR | 2 |
| 2010 | Closed-Loop Adaptation for Robust Tracking
Jialue Fan, Xiaohui Shen, Ying Wu 0001 |
ECCV (1) | 3 |
| 2010 | Discriminative Spatial Attention for Robust Tracking
Jialue Fan, Ying Wu 0001, Shengyang Dai |
ECCV (1) | 2 |
| 2010 | Automatic video-based analysis of animal behaviorsabstractVision-based animal behavior analysis is a critical and interesting problem for both biologists and computer vision scientists. In this paper, an automatic system for detecting behaviors of fruit flies is presented. Firstly, we propose an ellipse model to fit the contours of fruit flies, which efficiently detects fruit flies in a single frame. Then we associate the detection results together to form the trajectories. An AdaBoost classifier is used to analyze special behaviors of flies. The experiments show that our system can robustly track fruit flies and detect the fly behaviors with high recall rate in real time. This system has been adopted to aid biologists for research purposes. Jialue Fan, Nan Jiang 0016, Ying Wu 0001 |
ICIP | 3 |
| 2010 | L1 mattingabstractNatural image matting continues to play a large role in a wide variety of applications. As an ill-posed problem, matting is a very difficult to solve due to its underconstrained nature. Current approaches can require a lot of user input, restrict themselves to a sparse subset of the image, and often make assumptions that are unlikely to hold. In this paper, we pose a way to better satisfy smoothness assumptions of some of these methods utilizing the nonlinear median filter which arises naturally from the L1norm. The median has the property that it tends to smooth the foreground and background of the image while leaving any edges relatively unaltered. We then show that such an image is often more suitable as input than the original image, even when user interaction is minimal, suggesting that our method is more amenable to automatic matting. Philip Greggory Lee, Ying Wu 0001 |
ICIP | 2 |
| 2010 | Exploiting sparsity in dense optical flowabstractIn this paper we validated that the dense optical flow field is sparse in certain frequency domains, while the flow gradient field is also sparse in image domain. Based on this sparsity prior, the optical flow estimation problem is casted as sparse signal recovery from highly shorted measurements. By minimizing its l1-norm in frequency domain and gradient domain, the model can accurately estimate the dense flow field without other assumptions. Outliers are further identified and removed in the flow denoising process to improve the results. Experiments show that our method significantly outperforms traditional methods based on global or piecewise smoothness priors. Moreover, it can well handle the complexity incurred by motion discontinuities. Xiaohui Shen, Ying Wu 0001 |
ICIP | 2 |
| 2010 | Part-based initialization for hand trackingabstractInitializing hand/finger articulation for tracking is a very challenging problem, mainly because hand articulation is complicated and it has a large number of degrees of freedom. Most existing algorithms initialize tracking manually, or use a nearest-neighbor search with restricting the number of possible hand gestures. This paper presents a new solution to this problem by increasing the dimensionality but taking advantage of the sparseness. The basic idea is to divide the set of phalange joint angles into many overlapping subsets. As each subset has a much smaller number of joint angles, it is much easier to design a smaller-scale articulation estimator. The estimation of the whole hand is done by the collaboration of a network of dependent smaller-scale estimators. This paper describes a novel way of designing the smaller-scale estimators as well as a principled way of fusing the estimates. A tracking system is also shown by using this initialization technique. Jiang Xu 0002, Ying Wu 0001, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2010 | Bipolar groupingabstractMost affinity-based grouping methods only model the inclusive relation among the data. When the data set contains a significant amount of noise data that should not be included in any clusters, these methods are likely to lead to undesired results. To address this issue, this paper presents a new approach called bipolar grouping that is targeted on extracting the groups from the data while excluding the noise. This new approach incorporates both inclusive and exclusive relations among data, and a fixed-point procedure is proposed to find the stable groups. Its effectiveness and general applicability are demonstrated in two applications, including discovering common objects in images and tracking targets in clutter. Jiang Xu 0002, Junsong Yuan 0001, Ying Wu 0001 |
ICME | 3 |
| 2010 | Interactive visual object search through mutual information maximizationabstractSearching for small objects (e.g., logos) in images is a critical yet challenging problem. It becomes more difficult when target objects differ significantly from the query object due to changes in scale, viewpoint or style, not to mention partial occlusion or cluttered backgrounds. With the goal to retrieve and accurately locate the small object in the images, we formulate the object search as the problem of finding subimages with the largest mutual information toward the query object. Each image is characterized by a collection of local features. Instead of only using the query object for matching, we propose a discriminative matching using both positive and negative queries to obtain the mutual information score. The user can verify the retrieved subimages and improve the search results incrementally. Our experiments on a challenging logo database of 10,000 images highlight the effectiveness of this approach. Jingjing Meng, Junsong Yuan 0001, Yuning Jiang 0001, Nitya Narasimhan, Venu Vasudevan, Ying Wu 0001 |
ACM Multimedia | 6 |
| 2010 | AdaBoost-based face detection for embedded systems
Ming Yang 0007, Jim E. Crenshaw, Bruce Augustine, Russell Mareachen, Ying Wu 0001 |
Comput. Vis. Image Underst. | 5 |
| 2010 | Mining Compositional Features From GPS and Visual Cues for Event Recognition in Photo CollectionsabstractAs digital cameras with Global Positioning System (GPS) capability become available and people geotag their photos using other means, it is of great interest to annotate semantic events (e.g., hiking, skiing, party) characterized by a collection of geotagged photos with timestamps and GPS information at the capture. We address this emerging event classification problem by mining informative features derived from image contents and spatio-temporal traces of GPS coordinates that characterize the underlying movement patterns of various event types, both based on the entire collection as opposed to individual photos. Considering that events are better described by the co-occurrence of objects and scenes, we bundle primitive features such as color and texture histograms or GPS features to form the discriminative compositional feature. A data mining method is proposed to efficiently discover discriminative compositional features of small classification errors. A theoretical analysis is also presented to guide the selection of the data mining parameters. Upon compositional feature mining, we perform the multiclass AdaBoost to further integrate the mined compositional features. Finally, the GPS and visual modalities are united through a confidence-based fusion. Based on a dataset of more than 3000 geotagged images, experimental results have shown the synergy of all of the components in our proposed approach to event classification. Junsong Yuan 0001, Jiebo Luo 0001, Ying Wu 0001 |
IEEE Trans. Multim. | 3 |
| 2010 | Human tracking using convolutional neural networksabstractIn this paper, we treat tracking as a learning problem of estimating the location and the scale of an object given its previous location, scale, as well as current and previous image frames. Given a set of examples, we train convolutional neural networks (CNNs) to perform the above estimation task. Different from other learning methods, the CNNs learn both spatial and temporal features jointly from image pairs of two adjacent frames. We introduce multiple path ways in CNN to better fuse local and global information. A creative shift-variant CNN architecture is designed so as to alleviate the drift problem when the distracting objects are similar to the target in cluttered environment. Furthermore, we employ CNNs to estimate the scale through the accurate localization of some key points. These techniques are object-independent so that the proposed method can be applied to track other types of object. The capability of the tracker of handling complex situations is demonstrated in many testing sequences. Jialue Fan, Wei Xu 0007, Ying Wu 0001, Yihong Gong |
IEEE Trans. Neural Networks | 3 |
| 2010 | Connectivity-Based Skeleton Extraction in Wireless Sensor NetworksabstractMany sensor network applications are tightly coupled with the geometric environment where the sensor nodes are deployed. The topological skeleton extraction for the topology has shown great impact on the performance of such services as location, routing, and path planning in wireless sensor networks. Nonetheless, current studies focus on using skeleton extraction for various applications in wireless sensor networks. How to achieve a better skeleton extraction has not been thoroughly investigated. There are studies on skeleton extraction from the computer vision community; their centralized algorithms for continuous space, however, are not immediately applicable for the discrete and distributed wireless sensor networks. In this paper, we present a novel Connectivity-bAsed Skeleton Extraction (CASE) algorithm to compute skeleton graph that is robust to noise, and accurate in preservation of the original topology. In addition, CASE is distributed as no centralized operation is required, and is scalable as both its time complexity and its message complexity are linearly proportional to the network size. The skeleton graph is extracted by partitioning the boundary of the sensor network to identify the skeleton points, then generating the skeleton arcs, connecting these arcs, and finally refining the coarse skeleton graph. We believe that CASE has broad applications and present a skeleton-assisted segmentation algorithm as an example. Our evaluation shows that CASE is able to extract a well-connected skeleton graph in the presence of significant noise and shape variations, and outperforms the state-of-the-art algorithms. Hongbo Jiang 0001, Wenping Liu 0001, Dan Wang 0002, Chen Tian 0001, Xiang Bai, Xue (Steve) Liu, Ying Wu 0001, Wenyu Liu 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2009 | Removing partial blur in a single imageabstractRemoving image partial blur is of great practical importance. However, as existing recovery techniques usually assume a one-layer clear image model, they can not characterize the actual generation process of partial blurs. In this paper, a two-layer image model is investigated. Based on the study of partial blur generation process, a novel recovery technique is proposed for a single input image. Both foreground and background layers are recovered simultaneously with the help of the matting technique, powerful image prior models, and user assistance. The effectiveness of the proposed approach is demonstrated by extensive experiments on image recovery and synthesis on real data. Shengyang Dai, Ying Wu 0001 |
CVPR | 2 |
| 2009 | Contextual flowabstractMatching based on local brightness is quite limited, because small changes on local appearance invalidate the constancy in brightness. The root of this limitation is its treatment regardless of the information from the spatial contexts. This papers leaps from brightness constancy to context constancy, and thus from optical flow to contextual flow. It presents a new approach that incorporates contexts to constrain motion estimation for target tracking. In this approach, one individual spatial context of a given pixel is represented by the posterior density of the associated feature class in its contextual domain. Each individual context gives a linear contextual flow constraint to the motion, so that the motion can be estimated in an over-determined contextual system. Based on this contextual flow model, this paper presents a new and powerful target tracking method that integrates the processes of salient contextual point selection, robust contextual matching, and dynamic context selection. Extensive experiment results show the effectiveness of the proposed approach. Ying Wu 0001, Jialue Fan |
CVPR | 1 |
| 2009 | Discriminative subvolume search for efficient action detectionabstractActions are spatio-temporal patterns which can be characterized by collections of spatio-temporal invariant features. Detection of actions is to find the re-occurrences (e.g. through pattern matching) of such spatio-temporal patterns. This paper addresses two critical issues in pattern matching-based action detection: (1) efficiency of pattern search in 3D videos and (2) tolerance of intra-pattern variations of actions. Our contributions are two-fold. First, we propose a discriminative pattern matching called naive-Bayes based mutual information maximization (NBMIM) for multi-class action categorization. It improves the state-of-the-art results on standard KTH dataset. Second, a novel search algorithm is proposed to locate the optimal subvolume in the 3D video space for efficient action detection. Our method is purely data-driven and does not rely on object detection, tracking or background subtraction. It can well handle the intra-pattern variations of actions such as scale and speed variations, and is insensitive to dynamic and clutter backgrounds and even partial occlusions. The experiments on versatile datasets including KTH and CMU action datasets demonstrate the effectiveness and efficiency of our method. Junsong Yuan 0001, Zicheng Liu 0001, Ying Wu 0001 |
CVPR | 3 |
| 2009 | Multimodal partial estimates fusionabstractFusing partial estimates is a critical and common problem in many computer vision tasks such as part-based detection and tracking. It generally becomes complicated and intractable when there are a large number of multimodal partial estimates, and thus it is desirable to find an effective and scalable fusion method to integrate these partial estimates. This paper presents a novel and effective approach to fusing multimodal partial estimates in a principled way. In this new approach, fusion is related to a computational geometry problem of finding the minimum-volume orthotope, and an effective and scalable branch and bound search algorithm is designed to obtain the global optimal solution. Experiments on tracking articulated objects and occluded objects show the effectiveness of the proposed approach. Jiang Xu 0002, Junsong Yuan 0001, Ying Wu 0001 |
ICCV | 3 |
| 2009 | Detecting contextual anomalies of crowd motion in surveillance videoabstractMany works have been proposed on detecting individual anomalies in crowd scenes, i.e., human behaviors anomalous with respect to the rest of the behaviors. In this paper, we introduce a new concept of contextual anomaly into the field of crowd analysis, i.e., the behaviors themselves are normal but they are anomalous in a specific context. Our system follows an unsupervised approach. It automatically discovers important contextual information from the crowd video and detects the blobs corresponding to contextually anomalous behaviors. Our experiments show that the approach works well in detecting contextual anomalies from crowd video with different motion contexts. Ying Wu 0001, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2009 | CASE: Connectivity-Based Skeleton Extraction in Wireless Sensor NetworksabstractMany sensor network applications are tightly coupled with the geometric environment where the sensor nodes are deployed. The topological skeleton extraction has shown great impact on the performance of such services as location, routing, and path planning in sensor networks. Nonetheless, current studies focus on using skeleton extraction for various applications in sensor networks. How to achieve a better skeleton extraction has not been thoroughly investigated. There are studies on skeleton extraction from the computer vision community; their centralized algorithms for continuous space, however, is not immediately applicable for the discrete and distributed sensor networks. In this paper we present CASE: a novel connectivity-based skeleton extraction algorithm to compute skeleton graph that is robust to noise, and accurate in preservation of the original topology. In addition, no centralized operation is required. The skeleton graph is extracted by partitioning the boundary of the sensor network to identify the skeleton points, then generating the skeleton arcs, connecting these arcs, and finally refining the coarse skeleton graph. Our evaluation shows that CASE is able to extract a well-connected skeleton graph in the presence of significant noise and shape variations, and outperforms state-of-the-art algorithms. Hongbo Jiang 0001, Wenping Liu 0001, Dan Wang 0002, Chen Tian 0001, Xiang Bai, Xue (Steve) Liu, Ying Wu 0001, Wenyu Liu 0001 |
INFOCOM | 7 |
| 2009 | Context-Aware Visual TrackingabstractEnormous uncertainties in unconstrained environments lead to a fundamental dilemma that many tracking algorithms have to face in practice: Tracking has to be computationally efficient, but verifying whether or not the tracker is following the true target tends to be demanding, especially when the background is cluttered and/or when occlusion occurs. Due to the lack of a good solution to this problem, many existing methods tend to be either effective but computationally intensive by using sophisticated image observation models or efficient but vulnerable to false alarms. This greatly challenges long-duration robust tracking. This paper presents a novel solution to this dilemma by considering the context of the tracking scene. Specifically, we integrate into the tracking process a set of auxiliary objects that are automatically discovered in the video on the fly by data mining. Auxiliary objects have three properties, at least in a short time interval: 1) persistent co-occurrence with the target, 2) consistent motion correlation to the target, and 3) easy to track. Regarding these auxiliary objects as the context of the target, the collaborative tracking of these auxiliary objects leads to efficient computation as well as strong verification. Our extensive experiments have exhibited exciting performance in very challenging real-world testing cases. Ming Yang 0007, Ying Wu 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | SoftCuts: A Soft Edge Smoothness Prior for Color Image Super-ResolutionabstractDesigning effective image priors is of great interest to image super-resolution (SR), which is a severely under-determined problem. An edge smoothness prior is favored since it is able to suppress the jagged edge artifact effectively. However, for soft image edges with gradual intensity transitions, it is generally difficult to obtain analytical forms for evaluating their smoothness. This paper characterizes soft edge smoothness based on a novel SoftCuts metric by generalizing the Geocuts method . The proposed soft edge smoothness measure can approximate the average length of all level lines in an intensity image. Thus, the total length of all level lines can be minimized effectively by integrating this new form of prior. In addition, this paper presents a novel combination of this soft edge smoothness prior and the alpha matting technique for color image SR, by adaptively normalizing image edges according to their alpha-channel description. This leads to the adaptive SoftCuts algorithm, which represents a unified treatment of edges with different contrasts and scales. Experimental results are presented which demonstrate the effectiveness of the proposed method. Shengyang Dai, Wei Xu 0007, Ying Wu 0001, Yihong Gong, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 2009 | A Dynamic Hierarchical Clustering Method for Trajectory-Based Unusual Video Event DetectionabstractThe proposed unusual video event detection method is based on unsupervised clustering of object trajectories, which are modeled by hidden Markov models (HMM). The novelty of the method includes a dynamic hierarchical process incorporated in the trajectory clustering algorithm to prevent model overfitting and a 2-depth greedy search strategy for efficient clustering. Ying Wu 0001, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 2 |
| 2009 | Tracking Nonstationary Visual Appearances by Data-Driven AdaptationabstractWithout any prior about the target, the appearance is usually the only cue available in visual tracking. However, in general, the appearances are often nonstationary which may ruin the predefined visual measurements and often lead to tracking failure in practice. Thus, a natural solution is to adapt the observation model to the nonstationary appearances. However, this idea is threatened by the risk of adaptation drift that originates in its ill-posed nature, unless good data-driven constraints are imposed. Different from most existing adaptation schemes, we enforce three novel constraints for the optimal adaptation: 1) negative data, 2) bottom-up pair-wise data constraints, and 3) adaptation dynamics. Substantializing the general adaptation problem as a subspace adaptation problem, this paper presents a closed-form solution as well as a practical iterative algorithm for subspace tracking. Extensive experiments have demonstrated that the proposed approach can largely alleviate adaptation drift and achieve better tracking results for a large variety of nonstationary scenes. Ming Yang 0007, Zhimin Fan 0002, Jialue Fan, Ying Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2008 | Motion from blurabstractMotion blur retains some information about motion, based on which motion may be recovered from blurred images. This is a difficult problem, as the situations of motion blur can be quite complicated, such as they may be space-variant, nonlinear, and local. This paper addresses a very challenging problem: can we recover motion blindly from a single motion-blurred image? A major contribution of this paper is a new finding of an elegant motion blur constraint. Exhibiting a very similar mathematical form as the optical flow constraint, this linear constraint applies locally to pixels in the image. Therefore, a number of challenging problems can be addressed, including estimating global affine motion blur, estimating global rotational motion blur, estimating and segmenting multiple motion blur, and estimating nonparametric motion blur field. Extensive experiments on blur estimation and image deblurring on both synthesized and real data demonstrate the accuracy and general applicability of the proposed approach. Shengyang Dai, Ying Wu 0001 |
CVPR | 2 |
| 2008 | Vital sign estimation from passive thermal videoabstractConventional wired detection of vital signs limits the use of these important physiological parameters by many applications, such as airport health screening, elder care, and workplace preventive care. In this paper, we explore contact-free heart rate and respiratory rate detection through measuring infrared light modulation emitted near superficial blood vessels or a nasal area respectively. To deal with complications caused by subjects’ movements, facial expressions, and partial occlusions of the skin, we propose a novel algorithm based on contour segmentation and tracking, clustering of informative pixels, and dominant frequency component estimation. The proposed method achieves robust subject regions-of-interest alignment and motion compensation in infrared video with low SNR. It relaxes some strong assumptions used in previous work and substantially improves on previously reported performance. Preliminary experiments on heart rate estimation for 20 subjects and respiratory rate estimation for 8 subjects exhibit promising results. Ming Yang 0007, Qiong Liu 0003, Thea Turner, Ying Wu 0001 |
CVPR | 4 |
| 2008 | Granularity and elasticity adaptation in visual trackingabstractThe observation models in tracking algorithms are critical to both tracking performance and applicable scenarios but are often simplified to focus on fixed level of certain target properties such as appearances and structures. In this paper, we propose a unified tracking paradigm in which targets are represented by Markov random fields of interest regions and introduce a new way to adapt observation models by automatically tuning the feature granularity and model elasticity, i.e. the abstraction level of features and the model’s degree of flexibility to tolerate deformations. Specifically, we employ a multi-scale scheme to extract features from interest regions and adjust the parameters of the potential functions of the MRF model to maximize the likelihoods of tracking results. Experiments demonstrate the method can estimate translation, scaling and rotation and deal with deformation, partial occlusions, and camouflage objects within this unified framework. Ming Yang 0007, Ying Wu 0001 |
CVPR | 2 |
| 2008 | Distributed data association and filtering for multiple target trackingabstractThis paper presents a novel distributed framework for multi-target tracking with an efficient data association computation. A decentralized representation of trackers’ motion and association variables is adopted. Considering the interleaved nature of data association and tracker filtering, the multi-target tracking is formulated as a missing data problem, and the solution is found by the proposed variational EM algorithm. We analytically show that 1) the posteriori distributions of trackers’ motions (the real interests in terms of tracking applications) can be effectively computed in the E-step of the EM iterations, and 2) the solution of trackers’ association variables can be pursued under a derived graph-based discrete optimization formulation, thus efficiently estimated in the M-step by the recently emerging graph optimization algorithms. The proposed approach is very general such that sophisticated data association priori and likelihood function can be easily incorporated. This general framework is tested with both simulation data and real world surveillance video. The reported qualitative and quantitative studies verify the effectiveness and low computational cost of the algorithm. Ting Yu 0003, Ying Wu 0001, Nils Krahnstoever, Peter H. Tu |
CVPR | 2 |
| 2008 | Mining compositional features for boostingabstractThe selection of weak classifiers is critical to the success of boosting techniques. Poor weak classifiers do not perform better than random guess, thus cannot help decrease the training error during the boosting process. Therefore, when constructing the weak classifier pool, we prefer the quality rather than the quantity of the weak classifiers. In this paper, we present a data mining-driven approach to discovering compositional features from a given and possibly small feature pool. Compared with individual features (e.g. weak decision stumps) which are of limited discriminative ability, the mined compositional features have guaranteed power in terms of the descriptive and discriminative abilities, as well as bounded training error. To cope with the combinatorial cost of discovering compositional features, we apply data mining methods (frequent itemset mining) to efficiently find qualified compositional features of any possible order. These weak classifiers are further combined through a multi-class AdaBoost method for final multi-class classification. Experiments on a challenging 10-class event recognition problem show that boosting compositional features can lead to faster decrease of training error and significantly higher accuracy compared to conventional boosting decision stumps. Junsong Yuan 0001, Jiebo Luo 0001, Ying Wu 0001 |
CVPR | 3 |
| 2008 | Context-aware clusteringabstractMost existing methods of semi-supervised clustering introduce supervision from outside, e.g., manually label some data samples or introduce constrains into clustering results. This paper studies an interesting problem: can the supervision come from inside, i.e., the unsupervised training data themselves? If the data samples are not independent, we can capture the contextual information reflecting the dependency among the data samples, and use it as supervision to improve the clustering. This is called context-aware clustering. The investigation is substantialized on two scenarios of (1) clustering primitive visual features (e.g., SIFT features) with help of spatial contexts, and (2) clustering ‘0’–‘9’ hand written digits with help of contextual patterns among different types of features. Our context-aware clustering can be well formulated in a closed-form, where the contextual information serves as a regularization term to balance the data fidelity in original feature space and the influences of contextual patterns. A nested-EM algorithm is proposed to obtain an efficient solution, which proves to converge. By exploring the dependent structure of the data samples, this method is completely unsupervised, as no outside supervision is introduced. Junsong Yuan 0001, Ying Wu 0001 |
CVPR | 2 |
| 2008 | Image spam hunterabstractSpammers are constantly creating sophisticated new weapons in their arms race with anti-spam technology, the latest of which is image-based spam. The newest image-based spam uses simple image processing technologies to vary the content of individual messages, e.g. by changing foreground colors, backgrounds, font types, or even rotating and adding artifacts to the images. Thus, they pose great challenges to conventional spam filters. In this paper, we propose a system using a probabilistic boosting tree to determine whether an incoming image is a spam or not based on global image features, i.e. color and gradient orientation histograms. The system identifies spam without the need for OCR and is robust in the face of the kinds of variation found in current spam images. Evaluation results show the system correctly classifies 90% of spam images while mislabeling only 0.86% of non-spam images as spam. Yan Gao 0003, Ming Yang 0007, Xiaonan Zhao, Bryan Pardo, Ying Wu 0001, Thrasyvoulos N. Pappas, Alok N. Choudhary |
ICASSP | 5 |
| 2008 | Abnormal event detection based on trajectory clustering by 2-depth greedy searchabstractClustering-based approaches for abnormal video event detection have been proven to be effective in the recent literature. Based on the framework proposed in our previous work [1], we have developed in this paper a new strategy for unsupervised trajectory clustering. More specifically, an information-based trajectory dissimilarity measure is proposed, based on the Bayesian information criterion (BIC). In order to minimize BIC, the agglomerative hierarchical clustering is applied using a 2-depth greedy search process. This strategy achieves better clustering results compared to the traditional 1-depth greedy search. The increased computational complexity is addressed with several bounds on the trajectory dissimilarity. Ying Wu 0001, Aggelos K. Katsaggelos |
ICASSP | 2 |
| 2008 | Estimating space-variant motion blur without deblurringabstractIdentifying space-variant motion blurs is a very challenging task in blind blur identification research. This paper describes a novel method towards blind identification without deblurring. Based on the image gradients in the α-channel component of a blurred color image, an elegant α-motion blur constraint is proposed, which is a linear constraint for local motion blur parameters. It makes possible efficient blind identification of space-variant and even nonlinear motion blurs and the estimation of motion blur fields, without deblurring. Shengyang Dai, Ying Wu 0001 |
ICIP | 2 |
| 2008 | A bi-subspace model for robust visual trackingabstractThe changes of the target’s visual appearance often lead to tracking failure in practice. Hence, trackers need to be adaptive to non-stationary appearances to achieve robust visual tracking. However, the risk of adaptation drift is common in most existing adaptation schemes. This paper describes a bi-subspace model that stipulates the interactions of two different visual cues. The visual appearance of the target is represented by two interactive subspaces, each of which corresponds to a particular cue. The adaption of the subspaces is through the interaction of the two cues, which leads to robust tracking performance. Extensive experiments show that the proposed approach can largely alleviate adaptation drift and obtain better tracking results. Jialue Fan, Ming Yang 0007, Ying Wu 0001 |
ICIP | 3 |
| 2008 | Locality Versus Globality: Query-Driven Localized Linear Models for Facial Image ComputingabstractConventional subspace learning or recent feature extraction methods consider globality as the key criterion to design discriminative algorithms for image classification. We demonstrate in this paper that applying the local manner in sample space, feature space, and learning space via linear subspace learning can sufficiently boost the discriminating power, as measured by discriminating power coefficient (DPC). The proposed solution achieves good classification accuracy gains and shows computationally efficient. Particularly, we approximate the global nonlinearity through a multimodal localized piecewise subspace learning framework, in which three locality criteria can work individually or jointly for any new subspace learning algorithm design. It turns out that most existing subspace learning methods can be unified in such a common framework embodying either the global or local learning manner. On the other hand, we address the problem of numerical difficulty in the large-size pattern classification case, where many local variations cannot be adequately handled by a single global model. By localizing the modeling, the classification error rate estimation is also localized and thus it appears to be more robust and flexible for the model selection among different model candidates. As a new algorithm design based on the proposed framework, the query-driven locally adaptive (QDLA) mixture-of-experts model for robust face recognition and head pose estimation is presented. Experiments demonstrate the local approach to be effective, robust, and fast for large size, multiclass, and multivariance data sets. Yun Fu 0001, Zhu Li 0001, Junsong Yuan 0001, Ying Wu 0001, Thomas S. Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2008 | Mining Recurring Events Through Forest GrowingabstractRecurring events are short temporal patterns that consist of multiple instances in the target database. Without anya prioriknowledge of the recurring events, in terms of their lengths, temporal locations, the total number of such events, and possible variations, it is a challenging problem to discover them because of the enormous computational cost involved in analyzing huge databases and the difficulty in accommodating all the possible variations without even knowing the target. Junsong Yuan 0001, Jingjing Meng, Ying Wu 0001, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | Soft Edge Smoothness Prior for Alpha Channel Super ResolutionabstractEffective image prior is necessary for image super resolution, due to its severely under-determined nature. Although the edge smoothness prior can be effective, it is generally difficult to have analytical forms to evaluate the edge smoothness, especially for soft edges that exhibit gradual intensity transitions. This paper finds the connection between the soft edge smoothness and a soft cut metric on an image grid by generalizing the Geocuts method (Y. Boykov and V. Kolmogorov, 2003), and proves that the soft edge smoothness measure approximates the average length of all level lines in an intensity image. This new finding not only leads to an analytical characterization of the soft edge smoothness prior, but also gives an intuitive geometric explanation. Regularizing the super resolution problem by this new form of prior can simultaneously minimize the length of all level lines, and thus resulting in visually appealing results. In addition, this paper presents a novel combination of this soft edge smoothness prior and the alpha matting technique for color image super resolution, by normalizing edge segments with their alpha channel description, to achieve a unified treatment of edges with different contrast and scale. Shengyang Dai, Wei Xu 0007, Ying Wu 0001, Yihong Gong |
CVPR | 4 |
| 2007 | Detector EnsembleabstractComponent-based detection methods have demonstrated their promise by integrating a set of part-detectors to deal with large appearance variations of the target. However, an essential and critical issue, i.e., how to handle the imperfectness of part-detectors in the integration, is not well addressed in the literature. This paper proposes a detector ensemble model that consists of a set of substructure-detectors, each of which is composed of several part-detectors. Two important issues are studied both in theory and in practice, (1) finding an optimal detector ensemble, and (2) detecting targets based on an ensemble. Based on some theoretical analysis, a new model selection strategy is proposed to learn an optimal detector ensemble that has a minimum number of false positives and satisfies the design requirement on the capacity of tolerating missing parts. In addition, this paper also links ensemble-based detection to the inference in Markov random field, and shows that the target detection can be done by a max-product belief propagation algorithm. Shengyang Dai, Ming Yang 0007, Ying Wu 0001, Aggelos K. Katsaggelos |
CVPR | 3 |
| 2007 | Spatial selection for attentional visual trackingabstractLong-duration tracking of general targets is quite challenging for computer vision, because in practice target may undergo large uncertainties in its visual appearance and the unconstrained environments may be cluttered and distractive, although tracking has never been a challenge to the human visual system. Psychological and cognitive findings indicate that the human perception is attentional and selective, and both early attentional selection that may be innate and late attentional selection that may be learned are necessary for human visual tracking. This paper proposes a new visual tracking approach by reflecting some aspects of spatial selective attention, and presents a novel attentional visual tracking (AVT) algorithm. In AVT, the early selection process extracts a pool of attentional regions (ARs) that are defined as the salient image regions which have good localization properties, and the late selection process dynamically identifies a subset of discriminative attentional regions (D-ARs) through a discriminative learning on the historical data on the fly. The computationally demanding process of matching of the AR pool is done in an efficient and innovative way by using the idea in the locality-sensitive hashing (LSH) technique. The proposed AVT algorithm is general, robust and computationally efficient, as shown in extensive experiments on a large variety of real-world video. Ming Yang 0007, Junsong Yuan 0001, Ying Wu 0001 |
CVPR | 3 |
| 2007 | Discovery of Collocation Patterns: from Visual Words to Visual PhrasesabstractA visual word lexicon can be constructed by clustering primitive visual features, and a visual object can be described by a set of visual words. Such a "bag-of-words" representation has led to many significant results in various vision tasks including object recognition and categorization. However, in practice, the clustering of primitive visual features tends to result in synonymous visual words that over-represent visual patterns, as well as polysemous visual words that bring large uncertainties and ambiguities in the representation. This paper aims at generating a higher-level lexicon, i.e. visual phrase lexicon, where a visual phrase is a meaningful spatially co-occurrent pattern of visual words. This higher-level lexicon is much less ambiguous than the lower-level one. The contributions of this paper include: (1) a fast and principled solution to the discovery of significant spatial co-occurrent patterns using frequent itemset mining; (2) a pattern summarization method that deals with the compositional uncertainties in visual phrases; and (3) a top-down refinement scheme of the visual word lexicon by feeding back discovered phrases to tune the similarity measure through metric learning. Junsong Yuan 0001, Ying Wu 0001, Ming Yang 0007 |
CVPR | 2 |
| 2007 | False Positive Reduction in Lung GGO Nodule Detection with 3D Volume Shape DescriptorabstractLung nodule detection, especially ground glass opacity (GGO) detection, in helical computed tomography (CT) images is a challenging computer-aided detection (CAD) task due to the enormous variances in nodules' volumes, shapes, appearances, and the structures nearby. Most of the detection algorithms employ some efficient candidate generation (CG) algorithms to spot the suspicious volumes with high sensitivity at the cost of low specificity, e.g. tens even hundreds of false positives per volume. This paper proposes a learning based method to reduce the number of false positives given by CG based on a new general 3D volume shape descriptor. The 3D volume shape descriptor is constructed by concatenating spatial histograms of gradient orientations, which is robust to large variabilities in intensity levels, shapes, and appearances. The proposed method achieves promising performance on a difficult mixture lung nodule dataset with average 81% detection rate and 4.3 false positives per volume. Ming Yang 0007, Senthil Periaswamy, Ying Wu 0001 |
ICASSP (1) | 3 |
| 2007 | Game-Theoretic Multiple Target TrackingabstractVideo-based multiple target tracking (MTT) is a challenging task when similar targets are present in close vicinity. Because their visual observations are mixed and difficult to segment, their motions have to be estimated jointly. Most existing approaches perform this joint motion estimation in a centralized fashion and involve searching a rather high dimensional space, and thus leading to quite complicated joint trackers. This paper brings a new view to MTT from a game-theoretic perspective, bridging the joint motion estimation and the Nash equilibrium of a game. Instead of designing a centralized tracker, MTT is decentralized and a set of individual trackers is used, each of which tries to maximize its visual evidence for explaining its motion as well as generates interferences to others. Modelling this competition behavior, a special game is designed so that the difficult joint motion estimation is achieved at the Nash Equilibrium of this game where no individual tracker has incentives to change its motion estimate. This paper substantializes this novel idea in a solid case study where individual trackers are kernel-based trackers. An efficient best response updating procedure is designed to find the Nash equilibrium. The powerfulness of this game-theoretic MTT is shown by promising results on difficult real videos. Ming Yang 0007, Ting Yu 0003, Ying Wu 0001 |
ICCV | 3 |
| 2007 | Spatial Random Partition for Common Visual Pattern DiscoveryabstractAutomatically discovering common visual patterns from a collection of images is an interesting but yet challenging task, in part because it is computationally prohibiting. Although representing images as visual documents based on discrete visual words offers advantages in computation, the performance of these word-based methods largely depends on the quality of the visual word dictionary. This paper presents a novel approach base on spatial random partition and fast word-free image matching. Represented as a set of continuous visual primitives, each image is randomly partitioned many times to form a pool of subimages. Each subimage is queried and matched against the pool, and then common patterns can be localized by aggregating the set of matched subimages. The asymptotic property and the complexity of the proposed method are given in this paper, along with many real experiments. Both theoretical studies and experiment results show its advantages. Junsong Yuan 0001, Ying Wu 0001 |
ICCV | 2 |
| 2007 | Query-Driven Locally Adaptive Fisher Faces and Expert-Model for Face RecognitionabstractWe present a novel expert-model of Query-Driven Locally Adaptive (QDLA) Fisher faces for robust face recognition. For each query face, the proposed method first fits local Fisher models with different appearances. A hybrid expert model then integrates these local models and combines the classification results based on the estimated error rate for each local model. This approach addresses the large size recognition problem, where many local variations can not be adequately handled by a single global model in a single appearance space. To speed up the query process, Locality Sensitive Hash (LSH) is applied for fast nearest neighbor search. Experiments demonstrate the approach to be effective, robust, and fast for large size, multi-class, and multi-variance data sets. Yun Fu 0001, Junsong Yuan 0001, Zhu Li 0001, Thomas S. Huang, Ying Wu 0001 |
ICIP (1) | 5 |
| 2007 | Abnormal Event Detection from Surveillance Video by Dynamic Hierarchical ClusteringabstractThe clustering-based approach for detecting abnormalities in surveillance video requires the appropriate definition of similarity between events. The HMM-based similarity defined previously falls short in handling the overfitting problem. We propose in this paper a multi-sample-based similarity measure, where HMM training and distance measuring are based on multiple samples. These multiple training data are acquired by a novel dynamic hierarchical clustering (DHC) method. By iteratively reclassifying and retraining the data groups at different clustering levels, the initial training and clustering errors due to overfitting will be sequentially corrected in later steps. Experimental results on real surveillance video show an improvement of the proposed method over a baseline method that uses single-sample-based similarity measure and spectral clustering. Ying Wu 0001, Aggelos K. Katsaggelos |
ICIP (5) | 2 |
| 2007 | Mining Auxiliary Objects for Tracking by Multibody GroupingabstractOn-line discovery of some auxiliary objects to verify the tracking results is a novel approach to achieving robust tracking by balancing the need for strong verification and computational efficiency. However, the applicability and effectiveness of this approach highly depend on how to reliably validate the motion correlation between the target and the auxiliary objects so as to estimate the motion model. In this paper, we extend the algorithm of mining auxiliary objects for tracking by incorporating multibody grouping to detect the motion correlation and estimate the motion model, which imposes more general motion correlation constraints. The proposed method discovers the auxiliary objects that exhibit strong affine motion correlation and estimates the closed-form affine models. The proposed tracking algorithm shows good performance in real-world test sequences. Ming Yang 0007, Ying Wu 0001, Shihong Lao |
ICIP (3) | 2 |
| 2007 | Common Spatial Pattern Discovery by Efficient Candidate PruningabstractAutomatically discovering common visual patterns in images is very challenging due to the uncertainties in the visual appearances of such spatial patterns and the enormous computational cost involved in exploring the huge solution space. Instead of performing exhaustive search on all possible candidates of such spatial patterns at various locations and scales, this paper presents a novel and very efficient algorithm for discovering common visual patterns by designing a provably correct and computationally efficient pruning procedure that has a quadratic complexity. This new approach is able to efficiently search a set of images for unknown visual patterns that exhibit large appearance variations because of rotation, scale changes, slight view changes, color variations and partial occlusions. Junsong Yuan 0001, Zhu Li 0001, Yun Fu 0001, Ying Wu 0001, Thomas S. Huang |
ICIP (1) | 4 |
| 2007 | Bilateral Back-Projection for Single Image Super ResolutionabstractIn this paper, a novel algorithm for single image super resolution is proposed. Back-projection [1] can minimize the reconstruction error with an efficient iterative procedure. Although it can produce visually appealing result, this method suffers from the chessboard effect and ringing effect, especially along strong edges. The underlining reason is that there is no edge guidance in the error correction process. Bilateral filtering can achieve edge-preserving image smoothing by adding the extra information from the feature domain. The basic idea is to do the smoothing on the pixels which are nearby both in space domain and in feature domain. The proposed bilateral back-projection algorithm strives to integrate the bilateral filtering into the back-projection method. In our approach, the back-projection process can be guided by the edge information to avoid across-edge smoothing, thus the chessboard effect and ringing effect along image edges are removed. Promising results can be obtained by the proposed bilateral back-projection method efficiently. Shengyang Dai, Ying Wu 0001, Yihong Gong |
ICME | 3 |
| 2007 | Query Driven Localized Linear Discriminant Models for Head Pose EstimationabstractHead pose appearances under the pan and tilt variations span a high dimensional manifold that has complex structures and local variations. For pose estimation purpose, we need to discover the subspace structure of the manifold and learn discriminative subspaces/metrics for head pose recognition. The performance of the head pose estimation is heavily dependent on the accuracy of structure learnt and the discriminating power of the metric. In this work we develop a query point driven, localized linear subspace learning method that approximates the non-linearity of the head pose manifold structure with piece-wise linear discriminating subspaces/metrics. Simulation results demonstrate the effectiveness of the proposed solution in both accuracy and computational efficiency. Zhu Li 0001, Yun Fu 0001, Junsong Yuan 0001, Thomas S. Huang, Ying Wu 0001 |
ICME | 5 |
| 2007 | From frequent itemsets to semantically meaningful visual patternsabstractData mining techniques that are successful in transaction and text data may not be simply applied to image data that contain high-dimensional features and have spatial structures. It is not a trivial task to discover meaningful visual patterns in image databases, because the content variations and spatial dependency in the visual data greatly challenge most existing methods. This paper presents a novel approach to coping with these difficulties for mining meaningful visual patterns. Specifically, the novelty of this work lies in the following new contributions: (1) a principled solution to the discovery of meaningful itemsets based on frequent itemset mining; (2) a self-supervised clustering scheme of the high-dimensional visual features by feeding back discovered patterns to tune the similarity measure through metric learning; and (3) a pattern summarization method that deals with the measurement noises brought by the image data. The experimental results in the real images show that our method can discover semantically meaningful patterns efficiently and effectively. Junsong Yuan 0001, Ying Wu 0001, Ming Yang 0007 |
KDD | 2 |
| 2007 | Mining repetitive clips through finding continuous pathsabstractAutomatically discovering repetitive clips from large video database is a challenging problem due to the enormous computational cost involved in exploring the huge solution space. Without any a priori knowledge of the contents, lengths and total number of the repetitive clips, we need to discover all of them in the video database. To address the large computational cost, we propose a novel method which translates repetitive clip mining to the continuous path finding problem in a matching trellis, where sequence matching can be accelerated by taking advantage of the temporal redundancies in the videos. By applying the locality sensitive hashing (LSH) for efficient similarity query and the proposed continuous path finding algorithm, our method is of only quadratic complexity of the database size. Experiments conducted on a 10.5-hour TRECVID news dataset have shown the effectiveness, which can discover repetitive clips of various lengths and contents in only 25 minutes, with features extracted off-line. Junsong Yuan 0001, Wei Wang 0040, Jingjing Meng, Ying Wu 0001, Dongge Li |
ACM Multimedia | 4 |
| 2007 | A decentralized probabilistic approach to articulated body tracking
Gang Hua 0001, Ying Wu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2007 | Multiple Collaborative Kernel TrackingabstractThose motion parameters that cannot be recovered from image measurements are unobservable in the visual dynamic system. This paper studies this important issue of singularity in the context of kernel-based tracking and presents a novel approach that is based on a motion field representation which employs redundant but sparsely correlated local motion parameters instead of compact but uncorrelated global ones. This approach makes it easy to design fully observable kernel-based motion estimators. This paper shows that these high-dimensional motion fields can be estimated efficiently by the collaboration among a set of simpler local kernel-based motion estimators, which makes the new approach very practical. Zhimin Fan 0002, Ming Yang 0007, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2006 | Efficient Optimal Kernel Placement for Reliable Visual TrackingabstractThis paper describes a novel approach to optimal kernel placement in kernel-based tracking. If kernels are placed at arbitrary places, kernel-based methods are likely to be trapped in ill-conditioned locations, which prevents the reliable recovery of the motion parameters and jeopardizes the tracking performance. The theoretical analysis presented in this paper indicates that the optimal kernel placement can be evaluated based on a closed-form criterion, and achieved efficiently by a novel gradient-based algorithm. Based on that, new methods for temporal-stable multiple kernel placement and scale-invariant kernel placement are proposed. These new theoretical results and new algorithms greatly advance the study of kernel-based tracking in both theory and practice. Extensive real-time experimental results demonstrate the improved tracking reliability. Zhimin Fan 0002, Ming Yang 0007, Ying Wu 0001, Gang Hua 0001, Ting Yu 0003 |
CVPR (1) | 3 |
| 2006 | Measurement integration under inconsistency for robust trackingabstractThe solutions to many vision problems involve integrating measurements from multiple sources. Most existing methods rely on a hidden assumption, i.e., these measurements are consistent. In reality, unfortunately, this may not hold. The fact that naively fusing inconsistent measurements amounts to failing these methods indicates that this is not a trivial problem. This paper presents a novel approach to handling it. A new theorem is proven that gives two algebraic criteria to examine the consistency and inconsistency. In addition, a more general criterion is presented. Based on the theoretical analysis, a new information integration method is proposed and leads to encouraging results when applied to the task of visual tracking. Gang Hua 0001, Ying Wu 0001 |
CVPR (1) | 2 |
| 2006 | Intelligent Collaborative Tracking by Mining Auxiliary ObjectsabstractMany tracking methods face a fundamental dilemma in practice: tracking has to be computationally efficient but verifying if or not the tracker is following the true target tends to be demanding, especially when the background is cluttered and/or when occlusion occurs. Due to the lack of a good solution to this problem, many existing methods tend to be either computationally intensive with the use of sophisticated image observation models, or vulnerable to the false alarms. This greatly threatens long-duration robust tracking. This paper presents a novel solution to this dilemma by integrating into the tracking process a set of auxiliary objects that are automatically discovered in the video on the fly by data mining. Auxiliary objects have three properties at least in a short time interval: (1) persistent co-occurrence with the target; (2) consistent motion correlation with the target; and (3) easy to track. The collaborative tracking of these auxiliary objects leads to an efficient computation as well as a strong verification. Our extensive experiments have exhibited exciting performance in very challenging real-world testing cases. Ming Yang 0007, Ying Wu 0001, Shihong Lao |
CVPR (1) | 2 |
| 2006 | Differential Tracking based on Spatial-Appearance Model (SAM)abstractA fundamental issue in differential motion analysis is the compromise between the flexibility of the matching criterion for image regions and the ability of recovering the motion. Localized matching criteria, e.g., pixel-based SSD, may enable the recovery of all motion parameters, but it does not tolerate much appearance changes. On the other hand, global criteria, e.g., matching histograms, can accommodate dramatic appearance changes, but may be blind to some motion parameters, e.g., scaling and rotation. This paper presents a novel differential approach that integrates the advantages of both in a principled way based on a spatial-appearance model (SAM) that combines local appearances variations and global spatial structures. This model can capture a large variety of appearance variations that are attributed to the local non-rigidity. At the same time, this model enables efficient recovery of all motion parameters. A maximum likelihood matching criterion is defined and rigorous analytical results are obtained that lead to a closed form solution to motion tracking. Very encouraging results demonstrate the effectiveness and efficiency of the proposed method for tracking non-rigid objects that exhibit dramatic appearance deformations, large object scale changes and partial occlusions. Ting Yu 0003, Ying Wu 0001 |
CVPR (1) | 2 |
| 2006 | Tracking Motion-Blurred Targets in VideoabstractMany emerging applications require tracking targets in video. Most existing visual tracking methods do not work well when the target is motion-blurred (especially due to fast motion), because the imperfectness of the target's appearances invalidates the image matching model (or the measurement model) in tracking. This paper presents a novel method to track motion-blurred targets by taking advantage of the blurs without performing image restoration. Unlike the global blur induced by camera motion, this paper is concerned with the local blurs that are due to target's motion. This is a challenging task because the blurs need to be identified blindly. The proposed method addresses this difficulty by integrating signal processing and statistical learning techniques. The estimated blurs are used to reduce the search range by providing strong motion predictions and to localize the best match accurately by modifying the measurement models. Shengyang Dai, Ming Yang 0007, Ying Wu 0001, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2006 | Automatic Business Card Scanning with a CameraabstractIn this paper, we present a system to automatically extract, rectify and enhance business card images. First the business card image patch is automatically segmented by minimizing a novel local-global variational energy. Second a quadrangle is fitted to the segmented image patch. With the four corner points of the quadrangle, we then estimate the physical aspect ratio of the business card and obtain a homography to rectify the quadrangle back to rectangular shape. We finally enhance the contrast of the rectified business card image using a S-shaped curve. Extensive experiments demonstrated the efficacy and robustness of our system. Gang Hua 0001, Zicheng Liu 0001, Zhengyou Zhang, Ying Wu 0001 |
ICIP | 4 |
| 2006 | Face detection for automatic exposure control in handheld cameraabstractFace detection is a widely studied topic in computer vision, and advances in algorithms, low cost processing, and CMOS imagers make it practical for embedded consumer applications. As with graphics, the best cost-performance ratio is achieved with dedicated hardware. The challenges of face detection in embedded environments include bandwidth constraints set by low cost memory and a need to find parallelism. Consumer applications need reliability, calling for a hard real-time approach to guarantee that deadlines are met. We present a face detection system for automatic exposure control in a handheld digital camera or camera phone. Contributions include a complexity control scheme to meet hard real-time deadlines, a hardware pipeline design for Haar-like feature calculation, and a system design exploiting several levels of parallelism. The proposed architecture is verified by synthesis to Altera’s low cost Cyclone II FPGA. Simulation results show the algorithm can achieve about 80% detection rate for group portrait pictures. Ming Yang 0007, Ying Wu 0001, Jim E. Crenshaw, Bruce Augustine, Russell Mareachen |
ICVS | 2 |
| 2006 | Sequential mean field variational analysis of structured deformable shapes
Gang Hua 0001, Ying Wu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2006 | Multibody Grouping by Inference of Multiple Subspaces from High-Dimensional Data Using Oriented-FramesabstractRecently, subspace constraints have been widely exploited in many computer vision problems such as multibody grouping. Under linear projection models, feature points associated with multiple bodies reside in multiple subspaces. Most existing factorization-based algorithms can segment objects undergoing independent motions. However, intersections among the correlated motion subspaces will lead most previous factorization-based algorithms to erroneous segmentation. To overcome this limitation, in this paper, we formulate the problem of multibody grouping as inference of multiple subspaces from a high-dimensional data space. A novel and robust algorithm is proposed to capture the configuration of the multiple subspace structure and to find the segmentation of objects by clustering the feature points into these inferred subspaces, no matter whether they are independent or correlated. In the proposed method, an Oriented-Frame (OF), which is a multidimensional coordinate frame, is associated with each data point indicating the point's preferred subspace configuration. Based on the similarity between the subspaces, novel mechanisms of subspace evolution and voting are developed. By filtering the outliers due to their structural incompatibility, the subspace configurations will emerge. Compared with most existing factorization-based algorithms that cannot correctly segment correlated motions, such as motions of articulated objects, the proposed method has a robust performance in both independent and correlated motion segmentation. A number of controlled and real experiments show the effectiveness of the proposed method. However, the current approach does not deal with transparent motions and motion subspaces of different dimensions. Zhimin Fan 0002, Jie Zhou 0001, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2006 | Iterative Local-Global Energy Minimization for Automatic Extraction of Objects of InterestabstractWe propose a novel global-local variational energy to automatically extract objects of interest from images. Previous formulations only incorporate local region potentials, which are sensitive to incorrectly classified pixels during iteration. We introduce a global likelihood potential to achieve better estimation of the foreground and background models and, thus, better extraction results. Extensive experiments demonstrate its efficacy. Gang Hua 0001, Zicheng Liu 0001, Zhengyou Zhang, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2006 | A Field Model for Human Detection and TrackingabstractThe large shape variability and partial occlusions challenge most object detection and tracking methods for nonrigid targets such as pedestrians. This paper presents a new approach based on a two-layer statistical field model that characterizes the prior of the complex shape variations as a Boltzmann distribution and embeds this prior and the complex image likelihood into a Markov field. A probabilistic variational analysis of this model reveals a set of fixed-point equations characterizing the equilibrium of the field. It leads to computationally efficient methods for calculating the image likelihood and for training the model. Based on that, effective algorithms for detecting nonrigid objects are developed. This new approach has several advantages. First, it is intrinsically suitable for capturing local nonrigidity. In addition, due to the distributed likelihood, this approach is robust to partial occlusions. Moreover, the two-layer structure provides large flexibility of modeling the image observations, which makes the new method robust to clutters. Extensive experiments demonstrate its effectiveness. Ying Wu 0001, Ting Yu 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Multiple Collaborative Kernel TrackingabstractThis paper presents a novel multiple collaborative kernel approach to visual tracking. This approach treats kernel-based tracking in a more general setting, i.e., a relaxation and constraints formulation, in which a complex motion is represented by a set of inter-correlated simpler motions. With this formulation, we present a rigorous analysis on a critical issue of kernel observability and obtain a criterion, based on which we propose a new method using collaborative kernels that has the theoretical guarantee of enhanced observability. This new method has been shown to be computationally efficient in both theory and practice, which can be readily applied to complex motions such as articulated motions. Zhimin Fan 0002, Ying Wu 0001, Ming Yang 0007 |
CVPR (2) | 2 |
| 2005 | Learning to Estimate Human Pose with Data Driven Belief PropagationabstractWe propose a statistical formulation for 2D human pose estimation from single images. The human body configuration is modeled by a Markov network and the estimation problem is to infer pose parameters from image cues such as appearance, shape, edge, and color. From a set of hand labeled images, we accumulate prior knowledge of 2D body shapes by learning their low-dimensional representations for inference of pose parameters. A data driven belief propagation Monte Carlo algorithm, utilizing importance sampling functions built from bottom-up visual cues, is proposed for efficient probabilistic inference. Contrasted to the few sequential statistical formulations in the literature, our algorithm integrates both top-down as well as bottom-up reasoning mechanisms, and can carry out the inference tasks in parallel. Experimental results demonstrate the potency and effectiveness of the proposed algorithm in estimating 2D human pose from single images. Gang Hua 0001, Ming-Hsuan Yang 0001, Ying Wu 0001 |
CVPR (2) | 3 |
| 2005 | A Statistical Field Model for Pedestrian DetectionabstractThis paper presents a new statistical model for detecting and tracking deformable objects such as pedestrians, where large shape variations induced by local shape deformation can not be well captured by global methods such as PCA. The proposed model employs a Boltzmann distribution to capture the prior of local deformation, and embeds it into a Markov network which can be learned from data. A mean field variational analysis of this model provides computationally efficient algorithms for computing the likelihood of image observations and facilitate fast model training. Based on that, effective detection and tracking algorithms for deformable objects are proposed and applied to pedestrian detection and tracking. The proposed method has several advantages. Firstly, it captures local deformation well and thus is robust to occlusions and clutter. In addition, it is computationally tractable. Moreover, it divides deformation into local deformation and global deformation, then conquers them by combining bottom-up and top-down methodologies. Extensive experiments demonstrate the effectiveness of the proposed model for deformable objects. Ying Wu 0001, Ting Yu 0003, Gang Hua 0001 |
CVPR (1) | 1 |
| 2005 | Tracking Non-Stationary Appearances and Dynamic Feature SelectionabstractSince the appearance changes of the target jeopardize visual measurements and often lead to tracking failure in practice, trackers need to be adaptive to non-stationary appearances or to dynamically select features to track. However, this idea is threatened by the risk of adaptation drift that roots in its ill-posed nature, unless good constraints are imposed. Different from most existing adaptation schemes, we enforce three novel constraints for the optimal adaptation: (1) negative data, (2) bottom-up pair-wise data constraints, and (3) adaptation dynamics. Substantializing the general adaptation problem as a subspace adaptation problem, this paper gives a closed-form solution as well as a practical iterative algorithm. Extensive experiments have shown that the proposed approach can largely alleviate adaptation drift and achieve better tracking results. Ming Yang 0007, Ying Wu 0001 |
CVPR (2) | 2 |
| 2005 | Decentralized Multiple Target Tracking Using Netted Collaborative Autonomous TrackersabstractThis paper presents a decentralized approach to multiple target tracking. The novelty of this approach lies in the use of a set of autonomous while collaborative trackers to overcome the tracker coalescence problem with linear complexity. In this approach, the individual trackers are autonomous in the sense that they can select targets to track and evaluate themselves, and they are also collaborative since they need to compete for the targets against those trackers that are close to them through communication. The theoretical foundation of this new approach is based on the variational analysis of a Markov network that reveals the collaborative mechanism through fixed point iteration among these trackers and the existence of the equilibriums. In addition, a trained object detector is incorporated to help sense the potential newly appearing targets in the dynamic scene. Experimental results on challenging video sequences demonstrate the effectiveness and efficiency of the proposed method. Ting Yu 0003, Ying Wu 0001 |
CVPR (1) | 2 |
| 2005 | Variational Maximum A Posteriori by Annealed Mean Field AnalysisabstractThis paper proposes a novel probabilistic variational method with deterministic annealing for the maximum a posteriori (MAP) estimation of complex stochastic systems. Since the MAP estimation involves global optimization, in general, it is very difficult to achieve. Therefore, most probabilistic inference algorithms are only able to achieve either the exact or the approximate posterior distributions. Our method constrains the mean field variational distribution to be multivariate Gaussian. Then, a deterministic annealing scheme is nicely incorporated into the mean field fix-point iterations to obtain the optimal MAP estimate. This is based on the observation that when the covariance of the variational Gaussian distribution approaches to zero, the infimum point of the Kullback-Leibler (KL) divergence between the variational Gaussian and the real posterior will be the same as the supreme point of the real posterior. Although global optimality may not be guaranteed, our extensive synthetic and real experiments demonstrate the effectiveness and efficiency of the proposed method. Gang Hua 0001, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Analyzing and Capturing Articulated Hand Motion in Image SequencesabstractCapturing the human hand motion from video involves the estimation of the rigid global hand pose as well as the nonrigid finger articulation. The complexity induced by the high degrees of freedom of the articulated hand challenges many visual tracking techniques. For example, the particle filtering technique is plagued by the demanding requirement of a huge number of particles and the phenomenon of particle degeneracy. This paper presents a novel approach to tracking the articulated hand in video by learning and integrating natural hand motion priors. To cope with the finger articulation, this paper proposes a powerful sequential Monte Carlo tracking algorithm based on importance sampling techniques, where the importance function is based on an initial manifold model of the articulation configuration space learned from motion-captured data. In addition, this paper presents a divide-and-conquer strategy that decouples the hand poses and finger articulations and integrates them in an iterative framework to reduce the complexity of the problem. Our experiments show that this approach is effective and efficient for tracking the articulated hand. This approach can be extended to track other articulated targets. Ying Wu 0001, John Lin, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Self-supervised learning based on discriminative nonlinear features for image classification
Qi Tian 0001, Ying Wu 0001, Jie Yu 0001, Thomas S. Huang |
Pattern Recognit. | 2 |
| 2004 | Multibody Motion Segmentation Based on Simulated Annealing
Zhimin Fan 0002, Jie Zhou 0001, Ying Wu 0001 |
CVPR (1) | 3 |
| 2004 | Inference of Multiple Subspaces from High-Dimensional Data and Application to Multibody Grouping
Zhimin Fan 0002, Jie Zhou 0001, Ying Wu 0001 |
CVPR (2) | 3 |
| 2004 | Multi-Scale Visual Tracking by Sequential Belief Propagation
Gang Hua 0001, Ying Wu 0001 |
CVPR (1) | 2 |
| 2004 | Collaborative Tracking of Multiple Targets
Ting Yu 0003, Ying Wu 0001 |
CVPR (1) | 2 |
| 2004 | Learning based on kernel discriminant-EM algorithm for image classificationabstractIn image classification and other learning-based object recognition tasks, it is often tedious and expensive to label large training data sets. Discriminant-EM (DEM), proposed as a semi-supervised learning framework, takes both labeled and unlabeled data to learn classifiers. The paper extends the linear DEM to a nonlinear kernel algorithm, KDEM, and evaluates KDEM on both benchmark image databases and synthetic data. Various comparisons with other state-of-the-art learning techniques are investigated. Qi Tian 0001, Jie Yu 0001, Ying Wu 0001, Thomas S. Huang |
ICASSP (5) | 3 |
| 2004 | Robust Visual Tracking by Integrating Multiple Cues Based on Co-Inference Learning
Ying Wu 0001, Thomas S. Huang |
Int. J. Comput. Vis. | 1 |
| 2003 | Switching Observation Models for Contour Tracking in ClutterabstractWe propose a generative model approach to contour tracking against nonstationary clutter and to coping with occlusions by explicit modelling and inferring. The proposed dynamic Bayesian networks consist of multiple hidden processes, which model the target, the clutter and the occlusions. The image observation models, which depict the generation of the image features, are conditioned on all the hidden processes. Based on this framework, the tracker can automatically switch among different observation models according to the hidden states of the clutter and occlusions. In addition, the inference of these hidden states provides self-evaluations for the tracker. The tracking and inference are implemented based on sequence Monte Carlo techniques. The effectiveness of the proposed approach to robust tracking and inferring nonstationary clutter and occlusion is demonstrated for a variety of image sequences. Ying Wu 0001, Gang Hua 0001, Ting Yu 0003 |
CVPR (1) | 1 |
| 2003 | Tracking Appearances with OcclusionsabstractOcclusion is a difficult problem for appearance-based target tracking, especially when we need to track multiple targets simultaneously and maintain the target identities during tracking. To cope with the occlusion problem explicitly, this paper proposes a dynamic Bayesian network, which accommodates an extra hidden process for occlusion and stipulates the conditions on which the image observation likelihood is calculated. The statistical inference of such a hidden process can reveal the occlusion relations among different targets, which makes the tracker more robust against partial even complete occlusions. In addition, considering the fact that target appearances change with views, another generative model for multiple view representation is proposed by adding a switching variable to select from different view templates. The integration of the occlusion model and multiple view model results in a complex dynamic Bayesian network, where extra hidden processes describe the switch of targets' templates, the targets' dynamics, and the occlusions among different targets. The tracking and inferring algorithms are implemented by the sampling-based sequential Monte Carlo strategies. Our experiments show the effectiveness of the proposed probabilistic models and the algorithms. Ying Wu 0001, Ting Yu 0003, Gang Hua 0001 |
CVPR (1) | 1 |
| 2003 | Tracking Articulated Body by Dynamic Markov NetworkabstractA new method for visual tracking of articulated objects is presented. Analyzing articulated motion is challenging because the dimensionality increase potentially demands tremendous increase of computation. To ease this problem, we propose an approach that analyzes subparts locally while reinforcing the structural constraints at the mean time. The computational model of the proposed approach is based on a dynamic Markov network, a generative model which characterizes the dynamics and the image observations of each individual subpart as well as the motion constraints among different subparts. Probabilistic variational analysis of the model reveals a mean field approximation to the posterior densities of each subparts given visual evidence, and provides a computationally efficient way for such a difficult Bayesian inference problem. In addition, we design mean field Monte Carlo (MFMC) algorithms, in which a set of low dimensional particle filters interact with each other and solve the high dimensional problem collaboratively. Extensive experiments on tracking human body parts demonstrate the effectiveness, significance and computational efficiency of the proposed method. Ying Wu 0001, Gang Hua 0001, Ting Yu 0003 |
ICCV | 1 |
| 2002 | 3D model-based visual hand trackingabstractCapturing human hand motion through visual input is a challenging problem that involves the estimation of both global hand pose as well as the local finger articulation. This is a difficult task that requires a search in a high dimensional space due to the high degrees of freedom that fingers exhibit and the self occlusions caused by global hand motion. We propose a divide and conquer approach to estimate both global and local hand motion. The hand pose is determined from the palm using Iterative closed point (ICP) algorithm and factorization method. By incorporating natural hand motion constraints, we propose an efficient tracking algorithm based on a sequential Monte Carlo technique for tracking finger motion. Finally, the iteration step between the pose estimation and finger articulation tracking is performed in an EM fashion to obtain an accurate configuration estimation. Our experiments show that our approach is accurate and robust for natural hand movements. Thomas S. Huang, Ying Wu 0001, John Lin |
ICME (1) | 2 |
| 2002 | Nonstationary color tracking for vision-based human-computer interactionabstractSkin color offers a strong cue for efficient localization and tracking of human body parts in video sequences for vision-based human-computer interaction. Color-based target localization could be achieved by analyzing segmented skin color regions. However, one of the challenges of color-based target tracking is that color distributions would change in different lighting conditions such that fixed color models would be inadequate to capture nonstationary color distributions over time. Meanwhile, using a fixed skin color model trained by the data of a specific person would probably not work well for other people. Although some work has been done on adaptive color models, this problem still needs further studies. We present our investigation of color-based image segmentation and nonstationary color-based target tracking, by studying two different representations for color distributions. We propose the structure adaptive self-organizing map (SASOM) neural network that serves as a new color model. Our experiments show that such a representation is powerful for efficient image segmentation. Then, we formulate the nonstationary color tracking problem as a model transduction problem, the solution of which offers a way to adapt and transduce color classifiers in nonstationary color distributions. To fulfill model transduction, we propose two algorithms, the SASOM transduction and the discriminant expectation-maximization (EM), based on the SASOM color model and the Gaussian mixture color model, respectively. Our extensive experiments on the task of real-time face/hand localization show that these two algorithms can successfully handle some difficulties in nonstationary color tracking. We also implemented a real-time face/hand localization system based on such algorithms for vision-based human-computer interaction. Ying Wu 0001, Thomas S. Huang |
IEEE Trans. Neural Networks | 1 |
| 2001 | Multibody Grouping via Orthogonal Subspace DecompositionabstractMultibody structure from motion could be solved by the factorization approach. However, the noise measurements would make the segmentation difficult when analyzing the shape interaction matrix. This paper presents an orthogonal subspace decomposition and grouping technique to approach such a problem. We decompose the object shape spaces into signal subspaces and noise subspaces. We show that the signal subspaces of the object shape spaces are orthogonal to each other. Instead of using the shape interaction matrix contaminated by noise, we introduce the shape signal subspace distance matrix for shape space grouping. Outliers could be easily identified by this approach. The robustness of the proposed approach lies in the fact that the shape space decomposition alleviates the influence of noise, and has been verified with extensive experiments. Ying Wu 0001, Zhengyou Zhang, Thomas S. Huang, John Y. Lin |
CVPR (2) | 1 |
| 2001 | A Co-inference Approach to Robust Visual TrackingabstractVisual tracking could be treated as a parameter estimation problem of target representation based on observations in image sequences. A richer target representations would incur better chances of successful tracking in cluttered and dynamic environments. However, the dimensionality of target's state space also increases making tracking a formidable estimation problem. In this paper, the problem of tracking and integrating multiple cues is formulated in a probabilistic framework; and represented by factorized graphical model. Structured variational analysis of such graphical model factorizes different modalities and suggests a co-inference process among these modalities. A sequential Monte Carlo algorithm is proposed to give an efficient approximation of the co-inference based on the importance sampling technique. This algorithm is implemented in real-time at around 30 Hz. Specifically, tracking both position, shape and color distribution of a target is investigated in this paper. Our extensive experiments show that the proposed algorithm performs robustly in a large variety of trucking scenarios. The approach presented in this paper has the potential to solve other sensor fusion problems. Ying Wu 0001, Thomas S. Huang |
ICCV | 1 |
| 2001 | Self-Supervised Learning for Object Recognition based on Kernel Discriminant-EM Algorithm
Ying Wu 0001, Thomas S. Huang, Kentaro Toyama |
ICCV | 1 |
| 2001 | Capturing Natural Hand ArticulationabstractVision-based motion capturing of hand articulation is a challenging task, since the hand presents a motion of high degrees of freedom. Model-based approaches could be taken to approach this problem by searching in a high dimensional hand state space, and matching projections of a hand model and image observations. However, it is highly inefficient due to the curse of dimensionality. Fortunately, natural hand articulation is highly constrained, which largely reduces the dimensionality of hand state space. This paper presents a model-based method to capture hand articulation by learning hand natural constraints. Our study shows that natural hand articulation lies in a lower dimensional configurations space characterized by a union of liner manifolds spanned by a set of basis configurations. By integrating hand motion constraints, an efficient articulated motion-capturing algorithm is proposed based on sequential Monte Carlo techniques. Our experiments show that this algorithm is robust and accurate for tracking natural hand movements. This algorithm is easy to extend to other articulated motion capturing tasks. Ying Wu 0001, John Y. Lin, Thomas S. Huang |
ICCV | 1 |
| 2001 | Spoken Language Acquisition Via Human-Robot InteractionabstractThis paper presents a subproject of a challenging project that explores teaching a computer human-intelligence. In the subproject, a multisensory mobile robot is used as the interface for human-computer interaction, and spoken language is taught to the computer through natural human-robot interaction. Different from state-of-the-art speech recognizers, our approach associates speech patterns directly with sensory inputs of the robot. This approach allows our system to learn multilingual speech patterns online. Further investigation of this project will include human-computer interaction that involves more modalities, and applications that use the proposed idea to train home appliances. 1. Qiong Liu 0003, Thomas S. Huang, Ying Wu 0001, Stephen E. Levinson |
ICME | 3 |
| 2000 | Color Tracking by Transductive LearningabstractOne of the difficulties of color tracking is that color changes in different lighting conditions, and static color models would be inadequate to capture the nonstationary color distribution over time. Although some work has been done on adaptive color models, this problem still needs further investigation. Different from many other approaches, we formulate the nonstationary color tracking problem as a transductive learning problem, in which the generalization of a trained color classifier is only defined on the pixels in a specific image, rather than the whole color space. This formulation offers a way to design and transduce color classifiers through non-stationary color distribution. Instead of assuming a color transition model. We assume that some unlabeled pixels in a new image frame can be "confidently" labeled by a "weak classifier" according to a preset confidence level. The proposed Discriminant-EM (D-EM) algorithm offers an effective way to transduce color classifiers as well as automatically select a good color space. Experiments show that D-EM successfully handles some problems in color tracking. As a component our natural gesture interface, this algorithm gives tight bounding boxes of the hand or face regions in video sequences. Ying Wu 0001, Thomas S. Huang |
CVPR | 1 |
| 2000 | View-Independent Recognition of Hand PosturesabstractSince the human hand is highly articulated and deformable, hand posture recognition is a challenging example in the research on view-independent object recognition. Due to the difficulties of the model-based approach, the appearance-based learning approach is promising to handle large variation in visual inputs. However, the generalization of many proposed supervised learning methods to this problem often suffers from the insufficiency of labeled training data. This paper describes an approach to alleviate this difficulty by adding a large unlabeled training set. Combining supervised and unsupervised learning paradigms, a novel and powerful learning approach, the Discriminant-EM (D-EM) algorithm, is proposed in this paper to handle the case of a small labeled training set. Experiments show that D-EM outperforms many other learning methods. Based on this approach, we implement a gesture interface to recognize a set of predefined gesture commands, and it is also extended to hand detection. This algorithm can also apply to other object recognition tasks. Ying Wu 0001, Thomas S. Huang |
CVPR | 1 |
| 2000 | Discriminant-EM Algorithm with Application to Image RetrievalabstractIn many vision applications, the practice of supervised learning faces several difficulties, one of which is that insufficient labeled training data result in poor generalization. In image retrieval, we have very few labeled images from query and relevance feedback so that it is hard to automatically weight image features and select similarity metrics for image classification. This paper investigates the possibility of including an unlabeled data set to make up the insufficiency of labeled data. Different from most current research in image retrieval, the proposed approach tries to cast image retrieval as a transductive learning problem, in which the generalization of an image classifier is only defined on a set of images such as the given image database. Formulating this transductive problem in a probabilistic framework the proposed algorithm, Discriminant EM (D-EM) not only estimates the parameters of a generative model but also finds a linear transformation to relax the assumption of probabilistic structure of data distributions as well as select good features automatically. Our experiments show that D-EM has a satisfactory performance in image retrieval applications. D-EM algorithm has the potential to many other applications. Ying Wu 0001, Qi Tian 0001, Thomas S. Huang |
CVPR | 1 |
| 2000 | Bootstrap Initialization of Nonparametric Texture Models for Tracking
Kentaro Toyama, Ying Wu 0001 |
ECCV (2) | 2 |
| 2000 | Wide-Range, Person- and Illumination-Insensitive Head Orientation EstimationabstractWe present an algorithm for estimation of head orientation, given cropped images of a subject's head from any viewpoint. Our algorithm handles dramatic changes in illumination, applies to many people without per-user initialization, and covers a wider range (e.g., side and back) of head orientations than previous algorithms. The algorithm builds an ellipsoidal model of the head, where points on the model maintain probabilistic information about surface edge density. To collect data for each point on the model, edge-density features are extracted from hand-annotated training images and projected into the model. Each model point learns a probability density function from the training observations. During pose estimation, features are extracted from input images; then, the maximum a posteriori pose is sought, given the current observation. Ying Wu 0001, Kentaro Toyama |
FG | 1 |
| 2000 | Combine User Defined Region-of-Interest and Spatial Layout for Image RetrievalabstractContent-based image retrieval (CBIR) is one of the most active research areas. Many visual feature representations have been explored and many systems built. However, in most of current systems, only the global features such as overall color histogram and texture moments are used which ignore the actual composition of the image in terms of internal objects. Although relevance feedback was proposed (Rui and Huang 1998) to incrementally supply more information, this may fail due to the lack of higher-level information about what exactly was of interest. Since automatic segmentation of the region-of-interest (ROI) is not always reliable, human assistance is necessary. In this paper, a novel approach combining a user defined region-of-interest and spatial layout is proposed for CBIR. Better capture of the image object is achieved by the user rather than the computer. Therefore, more accurate relevance feedback is achieved and thus lends to a more powerful search engine. Qi Tian 0001, Ying Wu 0001, Thomas S. Huang |
ICIP | 2 |
| 2000 | Integrating Unlabeled Images for Image Retrieval Based on Relevance FeedbackabstractRetrieval techniques based on pure similarity metrics are often suffered from the scales of image features. An alternative approach is to learn a mapping based on queries and relevance feedback by supervised learning. However, the learning is plagued by the insufficiency of labeled training images. Different from most current research in image retrieval, this paper investigates the possibility of taking advantage of unlabeled images in the given image database to make a hybrid statistical learning feasible. Assuming a generative model of the database, the proposed approach casts image retrieval as a transductive learning problem in a probabilistic framework. Our experiments show that the proposed approach has a satisfactory performance in image retrieval applications. Ying Wu 0001, Qi Tian 0001, Thomas S. Huang |
ICPR | 1 |
| 1999 | Capturing Articulated Human Hand Motion: A Divide-and-Conquer ApproachabstractThe use of the human hand as a natural interface device serves as a motivating force for research in the modeling, analysis and capture of the motion of an articulated hand. Model-based hand motion capture can be formulated as a large nonlinear programming problem, but this approach is plagued by local minima. An alternative way is to use analysis-by-synthesis by searching a huge space, but the results are rough and the computation expensive. In this paper, articulated hand motion is decoupled, a new two-step iterative model-based algorithm is proposed to capture articulated human hand motion, and a proof of convergence of this iterative algorithm is also given. In our proposed work, the decoupled global hand motion and local finger motion are parameterized by the 3D hand pose and the state of the hand respectively. Hand pose determination is formulated as a least-median-of-squares (LMS) problem rather than the nonrobust least-squares (LS) problem, so that 3D hand pose can be reliably calculated even if there are outliers. Local finger motion is formulated as an inverse kinematics problem. A genetic algorithm-based method is proposed to find a sub-optimal solution of the inverse kinematics effectively. Our algorithm and the LS-based algorithm are compared in several experiments. Both algorithms converge when local finger motion between consecutive frames is small. When large finger motion is present, the LS-based method fails, but our algorithm can still estimate the global and local finger motion well. Ying Wu 0001, Thomas S. Huang |
ICCV | 1 |
| 1999 | Human Hand Modeling, Analysis, and Animation in the Context of HCIabstractThe use of human hand as a natural interface device serves as a motivating force for research in visual analysis of highly articulated hand movement. Since hand motion covers a huge domain, the scope of this paper is limited to the developments of 3D model-based approaches. Numerous 3D models that have been used to analyze hand motion are studied. Various approaches to articulated motion analysis are discussed. Some realistic synthesis methods are also included in this paper. We conclude with some thoughts about future research directions. Ying Wu 0001, Thomas S. Huang |
ICIP (3) | 1 |