EDBT 2026 Demo / reviewers in the wild / expert
Stephen J. Maybank
dblp:m/StephenJMaybank · also Stephen John Maybank, Steve J. Maybank
· DBLP profile ↗
173ranked-venue papers
28as first author
23since 2021 · last 2025
0000-0003-2113-9119ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 124 · 28 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 65 · 6 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 10Databases, data management, data science and information retrieval · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MI-TRQR: Mutual Information-Based Temporal Redundancy Quantification and Reduction for Energy-Efficient Spiking Neural NetworksabstractBrain-inspired spiking neural networks (SNNs) provide energy-efficient computation through event-driven processing. However, the shared weights across multiple timesteps lead to serious temporal feature redundancy, limiting both efficiency and performance. This issue is further aggravated when processing static images due to the duplicated input. To mitigate this problem, we propose a parameter-free and plug-and-play module named Mutual Information-based Temporal Redundancy Quantification and Reduction (MI-TRQR), constructing energy-efficient SNNs. Specifically, Mutual Information (MI) is properly introduced to quantify redundancy between discrete spike features at different timesteps on two spatial scales: pixel (local) and the entire spatial features (global). Based on the multi-scale redundancy quantification, we apply a probabilistic masking strategy to remove redundant spikes. The final representation is subsequently recalibrated to account for the spike removal. Extensive experimental results demonstrate that our MI-TRQR achieves sparser spiking firing, higher energy efficiency, and better performance concurrently with different SNN architectures in tasks of neuromorphic data classification, static data classification, and time-series forecasting. Notably, MI-TRQR increases accuracy by \textbf{1.7\%} on CIFAR10-DVS with 4 timesteps while reducing energy cost by \textbf{37.5\%}. Our codes are available at https://github.com/dfxue/MI-TRQR. Dengfeng Xue, Yifan Lu 0001, Chunfeng Yuan, Yufan Liu 0001, Wei Liu 0153, Man Yao, Li Yang 0014, Bing Li 0001, Stephen J. Maybank, Weiming Hu 0004, Zhetao Li |
NeurIPS | 11 |
| 2025 | Two-stream transformer tracking with messengers
Miaobo Qiu, Wenyang Luo, Tongfei Liu, Yanqin Jiang, Jiaming Yan, Weiming Hu 0004, Stephen J. Maybank |
Image Vis. Comput. | 9 |
| 2025 | FiGVCL: Fine-Grained Benchmark and Method for Video Copy LocalizationabstractContent-based video copy localization (VCL) aims to detect and locate copied segments in pairs of videos. VCL requires fine-grained video analysis to robustly identify copied segments that have been edited. Despite recent progress, the prohibitive cost of annotating copied segments and the lack of a fine-grained benchmark hinder the development of effective VCL systems. In this work, we annotate a new real-world dataset, FiGVCL, with challenging scenarios designed to evaluate VCL methods. FiGVCL is carefully annotated to preserve the temporal correspondences observed in copied segments. Moreover, we propose a novel fine-grained VCL benchmark metric based on temporal correspondences to improve discriminability. Finally, we design a simple but effective baseline model that uses fine-grained local embeddings for accurate copied segment localization. We also present an unsupervised training strategy that outperforms previous supervised VCL methods. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Task-Aware Attentional Dynamic Alignment for Few-Shot Compressed Video ClassificationabstractWe present a novel Task-aware Attentional Dynamic Alignment (TADA) framework for visual-based few-shot video classification (FSVC) that addresses two key challenges in this field: efficiency and nuanced spatio-temporal reasoning. Existing methods are often hindered by computationally expensive video decoding processes and neglect the temporal order of videos. In contrast, our method harnesses compressed domain data to extract rich spatio-temporal cues at a fraction of the cost of traditional video processing methods. Specifically, we propose an embedding module to extract informative features from compressed domain data while minimizing computational overheads. Furthermore, to exploit the temporal order of frames, we develop a prototypical ADA module to align and classify videos with an explicit temporal order constraint. Our framework also incorporates a contextual mixer to enrich video embeddings with task-specific context. Extensive experiments on multiple datasets demonstrate that TADA achieves state-of-the-art performance and outperforms existing methods in accuracy and efficiency. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Cross-Architecture Knowledge Distillation
Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jingting Ding, Liang Li 0006, Stephen J. Maybank |
Int. J. Comput. Vis. | 7 |
| 2024 | DCFNet: Discriminant Correlation Filters Network for Visual Tracking
Weiming Hu 0004, Qiang Wang 0051, Bing Li 0001, Stephen J. Maybank |
J. Comput. Sci. Technol. | 5 |
| 2024 | One-Stage Anchor-Free Online Multiple Target Tracking With Deformable Local Attention and Task-Aware PredictionabstractThe tracking-by-detection paradigm currently dominates multiple target tracking algorithms. It usually includes three tasks: target detection, appearance feature embedding, and data association. Carrying out these three tasks successively usually leads to lower tracking efficiency. In this paper, we propose a one-stage anchor-free multiple task learning framework which carries out target detection and appearance feature embedding in parallel to substantially increase the tracking speed. This framework simultaneously predicts a target detection and produces a feature embedding for each location, by sharing a pyramid of feature maps. We propose a deformable local attention module which utilizes the correlations between features at different locations within a target to obtain more discriminative features. We further propose a task-aware prediction module which utilizes deformable convolutions to select the most suitable locations for the different tasks. At the selected locations, classification of samples into foreground or background, appearance feature embedding, and target box regression are carried out. Two effective training strategies, regression range overlapping and sample reweighting, are proposed to reduce missed detections in dense scenes. Ambiguous samples whose identities are difficult to determine are effectively dealt with to obtain more accurate feature embedding of target appearance. An appearance-enhanced non-maximum suppression is proposed to reduce over-suppression of true targets in crowded scenes. Based on the one-stage anchor-free network with the deformable local attention module and the task-aware prediction module, we implement a new online multiple target tracker. Experimental results show that our tracker achieves a very fast speed while maintaining a high tracking accuracy. Weiming Hu 0004, Shaoru Wang, Zongwei Zhou, Yangxi Li, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Chinese Title Generation for Short Videos: Dataset, Metric and AlgorithmabstractPrevious work for video captioning aims to objectively describe the video content but the captions lack human interest and attractiveness, limiting its practical application scenarios. The intention of video title generation (video titling) is to produce attractive titles, but there is a lack of benchmarks. This work offers CREATE, the first large-scale Chinese shoRt vidEo retrievAl and Title gEneration dataset, to assist research and applications in video titling, video captioning, and video retrieval in Chinese. CREATE comprises a high-quality labeled 210 K dataset and two web-scale 3 M and 10 M pre-training datasets, covering 51 categories, 50K+ tags, 537K+ manually annotated titles and captions, and 10M+ short videos with original video information. This work presents ACTEr, a unique Attractiveness-Consensus-based Title Evaluation, to objectively evaluate the quality of video title generation. This metric measures the semantic correlation between the candidate (model-generated title) and references (manual-labeled titles) and introduces attractive consensus weights to assess the attractiveness and relevance of the video title. Accordingly, this work proposes a novel multi-modal ALignment WIth Generation model, ALWIG, as one strong baseline to aid future model development. With the help of a tag-driven video-text alignment module and a GPT-based generation module, this model achieves video titling, captioning, and retrieval simultaneously. We believe that the release of the CREATE dataset, ACTEr metric, and ALWIG model will encourage in-depth research on the analysis and creation of Chinese short videos. Ziqi Zhang 0010, Zongyang Ma, Chunfeng Yuan, Peijin Wang, Zhongang Qi, Chenglei Hao, Bing Li 0001, Ying Shan, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 11 |
| 2023 | Ranking-Based Color Constancy With Limited Training SamplesabstractComputational color constancy is an important component of Image Signal Processors (ISP) for white balancing in many imaging devices. Recently, deep convolutional neural networks (CNN) have been introduced for color constancy. They achieve prominent performance improvements comparing with those statistics or shallow learning-based methods. However, the need for a large number of training samples, a high computational cost and a huge model size make CNN-based methods unsuitable for deployment on low-resource ISPs for real-time applications. In order to overcome these limitations and to achieve comparable performance to CNN-based methods, an efficient method is defined for selecting the optimal simple statistics-based method (SM) for each image. To this end, we propose a novel ranking-based color constancy method (RCC) that formulates the selection of the optimal SM method as a label ranking problem. RCC designs a specific ranking loss function, and uses a low rank constraint to control the model complexity and a grouped sparse constraint for feature selection. Finally, we apply the RCC model to predict the order of the candidate SM methods for a test image, and then estimate its illumination using the predicted optimal SM method (or fusing the results estimated by the top k SM methods). Comprehensive experiment results show that the proposed RCC outperforms nearly all the shallow learning-based methods and achieves comparable performance to (sometimes even better performance than) deep CNN-based methods with only 1/2000 of the model size and training time. RCC also shows good robustness to limited training samples and good generalization crossing cameras. Furthermore, to remove the dependence on the ground truth illumination, we extend RCC to obtain a novel ranking-based method without ground truth illumination (RCC_NO) that learns the ranking model using simple partial binary preference annotations provided by untrained annotators rather than experts. RCC_NO also achieves better performance than the SM methods and most shallow learning-based methods with low costs of sample collection and illumination measurement. Bing Li 0001, Haina Qin, Weihua Xiong, Yangxi Li, Songhe Feng, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Learning to Explore Distillability and Sparsability: A Joint Framework for Model CompressionabstractDeep learning shows excellent performance usually at the expense of heavy computation. Recently, model compression has become a popular way of reducing the computation. Compression can be achieved using knowledge distillation or filter pruning. Knowledge distillation improves the accuracy of a lightweight network, while filter pruning removes redundant architecture in a cumbersome network. They are two different ways of achieving model compression, but few methods simultaneously consider both of them. In this paper, we revisit model compression and define two attributes of a model: distillability and sparsability, which reflect how much useful knowledge can be distilled and how many pruned ratios can be obtained, respectively. Guided by our observations and considering both accuracy and model size, a dynamically distillability-and-sparsability learning framework (DDSL) is introduced for model compression. DDSL consists of teacher, student and dean. Knowledge is distilled from the teacher to guide the student. The dean controls the training process by dynamically adjusting the distillation supervision and the sparsity supervision in a meta-learning framework. An alternating direction method of multiplier (ADMM)-based knowledge distillation-with-pruning (KDP) joint optimization algorithm is proposed to train the model. Extensive experimental results show that DDSL outperforms 24 state-of-the-art methods, including both knowledge distillation and filter pruning methods. Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Self-Prior Guided Pixel Adversarial Networks for Blind Image InpaintingabstractBlind image inpainting involves two critical aspects, i.e., "where to inpaint" and "how to inpaint". Knowing "where to inpaint" can eliminate the interference arising from corrupted pixel values; a good "how to inpaint" strategy yields high-quality inpainted results robust to various corruptions. In existing methods, these two aspects usually lack explicit and separate consideration. This paper fully explores these two aspects and proposes a self-prior guided inpainting network (SIN). The self-priors are obtained by detecting semantic-discontinuous regions and by predicting global semantic structures of the input image. On the one hand, the self-priors are incorporated into the SIN, which enables the SIN to perceive valid context information from uncorrupted regions and to synthesize semantic-aware textures for corrupted regions. On the other hand, the self-priors are reformulated to provide a pixel-wise adversarial feedback and a high-level semantic structure feedback, which can promote the semantic continuity of inpainted images. Experimental results demonstrate that our method achieves state-of-the-art performance in metric scores and in visual quality. It has an advantage over many existing methods that assume "where to inpaint" is known in advance. Extensive experiments on a series of related image restoration tasks validate the effectiveness of our method in obtaining high-quality inpainting. Juan Wang 0012, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Wide-Angle Image Rectification: A Survey
Jinlong Fan 0001, Jing Zhang 0037, Stephen J. Maybank, Dacheng Tao |
Int. J. Comput. Vis. | 3 |
| 2022 | Bridging Composite and Real: Towards End-to-End Deep Image Matting
Jizhizi Li, Jing Zhang 0037, Stephen J. Maybank, Dacheng Tao |
Int. J. Comput. Vis. | 3 |
| 2022 | Interaction-Aware Spatio-Temporal Pyramid Attention Networks for Action ClassificationabstractFor CNN-based visual action recognition, the accuracy may be increased if local key action regions are focused on. The task of self-attention is to focus on key features and ignore irrelevant information. So, self-attention is useful for action recognition. However, current self-attention methods usually ignore correlations among local feature vectors at spatial positions in CNN feature maps. In this paper, we propose an effective interaction-aware self-attention model which can extract information about the interactions between feature vectors to learn attention maps. Since the different layers in a network capture feature maps at different scales, we introduce a spatial pyramid with the feature maps at different layers for attention modeling. The multi-scale information is utilized to obtain more accurate attention scores. These attention scores are used to weight the local feature vectors of the feature maps and then calculate attentional feature maps. Since the number of feature maps input to the spatial pyramid attention layer is unrestricted, we easily extend this attention layer to a spatio-temporal version. Our model can be embedded in any general CNN to form a video-level end-to-end attention network for action recognition. Several methods are investigated to combine the RGB and flow streams to obtain accurate predictions of human actions. Experimental results show that our method achieves state-of-the-art results on the datasets UCF101, HMDB51, Kinetics-400, and untrimmed Charades. Weiming Hu 0004, Chunfeng Yuan, Bing Li 0001, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Robust Face Alignment via Deep Progressive Reinitialization and Adaptive Error-Driven LearningabstractRegression-based face alignment involves learning a series of mapping functions to predict the true landmarks from an initial estimation of the alignment. Most existing approaches focus on learning efficacious mapping functions from some feature representations to improve performance. The issues related to the initial alignment estimation and the final learning objective, however, receive less attention. This work proposes a deep regression architecture with progressive reinitialization and a new error-driven learning loss function to explicitly address the above two issues. Given an image with a rough face detection result, the full face region is first mapped by a supervised spatial transformer network to a normalized form and trained to regress coarse positions of landmarks. Then, different face parts are further respectively reinitialized to their own normalized states, followed by another regression sub-network to refine the landmark positions. To deal with the inconsistent annotations in existing training datasets, we further propose an adaptive landmark-weighted loss function. It dynamically adjusts the importance of different landmarks according to their learning errors during training without depending on any hyper-parameters manually set by trial and error. A high level of robustness to annotation inconsistencies is thus achieved. The whole deep architecture permits training from end to end, and extensive experimental analyses and comparisons demonstrate its effectiveness and efficiency. The source code, trained models, and experimental results are made available at https://github.com/shaoxiaohu/Face_Alignment_DPR.git. Xiaohu Shao, Junliang Xing, Jiangjing Lyu, Yu Shi 0003, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Exposure Trajectory Recovery From Motion BlurabstractMotion blur in dynamic scenes is an important yet challenging research topic. Recently, deep learning methods have achieved impressive performance for dynamic scene deblurring. However, the motion information contained in a blurry image has yet to be fully explored and accurately formulated because: (i) the ground truth of dynamic motion is difficult to obtain; (ii) the temporal ordering is destroyed during the exposure; and (iii) the motion estimation from a blurry image is highly ill-posed. By revisiting the principle of camera exposure, motion blur can be described by the relative motions of sharp content with respect to each exposed position. In this paper, we define exposure trajectories, which represent the motion information contained in a blurry image and explain the causes of motion blur. A novel motion offset estimation framework is proposed to model pixel-wise displacements of the latent sharp image at multiple timepoints. Under mild constraints, our method can recover dense, (non-)linear exposure trajectories, which significantly reduce temporal disorder and ill-posed problems. Finally, experiments demonstrate that the recovered exposure trajectories not only capture accurate and interpretable motion information from a blurry image, but also benefit motion-aware image deblurring and warping-based video extraction tasks. Codes are available on https://github.com/yjzhang96/Motion-ETR. Youjian Zhang, Stephen J. Maybank, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | QuadNet: Quadruplet loss for multi-view learning in baggage re-identification
Hao Yang 0010, Xiuxiu Chu, Li Zhang 0050, Yunda Sun, Dong Li 0040, Stephen J. Maybank |
Pattern Recognit. | 6 |
| 2022 | Single Image Haze Removal Based on a Simple Additive Model With Haze Smoothness PriorabstractSingle image haze removal, which is to recover the clear version of a hazy image, is a challenging task in computer vision. In this paper, an additive haze model is proposed to approximate the hazy image formation process. In contrast with the traditional optical model, it regards the haze as an additive layer to a clean image. The model thus avoids estimating the medium transmission rate and the global atmospherical light. In addition, based on a critical observation that haze changes gradually and smoothly across the image, a haze smoothness prior is proposed to constrain this model. This prior assumes that the haze layer is much smoother than the clear image. Benefiting from this prior, we can directly separate the clean image from a single hazy image. Experimental results and comparisons with synthetic images and real-world images demonstrate that the proposed method outperforms state-of-the-art single image haze removal algorithms. Xiaoqin Zhang 0002, Tao Wang 0052, Guiying Tang, Li Zhao 0005, Yuewang Xu, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | DUT: Learning Video Stabilization by Simply Watching Unstable VideosabstractPrevious deep learning-based video stabilizers require a large scale of paired unstable and stable videos for training, which are difficult to collect. Traditional trajectory-based stabilizers, on the other hand, divide the task into several sub-tasks and tackle them subsequently, which are fragile in textureless and occluded regions regarding the usage of hand-crafted features. In this paper, we attempt to tackle the video stabilization problem in a deep unsupervised learning manner, which borrows the divide-and-conquer idea from traditional stabilizers while leveraging the representation power of DNNs to handle the challenges in real-world scenarios. Technically, DUT is composed of a trajectory estimation stage and a trajectory smoothing stage. In the trajectory estimation stage, we first estimate the motion of keypoints, initialize and refine the motion of grids via a novel multi-homography estimation strategy and a motion refinement network, respectively, and get the grid-based trajectories via temporal association. In the trajectory smoothing stage, we devise a novel network to predict dynamic smoothing kernels for trajectory smoothing, which can well adapt to trajectories with different dynamic patterns. We exploit the spatial and temporal coherence of keypoints and grid vertices to formulate the training objectives, resulting in an unsupervised training scheme. Experiment results on public benchmarks show that DUT outperforms state-of-the-art methods both qualitatively and quantitatively. The source code is available at https://github.com/Annbless/DUTCode. Yufei Xu, Jing Zhang 0037, Stephen J. Maybank, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2022 | Feedback Graph Convolutional Network for Skeleton-Based Action RecognitionabstractSkeleton-based action recognition has attracted considerable attention since the skeleton data is more robust to the dynamic circumstances and complicated backgrounds than other modalities. Recently, many researchers have used the Graph Convolutional Network (GCN) to model spatial-temporal features of skeleton sequences by an end-to-end optimization. However, conventional GCNs are feedforward networks for which it is impossible for the shallower layers to access semantic information in the high-level layers. In this paper, we propose a novel network, named Feedback Graph Convolutional Network (FGCN). This is the first work that introduces a feedback mechanism into GCNs for action recognition. Compared with conventional GCNs, FGCN has the following advantages: (1) A multi-stage temporal sampling strategy is designed to extract spatial-temporal features for action recognition in a coarse to fine process; (2) A Feedback Graph Convolutional Block (FGCB) is proposed to introduce dense feedback connections into the GCNs. It transmits the high-level semantic features to the shallower layers and conveys temporal information stage by stage to model video level spatial-temporal features for action recognition; (3) The FGCN model provides predictions on-the-fly. In the early stages, its predictions are relatively coarse. These coarse predictions are treated as priors to guide the feature learning in later stages, to obtain more accurate predictions. Extensive experiments on three datasets, NTU-RGB+D, NTU-RGB+D120 and Northwestern-UCLA, demonstrate that the proposed FGCN is effective for action recognition. It achieves the state-of-the-art performance on all three datasets. Hao Yang 0010, Dan Yan, Li Zhang 0050, Yunda Sun, Dong Li 0040, Stephen J. Maybank |
IEEE Trans. Image Process. | 6 |
| 2021 | 3D-FUTURE: 3D Furniture Shape with TextURE
Huan Fu, Rongfei Jia, Lin Gao 0004, Mingming Gong, Binqiang Zhao, Stephen J. Maybank, Dacheng Tao |
Int. J. Comput. Vis. | 6 |
| 2021 | Knowledge Distillation: A Survey
Jianping Gou, Baosheng Yu, Stephen J. Maybank, Dacheng Tao |
Int. J. Comput. Vis. | 3 |
| 2021 | EDP: An Efficient Decomposition and Pruning Scheme for Convolutional Neural Network CompressionabstractModel compression methods have become popular in recent years, which aim to alleviate the heavy load of deep neural networks (DNNs) in real-world applications. However, most of the existing compression methods have two limitations: 1) they usually adopt a cumbersome process, including pretraining, training with a sparsity constraint, pruning/decomposition, and fine-tuning. Moreover, the last three stages are usually iterated multiple times. 2) The models are pretrained under explicit sparsity or low-rank assumptions, which are difficult to guarantee wide appropriateness. In this article, we propose an efficient decomposition and pruning (EDP) scheme via constructing a compressed-aware block that can automatically minimize the rank of the weight matrix and identify the redundant channels. Specifically, we embed the compressed-aware block by decomposing one network layer into two layers: a new weight matrix layer and a coefficient matrix layer. By imposing regularizers on the coefficient matrix, the new weight matrix learns to become a low-rank basis weight, and its corresponding channels become sparse. In this way, the proposed compressed-aware block simultaneously achieves low-rank decomposition and channel pruning by only one single data-driven training stage. Moreover, the network of architecture is further compressed and optimized by a novel Pruning & Merging (PM) module which prunes redundant channels and merges redundant decomposed layers. Experimental results (17 competitors) on different data sets and networks demonstrate that the proposed EDP achieves a high compression ratio with acceptable accuracy degradation and outperforms state-of-the-arts on compression rate, accuracy, inference time, and run-time memory. Xiaofeng Ruan, Yufan Liu 0001, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Yangxi Li, Stephen J. Maybank |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2020 | Spatiotemporal Attacks for Embodied Agents
Aishan Liu, Tairan Huang 0003, Xianglong Liu 0001, Yitao Xu 0002, Yuqing Ma, Stephen J. Maybank, Dacheng Tao |
ECCV (17) | 7 |
| 2020 | Dual L1-Normalized Context Aware Tensor Power Iteration and Its Applications to Multi-object Tracking and Multi-graph MatchingabstractAbstract The multi-dimensional assignment problem is universal for data association analysis such as data association-based visual multi-object tracking and multi-graph matching. In this paper, multi-dimensional assignment is formulated as a rank-1 tensor approximation problem. A dualL1-normalized context/hyper-context aware tensor power iteration optimization method is proposed. The method is applied to multi-object tracking and multi-graph matching. In the optimization method, tensor power iteration with the dual unit norm enables the capture of information across multiple sample sets. Interactions between sample associations are modeled as contexts or hyper-contexts which are combined with the global affinity into a unified optimization. The optimization is flexible for accommodating various types of contextual models. In multi-object tracking, the global affinity is defined according to the appearance similarity between objects detected in different frames. Interactions between objects are modeled as motion contexts which are encoded into the global association optimization. The tracking method integrates high order motion information and high order appearance variation. The multi-graph matching method carries out matching over graph vertices and structure matching over graph edges simultaneously. The matching consistency across multi-graphs is based on the high-order tensor optimization. Various types of vertex affinities and edge/hyper-edge affinities are flexibly integrated. Experiments on several public datasets, such as the MOT16 challenge benchmark, validate the effectiveness of the proposed methods. Weiming Hu 0004, Xinchu Shi, Zongwei Zhou, Junliang Xing, Haibin Ling, Stephen J. Maybank |
Int. J. Comput. Vis. | 6 |
| 2020 | Tracking-by-Fusion via Gaussian Process Regression Extended to Transfer LearningabstractThis paper presents a new Gaussian Processes (GPs)-based particle filter tracking framework. The framework non-trivially extends Gaussian process regression (GPR) to transfer learning, and, following the tracking-by-fusion strategy, integrates closely two tracking components, namely a GPs component and a CFs one. First, the GPs component analyzes and models the probability distribution of the object appearance by exploiting GPs. It categorizes the labeled samples into auxiliary and target ones, and explores unlabeled samples in transfer learning. The GPs component thus captures rich appearance information over object samples across time. On the other hand, to sample an initial particle set in regions of high likelihood through the direct simulation method in particle filtering, the powerful yet efficient correlation filters (CFs) are integrated, leading to the CFs component. In fact, the CFs component not only boosts the sampling quality, but also benefits from the GPs component, which provides re-weighted knowledge as latent variables for determining the impact of each correlation filter template from the auxiliary samples. In this way, the transfer learning based fusion enables effective interactions between the two components. Superior performance on four object tracking benchmarks (OTB-2015, Temple-Color, and VOT2015/2016), and in comparison with baselines and recent state-of-the-art trackers, has demonstrated clearly the effectiveness of the proposed framework. Qiang Wang 0051, Junliang Xing, Haibin Ling, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | Self-Taught Semisupervised Dictionary Learning With Nonnegative ConstraintabstractThis paper investigates classification by dictionary learning. A novel unified framework termed self-taught semisupervised dictionary learning with nonnegative constraint is proposed for simultaneously optimizing the components of a dictionary and a graph Laplacian. Specifically, an atom graph Laplacian regularization is built by using sparse coefficients to effectively capture the underlying manifold structure. It is more robust to noisy samples and outliers because atoms are more concise and representative than training samples. A nonnegative constraint imposed on the sparse coefficients guarantees that each sample is in the middle of its related atoms. In this way, the dependency between samples and atoms is made explicit. Furthermore, a self-taught mechanism is introduced to effectively feed back the manifold structure induced by atom graph Laplacian regularization and the supervised information hidden in unlabeled samples in order to learn a better dictionary. An efficient algorithm, combining a block coordinate descent method with the alternating direction method of multipliers, is derived to optimize the unified framework. Experimental results on several benchmark datasets show the effectiveness of the proposed model. Xiaoqin Zhang 0002, Di Wang 0008, Li Zhao 0005, Nannan Gu, Stephen J. Maybank |
IEEE Trans. Ind. Informatics | 6 |
| 2020 | Tangent Fisher Vector on Matrix Manifolds for Action RecognitionabstractIn this paper, we address the problem of representing and recognizing human actions from videos on matrix manifolds. For this purpose, we propose a new vector representation method, named tangent Fisher vector, to describe video sequences in the Fisher kernel framework. We first extract dense curved spatio-temporal cuboids from each video sequence. Compared with the traditional 'straight cuboids', the dense curved spatio-temporal cuboids contain much more local motion information. Each cuboid is then described using a linear dynamical system (LDS) to simultaneously capture the local appearance and dynamics. Furthermore, a simple yet efficient algorithm is proposed to learn the LDS parameters and approximate the observability matrix at the same time. Each video sequence is thus represented by a set of LDSs. Considering that each LDS can be viewed as a point in a Grassmann manifold, we propose to learn an intrinsic GMM on the manifold to cluster the LDS points. Finally a tangent Fisher vector is computed by first accumulating all the tangent vectors in each Gaussian component, and then concatenating the normalized results across all the Gaussian components. A kernel is defined to measure the similarity between tangent Fisher vectors for classification and recognition of a video sequence. This approach is evaluated on the state-of-the-art human action benchmark datasets. The recognition performance is competitive when compared with current state-of-the-art results. Guan Luo, Jiutong Wei, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Image Process. | 4 |
| 2020 | CDPM: Convolutional Deformable Part Models for Semantically Aligned Person Re-IdentificationabstractPart-level representations are essential for robust person re-identification. However, common errors that arise during pedestrian detection frequently result in severe misalignment problems for body parts, which degrade the quality of part representations. Accordingly, to deal with this problem, we propose a novel model named Convolutional Deformable Part Models (CDPM). CDPM works by decoupling the complex part alignment procedure into two easier steps: first, a vertical alignment step detects each body part in the vertical direction, with the help of a multi-task learning model; second, a horizontal refinement step based on attention suppresses the background information around each detected body part. Since these two steps are performed orthogonally and sequentially, the difficulty of part alignment is significantly reduced. In the testing stage, CDPM is able to accurately align flexible body parts without any need for outside information. Extensive experimental results demonstrate the effectiveness of the proposed CDPM for part alignment. Most impressively, CDPM achieves state-of-the-art performance on three large-scale datasets: Market-1501, DukeMTMC-ReID, and CUHK03. Kan Wang 0004, Changxing Ding, Stephen J. Maybank, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2020 | STA-CNN: Convolutional Spatial-Temporal Attention Learning for Action RecognitionabstractConvolutional Neural Networks have achieved excellent successes for object recognition in still images. However, the improvement of Convolutional Neural Networks over the traditional methods for recognizing actions in videos is not so significant, because the raw videos usually have much more redundant or irrelevant information than still images. In this paper, we propose a Spatial-Temporal Attentive Convolutional Neural Network (STA-CNN) which selects the discriminative temporal segments and focuses on the informative spatial regions automatically. The STA-CNN model incorporates a Temporal Attention Mechanism and a Spatial Attention Mechanism into a unified convolutional network to recognize actions in videos. The novel Temporal Attention Mechanism automatically mines the discriminative temporal segments from long and noisy videos. The Spatial Attention Mechanism firstly exploits the instantaneous motion information in optical flow features to locate the motion salient regions and it is then trained by an auxiliary classification loss with a Global Average Pooling layer to focus on the discriminative non-motion regions in the video frame. The STA-CNN model achieves the state-of-the-art performance on two of the most challenging datasets, UCF-101 (95.8%) and HMDB-51 (71.5%). Hao Yang 0010, Chunfeng Yuan, Li Zhang 0050, Yunda Sun, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Image Process. | 6 |
| 2020 | Anomaly Detection Using Local Kernel Density Estimation and Context-Based RegressionabstractCurrent local density-based anomaly detection methods are limited in that the local density estimation and the neighborhood density estimation are not accurate enough for complex and large databases, and the detection performance depends on the size parameter of the neighborhood. In this paper, we propose a new kernel function to estimate samples' local densities and propose a weighted neighborhood density estimation to increase the robustness to changes in the neighborhood size. We further propose a local kernel regression estimator and a hierarchical strategy for combining information from the multiple scale neighborhoods to refine anomaly factors of samples. We apply our general anomaly detection method to image saliency detection by regarding salient pixels in objects as anomalies to the background regions. Local density estimation in the visual feature space and kernel-based saliency score propagation in the image enable the assignment of similar saliency values to homogenous object regions. Experimental results on several benchmark datasets demonstrate that our anomaly detection methods overall outperform several state-of-art anomaly detection methods. The effectiveness of our image saliency detection method is validated by comparison with several state-of-art saliency detection methods. Weiming Hu 0004, Bing Li 0001, Ou Wu 0001, Junping Du 0001, Stephen J. Maybank |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2019 | The Fisher-Rao Metric in Computer VisionabstractThe Fisher-Rao metric is a Riemannian metric defined on any manifold that forms the parameter space for a family of probability distributions. The metric is specified by quadratic forms defined on the tangent spaces of the manifold. If a parameterisation of the manifold is chosen then each quadratic form is given by a symmetric positive definite matrix. Lengths, areas, volumes and hyper-volumes calculated using the Fisher-Rao metric are invariant under reparameterisation. This invariance is essential in practice because the parameterisation can be changed arbitrarily while keeping the data unchanged. The Fisher-Rao metric is obtained as a limit of the expected value of the log likelihood ratio for two nearby probability distributions. The inverse of the Fisher-Rao matrix is the Cramer-Rao lower bound on the covariance of an unbiased estimate of a parameter. The Fisher-Rao metric is used to divide the parameter space for the Hough transform method for detecting structures in data. Each division or accumulator is invariant under reparametrerisation, and the number of accumulators is proportional to the volume of the parameter space. Accurate approximations to the Fisher Rao metric are obtained for lines, catadioptric images of lines, circles, ellipses and the cross ratio. It is shown that the Fisher-Rao metric can be used to compare the amount of information in point features with the amount of information in edge element features. Stephen J. Maybank |
CIKM | 1 |
| 2019 | World From BlurabstractWhat can we tell from a single motion-blurred image? We show in this paper that a 3D scene can be revealed. Unlike prior methods that focus on producing a deblurred image, we propose to estimate and take advantage of the hidden message of a blurred image, the relative motion trajectory, to restore the 3D scene collapsed during the exposure process. To this end, we train a deep network that jointly predicts the motion trajectory, the deblurred image, and the depth one, all of which in turn form a collaborative and self-supervised cycle that supervise one another to reproduce the input blurred image, enabling plausible 3D scene reconstruction from a single blurred image. We test the proposed model on several large-scale datasets we constructed based on benchmarks, as well as real-world blurred images, and show that it yields very encouraging quantitative and qualitative results. Jiayan Qiu, Xinchao Wang, Stephen J. Maybank, Dacheng Tao |
CVPR | 3 |
| 2019 | Asymmetric 3D Convolutional Neural Networks for action recognition
Hao Yang 0010, Chunfeng Yuan, Bing Li 0001, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank |
Pattern Recognit. | 7 |
| 2018 | Deep Cost-Sensitive and Order-Preserving Feature Learning for Cross-Population Age EstimationabstractFacial age estimation from a face image is an important yet very challenging task in computer vision, since humans with different races and/or genders, exhibit quite different patterns in their facial aging processes. To deal with the influence of race and gender, previous methods perform age estimation within each population separately. In practice, however, it is often very difficult to collect and label sufficient data for each population. Therefore, it would be helpful to exploit an existing large labeled dataset of one (source) population to improve the age estimation performance on another (target) population with only a small labeled dataset available. In this work, we propose a Deep Cross-Population (DCP) age estimation model to achieve this goal. In particular, our DCP model develops a two-stage training strategy. First, a novel cost-sensitive multitask loss function is designed to learn transferable aging features by training on the source population. Second, a novel order-preserving pair-wise loss function is designed to align the aging features of the two populations. By doing so, our DCP model can transfer the knowledge encoded in the source population to the target population. Extensive experiments on the two of the largest benchmark datasets show that our DCP model outperforms several strong baseline methods and many state-of-the-art methods. Kai Li 0022, Junliang Xing, Chi Su, Weiming Hu 0004, Stephen J. Maybank |
CVPR | 6 |
| 2018 | Learning Attentions: Residual Attentional Siamese Network for High Performance Online Visual TrackingabstractOffline training for object tracking has recently shown great potentials in balancing tracking accuracy and speed. However, it is still difficult to adapt an offline trained model to a target tracked online. This work presents a Residual Attentional Siamese Network (RASNet) for high performance object tracking. The RASNet model reformulates the correlation filter within a Siamese tracking framework, and introduces different kinds of the attention mechanisms to adapt the model without updating the model online. In particular, by exploiting the offline trained general attention, the target adapted residual attention, and the channel favored feature attention, the RASNet not only mitigates the over-fitting problem in deep network training, but also enhances its discriminative capacity and adaptability due to the separation of representation learning and discriminator learning. The proposed deep architecture is trained from end to end and takes full advantage of the rich spatial temporal information to achieve robust visual tracking. Experimental results on two latest benchmarks, OTB-2015 and VOT2017, show that the RASNet tracker has the state-of-the-art tracking accuracy while runs at more than 80 frames per second. Qiang Wang 0051, Zhu Teng, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank |
CVPR | 6 |
| 2018 | Visual Tracking via Spatially Aligned Correlation Filters Network
Mengdan Zhang, Qiang Wang 0051, Junliang Xing, Peixi Peng, Weiming Hu 0004, Stephen J. Maybank |
ECCV (3) | 7 |
| 2018 | Do not Lose the Details: Reinforced Representation Learning for High Performance Visual TrackingabstractThis work presents a novel end-to-end trainable CNN model for high performance visual object tracking. It learns both low-level fine-grained representations and a high-level semantic embedding space in a mutual reinforced way, and a multi-task learning strategy is proposed to perform the correlation analysis on representations from both levels. In particular, a fully convolutional encoder-decoder network is designed to reconstruct the original visual features from the semantic projections to preserve all the geometric information. Moreover, the correlation filter layer working on the fine-grained representations leverages a global context constraint for accurate object appearance modeling. The correlation filter in this layer is updated online efficiently without network fine-tuning. Therefore, the proposed tracker benefits from two complementary effects: the adaptability of the fine-grained correlation analysis and the generalization capability of the semantic embedding. Extensive experimental evaluations on four popular benchmarks demonstrate its state-of-the-art performance. Qiang Wang 0051, Mengdan Zhang, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank |
IJCAI | 6 |
| 2018 | Visible and infrared image registration based on region features and edginess
Yanjia Chen, Xiuwei Zhang 0001, Yanning Zhang 0001, Stephen J. Maybank, Zhipeng Fu |
Mach. Vis. Appl. | 4 |
| 2018 | Dual Sticky Hierarchical Dirichlet Process Hidden Markov Model and Its Application to Natural Language Description of MotionsabstractIn this paper, a new nonparametric Bayesian model called the dual sticky hierarchical Dirichlet process hidden Markov model (HDP-HMM) is proposed for mining activities from a collection of time series data such as trajectories. All the time series data are clustered. Each cluster of time series data, corresponding to a motion pattern, is modeled by an HMM. Our model postulates a set of HMMs that share a common set of states (topics in an analogy with topic models for document processing), but have unique transition distributions. The number of HMMs and the number of topics are both automatically determined. The sticky prior avoids redundant states and makes our HDP-HMM more effective to model multimodal observations. For the application to motion trajectory modeling, topics correspond to motion activities. The learnt topics are clustered into atomic activities which are assigned predicates. We propose a Bayesian inference method to decompose a given trajectory into a sequence of atomic activities. The sources and sinks in the scene are learnt by clustering endpoints (origins and destinations) of trajectories. The semantic motion regions are learnt using the points in trajectories. On combining the learnt sources and sinks, the learnt semantic motion regions, and the learnt sequence of atomic activities, the action represented by a trajectory can be described in natural language in as automatic a way as possible. The effectiveness of our dual sticky HDP-HMM is validated on several trajectory datasets. The effectiveness of the natural language descriptions for motions is demonstrated on the vehicle trajectories extracted from a traffic scene. Weiming Hu 0004, Guodong Tian, Yongxin Kang, Chunfeng Yuan, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | Deep Constrained Siamese Hash Coding Network and Load-Balanced Locality-Sensitive Hashing for Near Duplicate Image DetectionabstractWe construct a new efficient near duplicate image detection method using a hierarchical hash code learning neural network and load-balanced locality-sensitive hashing (LSH) indexing. We propose a deep constrained siamese hash coding neural network combined with deep feature learning. Our neural network is able to extract effective features for near duplicate image detection. The extracted features are used to construct a LSH-based index. We propose a load-balanced LSH method to produce load-balanced buckets in the hashing process. The load-balanced LSH significantly reduces the query time. Based on the proposed load-balanced LSH, we design an effective and feasible algorithm for near duplicate image detection. Extensive experiments on three benchmark data sets demonstrate the effectiveness of our deep siamese hash encoding network and load-balanced LSH. Weiming Hu 0004, Yabo Fan, Junliang Xing, Zhaoquan Cai 0001, Stephen J. Maybank |
IEEE Trans. Image Process. | 6 |
| 2018 | Context-Dependent Random Walk Graph Kernels and Tree Pattern Graph Matching Kernels With Applications to Action RecognitionabstractGraphs are effective tools for modeling complex data. Setting out from two basic substructures, random walks and trees, we propose a new family of context-dependent random walk graph kernels and a new family of tree pattern graph matching kernels. In our context-dependent graph kernels, context information is incorporated into primary random walk groups. A multiple kernel learning algorithm with a proposed l1,2-norm regularization is applied to combine context-dependent graph kernels of different orders. This improves the similarity measurement between graphs. In our tree-pattern graph matching kernel, a quadratic optimization with a sparse constraint is proposed to select the correctly matched tree-pattern groups. This augments the discriminative power of the tree-pattern graph matching. We apply the proposed kernels to human action recognition, where each action is represented by two graphs which record the spatiotemporal relations between local feature vectors. Experimental comparisons with state-of-the-art algorithms on several benchmark datasets demonstrate the effectiveness of the proposed kernels for recognizing human actions. It is shown that our kernel based on tree-pattern groups, which have more complex structures and exploit more local topologies of graphs than random walks, yields more accurate results but requires more runtime than the context-dependent walk graph kernel. Weiming Hu 0004, Baoxin Wu, Chunfeng Yuan, Yangxi Li, Stephen J. Maybank |
IEEE Trans. Image Process. | 6 |
| 2017 | Spatio-Temporal Self-Organizing Map Deep Network for Dynamic Object Detection from Videos
Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
CVPR | 5 |
| 2017 | Supervised classification via constrained subspace and tensor sparse representationabstractSRC, a supervised classifier via sparse representation, has rapidly gained popularity in recent years and can be adapted to a wide range of applications based on the sparse solution of a linear system. First, we offer an intuitive geometric model called constrained subspace to explain the mechanism of SRC. The constrained subspace model connects the dots of NN, NFL, NS, NM. Then, inspired from the constrained subspace model, we extend SRC to its tensor-based variant, which takes as input samples of high-order tensors which are elements of an algebraic ring. A tensor sparse representation is used for query tensors. We verify in our experiments on several publicly available databases that the tensor-based SRC called tSRC outperforms traditional SRC in classification accuracy. Although demonstrated for image recognition, tSRC is easily adapted to other applications involving underdetermined linear systems. Stephen J. Maybank, Yanning Zhang 0001 |
IJCNN | 2 |
| 2017 | GRMA: Generalized Range Move Algorithms for the Efficient Optimization of MRFs
Junge Zhang, Peipei Yang, Stephen J. Maybank, Kaiqi Huang |
Int. J. Comput. Vis. | 4 |
| 2017 | Hyperspectral Image Spectral-Spatial Feature Extraction via Tensor Principal Component AnalysisabstractWe consider the tensor-based spectral-spatial feature extraction problem for hyperspectral image classification. First, a tensor framework based on circular convolution is proposed. Based on this framework, we extend the traditional principal component analysis (PCA) to its tensorial version tensor PCA (TPCA), which is applied to the spectral-spatial features of hyperspectral image data. The experiments show that the classification accuracy obtained using TPCA features is significantly higher than the accuracies obtained by its rivals. Yuemei Ren, Stephen J. Maybank, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Semi-Supervised Tensor-Based Graph Embedding Learning and Its Application to Visual Discriminant TrackingabstractAn appearance model adaptable to changes in object appearance is critical in visual object tracking. In this paper, we treat an image patch as a two-order tensor which preserves the original image structure. We design two graphs for characterizing the intrinsic local geometrical structure of the tensor samples of the object and the background. Graph embedding is used to reduce the dimensions of the tensors while preserving the structure of the graphs. Then, a discriminant embedding space is constructed. We prove two propositions for finding the transformation matrices which are used to map the original tensor samples to the tensor-based graph embedding space. In order to encode more discriminant information in the embedding space, we propose a transfer-learning- based semi-supervised strategy to iteratively adjust the embedding space into which discriminative information obtained from earlier times is transferred. We apply the proposed semi-supervised tensor-based graph embedding learning algorithm to visual tracking. The new tracking algorithm captures an object's appearance characteristics during tracking and uses a particle filter to estimate the optimal object state. Experimental results on the CVPR 2013 benchmark dataset demonstrate the effectiveness of the proposed tracking algorithm. Weiming Hu 0004, Junliang Xing, Chao Zhang 0089, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2017 | Multi-View Multi-Instance Learning Based on Joint Sparse Representation and Multi-View Dictionary LearningabstractIn multi-instance learning (MIL), the relations among instances in a bag convey important contextual information in many applications. Previous studies on MIL either ignore such relations or simply model them with a fixed graph structure so that the overall performance inevitably degrades in complex environments. To address this problem, this paper proposes a novel multi-view multi-instance learning algorithm (MIL) that combines multiple context structures in a bag into a unified framework. The novel aspects are: (i) we propose a sparse -graph model that can generate different graphs with different parameters to represent various context relations in a bag, (ii) we propose a multi-view joint sparse representation that integrates these graphs into a unified framework for bag classification, and (iii) we propose a multi-view dictionary learning algorithm to obtain a multi-view graph dictionary that considers cues from all views simultaneously to improve the discrimination of the MIL. Experiments and analyses in many practical applications prove the effectiveness of the M IL. Bing Li 0001, Chunfeng Yuan, Weihua Xiong, Weiming Hu 0004, Houwen Peng, Xinmiao Ding, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2017 | Algorithm-Dependent Generalization Bounds for Multi-Task LearningabstractOften, tasks are collected for multi-task learning (MTL) because they share similar feature structures. Based on this observation, in this paper, we present novel algorithm-dependent generalization bounds for MTL by exploiting the notion of algorithmic stability. We focus on the performance of one particular task and the average performance over multiple tasks by analyzing the generalization ability of a common parameter that is shared in MTL. When focusing on one particular task, with the help of a mild assumption on the feature structures, we interpret the function of the other tasks as a regularizer that produces a specific inductive bias. The algorithm for learning the common parameter, as well as the predictor, is thereby uniformly stable with respect to the domain of the particular task and has a generalization bound with a fast convergence rate of order O(1/n), where n is the sample size of the particular task. When focusing on the average performance over multiple tasks, we prove that a similar inductive bias exists under certain conditions on the feature structures. Thus, the corresponding algorithm for learning the common parameter is also uniformly stable with respect to the domains of the multiple tasks, and its generalization bound is of the order O(1/T), where T is the number of tasks. These theoretical analyses naturally show that the similarity of feature structures in MTL will lead to specific regularizations for predicting, which enables the learning algorithms to generalize fast and correctly from a few examples. Tongliang Liu, Dacheng Tao, Mingli Song, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Salient Object Detection via Structured Matrix DecompositionabstractLow-rank recovery models have shown potential for salient object detection, where a matrix is decomposed into a low-rank matrix representing image background and a sparse matrix identifying salient objects. Two deficiencies, however, still exist. First, previous work typically assumes the elements in the sparse matrix are mutually independent, ignoring the spatial and pattern relations of image regions. Second, when the low-rank and sparse matrices are relatively coherent, e.g., when there are similarities between the salient objects and background or when the background is complicated, it is difficult for previous models to disentangle them. To address these problems, we propose a novel structured matrix decomposition model with two structural regularizations: (1) a tree-structured sparsity-inducing regularization that captures the image structure and enforces patches from the same object to have similar saliency values, and (2) a Laplacian regularization that enlarges the gaps between salient objects and the background in feature space. Furthermore, high-level priors are integrated to guide the matrix decomposition and boost the detection. We evaluate our model for salient object detection on five challenging datasets including single object, multiple objects and complex scene images, and show competitive results as compared with 24 state-of-the-art methods in terms of seven performance metrics. Houwen Peng, Bing Li 0001, Haibin Ling, Weiming Hu 0004, Weihua Xiong, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2017 | D2C: Deep cumulatively and comparatively learning for human age estimation
Kai Li 0022, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank |
Pattern Recognit. | 4 |
| 2017 | Guest Editorial Introduction to the Special Issue on Large-Scale Video Analytics for Enhanced Security: Algorithms and SystemsabstractDue to the rapid increase of the number of cameras used in the video surveillance and the huge needs of the smart city and public security, video surveillance by human beings is no longer suitable. Hence, since the end of the last century, video analytics for security or visual surveillance has become one of the hottest research topics. Wide-area video surveillance systems can have extremely high data rates and high data volumes. Therefore, the challenge of video analytics is to extract meaningful information efficiently from the huge flow of video data in order to produce high-level semantic descriptions of the activities occurring in the area under surveillance. Kaiqi Huang, Tieniu Tan, Stephen J. Maybank, Rama Chellappa, Jake Aggarval |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2016 | Non-parametric Hidden Conditional Random Fields for action classificationabstractConditional Random Fields (CRF), a structured prediction method, combines probabilistic graphical models and discriminative classification techniques in order to predict class labels in sequence recognition problems. Its extension the Hidden Conditional Random Fields (HCRF) uses hidden state variables in order to capture intermediate structures. The number of hidden states in an HCRF must be specified a priori. This number is often not known in advance. A non-parametric extension to the HCRF, with the number of hidden states automatically inferred from data, is proposed here. This is a significant advantage over the classical HCRF since it avoids ad hoc model selection procedures. Further, the training and inference procedure is fully Bayesian eliminating the over fitting problem associated with frequentist methods. In particular, our construction is based on scale mixtures of Gaussians as priors over the HCRF parameters and makes use of Hierarchical Dirichlet Process (HDP) and Laplace distribution. The proposed inference procedure uses elliptical slice sampling, a Markov Chain Monte Carlo (MCMC) method, in order to sample optimal and sparse posterior HCRF parameters. The above technique is applied for classifying human actions that occur in depth image sequences - a challenging computer vision problem. Experiments with real world video datasets confirm the efficacy of our classification approach. Natraj Raman, Stephen J. Maybank |
IJCNN | 2 |
| 2016 | Fusing ℝ Features and Local Features with Context-Aware Kernels for Action Recognition
Chunfeng Yuan, Baoxin Wu, Xi Li 0001, Weiming Hu 0004, Stephen J. Maybank, Fangshi Wang |
Int. J. Comput. Vis. | 5 |
| 2016 | Activity recognition using a supervised non-parametric hierarchical HMM
Natraj Raman, Stephen J. Maybank |
Neurocomputing | 2 |
| 2016 | Handcrafted vs. learned representations for human action recognition
Xiantong Zhen, Ling Shao 0001, Stephen J. Maybank, Rama Chellappa |
Image Vis. Comput. | 3 |
| 2016 | Facial expression transfer method based on frequency analysis
Wei Wei 0008, Chunna Tian, Stephen J. Maybank, Yanning Zhang 0001 |
Pattern Recognit. | 3 |
| 2016 | Hierarchical aesthetic quality assessment using deep convolutional neural networks
Yueying Kao, Kaiqi Huang, Stephen J. Maybank |
Signal Process. Image Commun. | 3 |
| 2016 | Multi-Modal Curriculum Learning for Semi-Supervised Image ClassificationabstractSemi-supervised image classification aims to classify a large quantity of unlabeled images by typically harnessing scarce labeled images. Existing semi-supervised methods often suffer from inadequate classification accuracy when encountering difficult yet critical images, such as outliers, because they treat all unlabeled images equally and conduct classifications in an imperfectly ordered sequence. In this paper, we employ the curriculum learning methodology by investigating the difficulty of classifying every unlabeled image. The reliability and the discriminability of these unlabeled images are particularly investigated for evaluating their difficulty. As a result, an optimized image sequence is generated during the iterative propagations, and the unlabeled images are logically classified from simple to difficult. Furthermore, since images are usually characterized by multiple visual feature descriptors, we associate each kind of features with a teacher, and design a multi-modal curriculum learning (MMCL) strategy to integrate the information from different feature modalities. In each propagation, each teacher analyzes the difficulties of the currently unlabeled images from its own modality viewpoint. A consensus is subsequently reached among all the teachers, determining the currently simplest images (i.e., a curriculum), which are to be reliably classified by the multi-modal learner. This well-organized propagation process leveraging multiple teachers and one learner enables our MMCL to outperform five state-of-the-art methods on eight popular image data sets. Chen Gong 0002, Dacheng Tao, Stephen J. Maybank, Wei Liu 0005, Guoliang Kang, Jie Yang 0002 |
IEEE Trans. Image Process. | 3 |
| 2016 | Multi-Perspective Cost-Sensitive Context-Aware Multi-Instance Sparse Coding and Its Application to Sensitive Video RecognitionabstractWith the development of video-sharing websites, P2P, micro-blog, mobile WAP websites, and so on, sensitive videos can be more easily accessed. Effective sensitive video recognition is necessary for web content security. Among web sensitive videos, this paper focuses on violent and horror videos. Based on color emotion and color harmony theories, we extract visual emotional features from videos. A video is viewed as a bag and each shot in the video is represented by a key frame which is treated as an instance in the bag. Then, we combine multi-instance learning (MIL) with sparse coding to recognize violent and horror videos. The resulting MIL-based model can be updated online to adapt to changing web environments. We propose a cost-sensitive context-aware multi- instance sparse coding (MI-SC) method, in which the contextual structure of the key frames is modeled using a graph, and fusion between audio and visual features is carried out by extending the classic sparse coding into cost-sensitive sparse coding. We then propose a multi-perspective multi- instance joint sparse coding (MI-J-SC) method that handles each bag of instances from an independent perspective, a contextual perspective, and a holistic perspective. The experiments demonstrate that the features with an emotional meaning are effective for violent and horror video recognition, and our cost-sensitive context-aware MI-SC and multi-perspective MI-J-SC methods outperform the traditional MIL methods and the traditional SVM and KNN-based methods. Weiming Hu 0004, Xinmiao Ding, Bing Li 0001, Fangshi Wang, Stephen J. Maybank |
IEEE Trans. Multim. | 7 |
| 2015 | Saliency propagation from simple to difficultabstractSaliency propagation has been widely adopted for identifying the most attractive object in an image. The propagation sequence generated by existing saliency detection methods is governed by the spatial relationships of image regions, i.e., the saliency value is transmitted between two adjacent regions. However, for the inhomogeneous difficult adjacent regions, such a sequence may incur wrong propagations. In this paper, we attempt to manipulate the propagation sequence for optimizing the propagation quality. Intuitively, we postpone the propagations to difficult regions and meanwhile advance the propagations to less ambiguous simple regions. Inspired by the theoretical results in educational psychology, a novel propagation algorithm employing the teaching-to-learn and learning-to-teach strategies is proposed to explicitly improve the propagation quality. In the teaching-to-learn step, a teacher is designed to arrange the regions from simple to difficult and then assign the simplest regions to the learner. In the learning-to-teach step, the learner delivers its learning confidence to the teacher to assist the teacher to choose the subsequent simple regions. Due to the interactions between the teacher and learner, the uncertainty of original difficult regions is gradually reduced, yielding manifest salient objects with optimized background suppression. Extensive experimental results on benchmark saliency datasets demonstrate the superiority of the proposed algorithm over twelve representative saliency detectors. Chen Gong 0002, Dacheng Tao, Wei Liu 0005, Stephen J. Maybank, Keren Fu, Jie Yang 0002 |
CVPR | 4 |
| 2015 | A Robust Tracking System for Low Frame Rate Video
Xiaoqin Zhang 0002, Weiming Hu 0004, Nianhua Xie, Hujun Bao, Stephen J. Maybank |
Int. J. Comput. Vis. | 5 |
| 2015 | Erratum to: A Robust Tracking System for Low Frame Rate Video
Xiaoqin Zhang 0002, Weiming Hu 0004, Nianhua Xie, Hujun Bao, Stephen J. Maybank |
Int. J. Comput. Vis. | 5 |
| 2015 | Action classification using a discriminative multilevel HDP-HMM
Natraj Raman, Stephen J. Maybank |
Neurocomputing | 2 |
| 2015 | Robust hand tracking via novel multi-cue integration
Xiaoqin Zhang 0002, Wei Li 0034, Xiuzi Ye, Stephen J. Maybank |
Neurocomputing | 4 |
| 2015 | Stereo matching-based definition of saliency via sample-based Kullback-Leibler divergence estimation
Stephen J. Maybank, Yanning Zhang 0001 |
Mach. Vis. Appl. | 2 |
| 2015 | Single and Multiple Object Tracking Using a Multi-Feature Joint Sparse RepresentationabstractIn this paper, we propose a tracking algorithm based on a multi-feature joint sparse representation. The templates for the sparse representation can include pixel values, textures, and edges. In the multi-feature joint optimization, noise or occlusion is dealt with using a set of trivial templates. A sparse weight constraint is introduced to dynamically select the relevant templates from the full set of templates. A variance ratio measure is adopted to adaptively adjust the weights of different features. The multi-feature template set is updated adaptively. We further propose an algorithm for tracking multi-objects with occlusion handling based on the multi-feature joint sparse reconstruction. The observation model based on sparse reconstruction automatically focuses on the visible parts of an occluded object by using the information in the trivial templates. The multi-object tracking is simplified into a joint Bayesian inference. The experimental results show the superiority of our algorithm over several state-of-the-art tracking algorithms. Weiming Hu 0004, Wei Li 0034, Xiaoqin Zhang 0002, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Horror Image Recognition Based on Context-Aware Multi-Instance LearningabstractHorror content sharing on the Web is a growing phenomenon that can interfere with our daily life and affect the mental health of those involved. As an important form of expression, horror images have their own characteristics that can evoke extreme emotions. In this paper, we present a novel context-aware multi-instance learning (CMIL) algorithm for horror image recognition. The CMIL algorithm identifies horror images and picks out the regions that cause the sensation of horror in these horror images. It obtains contextual cues among adjacent regions in an image using a random walk on a contextual graph. Borrowing the strength of the fuzzy support vector machine (FSVM), we define a heuristic optimization procedure based on the FSVM to search for the optimal classifier for the CMIL. To improve the initialization of the CMIL, we propose a novel visual saliency model based on the tensor analysis. The average saliency value of each segmented region is set as its initial fuzzy membership in the CMIL. The advantage of the tensor-based visual saliency model is that it not only adaptively selects features, but also dynamically determines fusion weights for saliency value combination from different feature subspaces. The effectiveness of the proposed CMIL model is demonstrated by its use in horror image recognition on two large-scale image sets collected from the Internet. Bing Li 0001, Weihua Xiong, Ou Wu 0001, Weiming Hu 0004, Stephen J. Maybank, Shuicheng Yan |
IEEE Trans. Image Process. | 5 |
| 2015 | Large-Scale Weakly Supervised Object Localization via Latent Category LearningabstractLocalizing objects in cluttered backgrounds is challenging under large-scale weakly supervised conditions. Due to the cluttered image condition, objects usually have large ambiguity with backgrounds. Besides, there is also a lack of effective algorithm for large-scale weakly supervised localization in cluttered backgrounds. However, backgrounds contain useful latent information, e.g., the sky in the aeroplane class. If this latent information can be learned, object-background ambiguity can be largely reduced and background can be suppressed effectively. In this paper, we propose the latent category learning (LCL) in large-scale cluttered conditions. LCL is an unsupervised learning method which requires only image-level class labels. First, we use the latent semantic analysis with semantic object representation to learn the latent categories, which represent objects, object parts or backgrounds. Second, to determine which category contains the target object, we propose a category selection strategy by evaluating each category's discrimination. Finally, we propose the online LCL for use in large-scale conditions. Evaluation on the challenging PASCAL Visual Object Class (VOC) 2007 and the large-scale imagenet large-scale visual recognition challenge 2013 detection data sets shows that the method can improve the annotation precision by 10% over previous methods. More importantly, we achieve the detection precision which outperforms previous results by a large margin and can be competitive to the supervised deformable part model 5.0 baseline on both data sets. Kaiqi Huang, Weiqiang Ren, Junge Zhang, Stephen J. Maybank |
IEEE Trans. Image Process. | 5 |
| 2014 | A multi-modal moving object detection method based on GrowCut segmentationabstractCommonly-used motion detection methods, such as background subtraction, optical flow and frame subtraction are all based on the differences between consecutive image frames. There are many difficulties, including similarities between objects and background, shadows, low illumination, thermal halo. Visible light images and thermal images are complementary. Many difficulties in motion detection do not occur simultaneously in visible and thermal images. The proposed multimodal detection method combines the advantages of multi-modal image and GrowCut segmentation, overcomes the difficulties mentioned above and works well in complicated outdoor surveillance environments. Experiments showed our method yields better results than commonly-used fusion methods. Xiuwei Zhang 0001, Yanning Zhang 0001, Stephen J. Maybank |
CIMSIVP | 3 |
| 2014 | Image recognition via two-dimensional random projection and nearest constrained subspace
Yanning Zhang 0001, Stephen J. Maybank, Zhoufeng Liu |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Bin Ratio-Based Histogram Distances and Their Application to Image ClassificationabstractLarge variations in image background may cause partial matching and normalization problems for histogram-based representations, i.e., the histograms of the same category may have bins which are significantly different, and normalization may produce large changes in the differences between corresponding bins. In this paper, we deal with this problem by using the ratios between bin values of histograms, rather than bin values' differences which are used in the traditional histogram distances. We propose a bin ratio-based histogram distance (BRD), which is an intra-cross-bin distance, in contrast with previous bin-to-bin distances and cross-bin distances. The BRD is robust to partial matching and histogram normalization, and captures correlations between bins with only a linear computational complexity. We combine the BRD with the ℓ1 histogram distance and the χ(2) histogram distance to generate the ℓ1 BRD and the χ(2) BRD, respectively. These combinations exploit and benefit from the robustness of the BRD under partial matching and the robustness of the ℓ1 and χ(2) distances to small noise. We propose a method for assessing the robustness of histogram distances to partial matching. The BRDs and logistic regression-based histogram fusion are applied to image classification. The experimental results on synthetic data sets show the robustness of the BRDs to partial matching, and the experiments on seven benchmark data sets demonstrate promising results of the BRDs for image classification. Weiming Hu 0004, Nianhua Xie, Ruiguang Hu, Haibin Ling, Qiang Chen 0007, Shuicheng Yan, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2014 | Learning Human Actions by Combining Global Dynamics and Local AppearanceabstractIn this paper, we address the problem of human action recognition through combining global temporal dynamics and local visual spatio-temporal appearance features. For this purpose, in the global temporal dimension, we propose to model the motion dynamics with robust linear dynamical systems (LDSs) and use the model parameters as motion descriptors. Since LDSs live in a non-Euclidean space and the descriptors are in non-vector form, we propose a shift invariant subspace angles based distance to measure the similarity between LDSs. In the local visual dimension, we construct curved spatio-temporal cuboids along the trajectories of densely sampled feature points and describe them using histograms of oriented gradients (HOG). The distance between motion sequences is computed with the Chi-Squared histogram distance in the bag-of-words framework. Finally we perform classification using the maximum margin distance learning method by combining the global dynamic distances and the local visual distances. We evaluate our approach for action recognition on five short clips data sets, namely Weizmann, KTH, UCF sports, Hollywood2 and UCF50, as well as three long continuous data sets, namely VIRAT, ADL and CRIM13. We show competitive results as compared with current state-of-the-art methods. Guan Luo, Guodong Tian, Chunfeng Yuan, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2014 | Intrinsic dimension estimation via nearest constrained subspace classifier
Yanning Zhang 0001, Stephen J. Maybank, Zhoufeng Liu |
Pattern Recognit. | 3 |
| 2014 | Online Adaboost-Based Parameterized Methods for Dynamic Distributed Network Intrusion DetectionabstractCurrent network intrusion detection systems lack adaptability to the frequently changing network environments. Furthermore, intrusion detection in the new distributed architectures is now a major requirement. In this paper, we propose two online Adaboost-based intrusion detection algorithms. In the first algorithm, a traditional online Adaboost process is used where decision stumps are used as weak classifiers. In the second algorithm, an improved online Adaboost process is proposed, and online Gaussian mixture models (GMMs) are used as weak classifiers. We further propose a distributed intrusion detection framework, in which a local parameterized detection model is constructed in each node using the online Adaboost algorithm. A global detection model is constructed in each node by combining the local parametric models using a small number of samples in the node. This combination is achieved using an algorithm based on particle swarm optimization (PSO) and support vector machines. The global model in each node is used to detect intrusions. Experimental results show that the improved online Adaboost process with GMMs obtains a higher detection rate and a lower false alarm rate than the traditional online Adaboost process that uses decision stumps. Both the algorithms outperform existing intrusion detection algorithms. It is also shown that our PSO, and SVM-based algorithm effectively combines the local detection models into the global model in each node; the global model in a node can handle the intrusion types that are found in other nodes, without sharing the samples of these intrusion types. Weiming Hu 0004, Yanguo Wang, Ou Wu 0001, Stephen J. Maybank |
IEEE Trans. Cybern. | 5 |
| 2014 | Image Classification Using Multiscale Information Fusion Based on Saliency Driven Nonlinear Diffusion FilteringabstractIn this paper, we propose saliency driven image multiscale nonlinear diffusion filtering. The resulting scale space in general preserves or even enhances semantically important structures such as edges, lines, or flow-like structures in the foreground, and inhibits and smoothes clutter in the background. The image is classified using multiscale information fusion based on the original image, the image at the final scale at which the diffusion process converges, and the image at a midscale. Our algorithm emphasizes the foreground features, which are important for image classification. The background image regions, whether considered as contexts of the foreground or noise to the foreground, can be globally handled by fusing information from different scales. Experimental tests of the effectiveness of the multiscale space for the image classification are conducted on the following publicly available datasets: 1) the PASCAL 2005 dataset; 2) the Oxford 102 flowers dataset; and 3) the Oxford 17 flowers dataset, with high classification rates. Weiming Hu 0004, Ruiguang Hu, Nianhua Xie, Haibin Ling, Stephen J. Maybank |
IEEE Trans. Image Process. | 5 |
| 2014 | Robust 3D Face Landmark Localization Based on Local Coordinate CodingabstractIn the 3D facial animation and synthesis community, input faces are usually required to be labeled by a set of landmarks for parameterization. Because of the variations in pose, expression and resolution, automatic 3D face landmark localization remains a challenge. In this paper, a novel landmark localization approach is presented. The approach is based on local coordinate coding (LCC) and consists of two stages. In the first stage, we perform nose detection, relying on the fact that the nose shape is usually invariant under the variations in the pose, expression, and resolution. Then, we use the iterative closest points algorithm to find a 3D affine transformation that aligns the input face to a reference face. In the second stage, we perform resampling to build correspondences between the input 3D face and the training faces. Then, an LCC-based localization algorithm is proposed to obtain the positions of the landmarks in the input face. Experimental results show that the proposed method is comparable to state of the art methods in terms of its robustness, flexibility, and accuracy. Mingli Song, Dacheng Tao, Shengpeng Sun, Chun Chen 0001, Stephen J. Maybank |
IEEE Trans. Image Process. | 5 |
| 2014 | Modeling Geometric-Temporal Context With Directional Pyramid Co-Occurrence for Action RecognitionabstractIn this paper, we present a new geometric-temporal representation for visual action recognition based on local spatio-temporal features. First, we propose a modified covariance descriptor under the log-Euclidean Riemannian metric to represent the spatio-temporal cuboids detected in the video sequences. Compared with previously proposed covariance descriptors, our descriptor can be measured and clustered in Euclidian space. Second, to capture the geometric-temporal contextual information, we construct a directional pyramid co-occurrence matrix (DPCM) to describe the spatio-temporal distribution of the vector-quantized local feature descriptors extracted from a video. DPCM characterizes the co-occurrence statistics of local features as well as the spatio-temporal positional relationships among the concurrent features. These statistics provide strong descriptive power for action recognition. To use DPCM for action recognition, we propose a directional pyramid co-occurrence matching kernel to measure the similarity of videos. The proposed method achieves the state-of-the-art performance and improves on the recognition performance of the bag-of-visual-words (BOVWs) models by a large margin on six public data sets. For example, on the KTH data set, it achieves 98.78% accuracy while the BOVW approach only achieves 88.06%. On both Weizmann and UCF CIL data sets, the highest possible accuracy of 100% is achieved. Chunfeng Yuan, Xi Li 0001, Weiming Hu 0004, Haibin Ling, Stephen J. Maybank |
IEEE Trans. Image Process. | 5 |
| 2014 | Human Pose Estimation and Tracking via Parsing a Tree Structure Based Human ModelabstractHuman pose estimation and tracking is the task of determining the states (location, orientation, and scale) of each body part over time. It is important for many vision understanding applications, such as visual interactive gaming, immersive virtual reality, visual surveillance, and content-based image retrieval. However, it remains a challenging task due to unknown image background, presence of clutter and especially the high dimensional state space (usually 30+ dimensions). In this paper, we contribute to human pose estimation and tracking in two aspects. First, we design two efficient Markov Chain dynamics under the data-driven Markov Chain Monte Carlo framework to effectively explore the high dimensional state space. Second, we parse the tree structure state space into a lexicographic order according to the image observations and body topology, and the optimization process is conducted in this order. This realizes a much more efficient exploration of the state space than the sampling based search or exhaustive search, and thus achieves a tremendous speed-up. Experimental results demonstrate the efficiency and effectiveness of the proposed method in estimating and tracking various kinds of human poses, even against cluttered backgrounds, in poor illumination or under partial self-occlusion. Xiaoqin Zhang 0002, Weiming Hu 0004, Xiaofeng Tong, Stephen J. Maybank, Yimin Zhang 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2013 | 3D R Transform on Spatio-temporal Interest Points for Action RecognitionabstractSpatio-temporal interest points serve as an elementary building block in many modern action recognition algorithms, and most of them exploit the local spatio-temporal volume features using a Bag of Visual Words (BOVW) representation. Such representation, however, ignores potentially valuable information about the global spatio-temporal distribution of interest points. In this paper, we propose a new global feature to capture the detailed geometrical distribution of interest points. It is calculated by using the R transform which is defined as an extended 3D discrete Radon transform, followed by applying a two-directional two-dimensional principal component analysis. Such R feature captures the geometrical information of the interest points and keeps invariant to geometry transformation and robust to noise. In addition, we propose a new fusion strategy to combine the R feature with the BOVW representation for further improving recognition accuracy. We utilize a context-aware fusion method to capture both the pairwise similarities and higher-order contextual interactions of the videos. Experimental results on several publicly available datasets demonstrate the effectiveness of the proposed approach for action recognition. Chunfeng Yuan, Xi Li 0001, Weiming Hu 0004, Haibin Ling, Stephen J. Maybank |
CVPR | 5 |
| 2013 | Discriminant Tracking Using Tensor Representation with Semi-supervised ImprovementabstractVisual tracking has witnessed growing methods in object representation, which is crucial to robust tracking. The dominant mechanism in object representation is using image features encoded in a vector as observations to perform tracking, without considering that an image is intrinsically a matrix, or a 2^nd-order tensor. Thus approaches following this mechanism inevitably lose a lot of useful information, and therefore cannot fully exploit the spatial correlations within the 2D image ensembles. In this paper, we address an image as a 2^nd-order tensor in its original form, and find a discriminative linear embedding space approximation to the original nonlinear sub manifold embedded in the tensor space based on the graph embedding framework. We specially design two graphs for characterizing the intrinsic local geometrical structure of the tensor space, so as to retain more discriminant information when reducing the dimension along certain tensor dimensions. However, spatial correlations within a tensor are not limited to the elements along these dimensions. This means that some part of the discriminant information may not be encoded in the embedding space. We introduce a novel technique called semi-supervised improvement to iteratively adjust the embedding space to compensate for the loss of discriminant information, hence improving the performance of our tracker. Experimental results on challenging videos demonstrate the effectiveness and robustness of the proposed tracker. Junliang Xing, Weiming Hu 0004, Stephen J. Maybank |
ICCV | 4 |
| 2013 | Action classification using a discriminative non-parametric Hidden Markov ModelabstractWe classify human actions occurring in videos, using the skeletal joint positions extracted from a depth image sequence as features. Each action class is represented by a non-parametric Hidden Markov Model (NP-HMM) and the model parameters are learnt in a discriminative way. Specifically, we use a Bayesian framework based on Hierarchical Dirichlet Process (HDP) to automatically infer the cardinality of hidden states and formulate a discriminative function based on distance between Gaussian distributions to improve classification performance. We use elliptical slice sampling to efficiently sample parameters from the complex posterior distribution induced by our discriminative likelihood function. We illustrate our classification results for action class models trained using this technique. Natraj Raman, Stephen J. Maybank, Dell Zhang |
ICMV | 2 |
| 2013 | An Improved Hierarchical Dirichlet Process-Hidden Markov Model and Its Application to Trajectory Modeling and Retrieval
Weiming Hu 0004, Guodong Tian, Xi Li 0001, Stephen J. Maybank |
Int. J. Comput. Vis. | 4 |
| 2013 | Dimension estimation of image manifolds by minimal cover approximation
Mingyu Fan, Xiaoqin Zhang 0002, Shengyong Chen, Hujun Bao, Stephen J. Maybank |
Neurocomputing | 5 |
| 2013 | A probabilistic definition of salient regions for image matching
Stephen J. Maybank |
Neurocomputing | 1 |
| 2013 | An IR and visible image sequence automatic registration method based on optical flow
Yanning Zhang 0001, Xiuwei Zhang 0001, Stephen J. Maybank |
Mach. Vis. Appl. | 3 |
| 2013 | An Incremental DPMM-Based Method for Trajectory Clustering, Modeling, and RetrievalabstractTrajectory analysis is the basis for many applications, such as indexing of motion events in videos, activity recognition, and surveillance. In this paper, the Dirichlet process mixture model (DPMM) is applied to trajectory clustering, modeling, and retrieval. We propose an incremental version of a DPMM-based clustering algorithm and apply it to cluster trajectories. An appropriate number of trajectory clusters is determined automatically. When trajectories belonging to new clusters arrive, the new clusters can be identified online and added to the model without any retraining using the previous data. A time-sensitive Dirichlet process mixture model (tDPMM) is applied to each trajectory cluster for learning the trajectory pattern which represents the time-series characteristics of the trajectories in the cluster. Then, a parameterized index is constructed for each cluster. A novel likelihood estimation algorithm for the tDPMM is proposed, and a trajectory-based video retrieval model is developed. The tDPMM-based probabilistic matching method and the DPMM-based model growing method are combined to make the retrieval model scalable and adaptable. Experimental comparisons with state-of-the-art algorithms demonstrate the effectiveness of our algorithm. Weiming Hu 0004, Xi Li 0001, Guodong Tian, Stephen J. Maybank, Zhongfei Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2013 | Block covariance based l1 tracker with a subtle template dictionary
Xiaoqin Zhang 0002, Wei Li 0034, Weiming Hu 0004, Haibin Ling, Stephen J. Maybank |
Pattern Recognit. | 5 |
| 2013 | Robust Head Tracking Based on Multiple Cues Fusion in the Kernel-Bayesian FrameworkabstractThis paper presents a robust head tracking algorithm based on multiple cues fusion in a kernel-Bayesian framework. In this algorithm, the object to be tracked is characterized using a spatial-constraint mixture of the Gaussians-based appearance model and a multichannel chamfer matching-based shape model. These two models complement each other and their combination is discriminative in distinguishing the object from the background. A selective updating technique for the appearance model is employed to accommodate appearance and illumination changes. Meantime, the kernel method-mean shift algorithm is embedded into the Bayesian framework to give a heuristic prediction in the hypotheses generation process. This alleviates the great computational load suffered by conventional Bayesian trackers. Experimental results demonstrate that the proposed algorithm is effective. Xiaoqin Zhang 0002, Weiming Hu 0004, Hujun Bao, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2013 | Active Contour-Based Visual Tracking by Integrating Colors, Shapes, and MotionsabstractIn this paper, we present a framework for active contour-based visual tracking using level sets. The main components of our framework include contour-based tracking initialization, color-based contour evolution, adaptive shape-based contour evolution for non-periodic motions, dynamic shape-based contour evolution for periodic motions, and the handling of abrupt motions. For the initialization of contour-based tracking, we develop an optical flow-based algorithm for automatically initializing contours at the first frame. For the color-based contour evolution, Markov random field theory is used to measure correlations between values of neighboring pixels for posterior probability estimation. For adaptive shape-based contour evolution, the global shape information and the local color information are combined to hierarchically evolve the contour, and a flexible shape updating model is constructed. For the dynamic shape-based contour evolution, a shape mode transition matrix is learnt to characterize the temporal correlations of object shapes. For the handling of abrupt motions, particle swarm optimization is adopted to capture the global motion which is applied to the contour in the current frame to produce an initial contour in the next frame. Weiming Hu 0004, Wei Li 0034, Wenhan Luo, Xiaoqin Zhang 0002, Stephen J. Maybank |
IEEE Trans. Image Process. | 6 |
| 2013 | Manifold Regularized Multitask Learning for Semi-Supervised Multilabel Image ClassificationabstractIt is a significant challenge to classify images with multiple labels by using only a small number of labeled samples. One option is to learn a binary classifier for each label and use manifold regularization to improve the classification performance by exploring the underlying geometric structure of the data distribution. However, such an approach does not perform well in practice when images from multiple concepts are represented by high-dimensional visual features. Thus, manifold regularization is insufficient to control the model complexity. In this paper, we propose a manifold regularized multitask learning (MRMTL) algorithm. MRMTL learns a discriminative subspace shared by multiple classification tasks by exploiting the common structure of these tasks. It effectively controls the model complexity because different tasks limit one another's search volume, and the manifold regularization ensures that the functions in the shared hypothesis space are smooth along the data manifold. We conduct extensive experiments, on the PASCAL VOC'07 dataset with 20 classes and the MIR dataset with 38 classes, by comparing MRMTL with popular image classification algorithms. The results suggest that MRMTL is effective for image classification. Yong Luo 0002, Dacheng Tao, Bo Geng, Chao Xu 0006, Stephen J. Maybank |
IEEE Trans. Image Process. | 5 |
| 2012 | A Fisher-Rao Metric for Paracatadioptric Images of Lines
Stephen J. Maybank, Sio-Hoi Ieng, Ryad Benosman |
Int. J. Comput. Vis. | 1 |
| 2012 | Single and Multiple Object Tracking Using Log-Euclidean Riemannian Subspace and Block-Division Appearance ModelabstractObject appearance modeling is crucial for tracking objects, especially in videos captured by nonstationary cameras and for reasoning about occlusions between multiple moving objects. Based on the log-euclidean Riemannian metric on symmetric positive definite matrices, we propose an incremental log-euclidean Riemannian subspace learning algorithm in which covariance matrices of image features are mapped into a vector space with the log-euclidean Riemannian metric. Based on the subspace learning algorithm, we develop a log-euclidean block-division appearance model which captures both the global and local spatial layout information about object appearances. Single object tracking and multi-object tracking with occlusion reasoning are then achieved by particle filtering-based Bayesian state inference. During tracking, incremental updating of the log-euclidean block-division appearance model captures changes in object appearance. For multi-object tracking, the appearance models of the objects can be updated even in the presence of occlusions. Experimental results demonstrate that the proposed tracking algorithm obtains more accurate results than six state-of-the-art tracking algorithms. Weiming Hu 0004, Xi Li 0001, Wenhan Luo, Xiaoqin Zhang 0002, Stephen J. Maybank, Zhongfei Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2012 | Efficient Clustering Aggregation Based on Data FragmentsabstractClustering aggregation, known as clustering ensembles, has emerged as a powerful technique for combining different clustering results to obtain a single better clustering. Existing clustering aggregation algorithms are applied directly to data points, in what is referred to as the point-based approach. The algorithms are inefficient if the number of data points is large. We define an efficient approach for clustering aggregation based on data fragments. In this fragment-based approach, a data fragment is any subset of the data that is not split by any of the clustering results. To establish the theoretical bases of the proposed approach, we prove that clustering aggregation can be performed directly on data fragments under two widely used goodness measures for clustering aggregation taken from the literature. Three new clustering aggregation algorithms are described. The experimental results obtained using several public data sets show that the new algorithms have lower computational complexity than three well-known existing point-based clustering aggregation algorithms (Agglomerative, Furthest, and LocalSearch); nevertheless, the new algorithms do not sacrifice the accuracy. Ou Wu 0001, Weiming Hu 0004, Stephen J. Maybank, Mingliang Zhu, Bing Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2011 | Shared feature extraction for semi-supervised image classificationabstractMulti-task learning (MTL) plays an important role in image analysis applications, e.g. image classification, face recognition and image annotation. That is because MTL can estimate the latent shared subspace to represent the common features given a set of images from different tasks. However, the geometry of the data probability distribution is always supported on an intrinsic image sub-manifold that is embedded in a high dimensional Euclidean space. Therefore, it is improper to directly apply MTL to multiclass image classification. In this paper, we propose a manifold regularized MTL (MRMTL) algorithm to discover the latent shared subspace by treating the high-dimensional image space as a sub-manifold embedded in an ambient space. We conduct experiments on the PASCAL VOC'07 dataset with 20 classes and the MIR dataset with 38 classes by comparing MRMTL with conventional MTL and several representative image classification algorithms. The results suggest that MRMTL can properly extract the common features for image representation and thus improve the generalization performance of the image classification models. Yong Luo 0002, Dacheng Tao, Bo Geng, Chao Xu 0006, Stephen J. Maybank |
ACM Multimedia | 5 |
| 2011 | Incremental Tensor Subspace Learning and Its Applications to Foreground Segmentation and TrackingabstractAppearance modeling is very important for background modeling and object tracking. Subspace learning-based algorithms have been used to model the appearances of objects or scenes. Current vector subspace-based algorithms cannot effectively represent spatial correlations between pixel values. Current tensor subspace-based algorithms construct an offline representation of image ensembles, and current online tensor subspace learning algorithms cannot be applied to background modeling and object tracking. In this paper, we propose an online tensor subspace learning algorithm which models appearance changes by incrementally learning a tensor subspace representation through adaptively updating the sample mean and an eigenbasis for each unfolding matrix of the tensor. The proposed incremental tensor subspace learning algorithm is applied to foreground segmentation and object tracking for grayscale and color image sequences. The new background models capture the intrinsic spatiotemporal characteristics of scenes. The new tracking algorithm captures the appearance characteristics of an object during tracking and uses a particle filter to estimate the optimal object state. Experimental evaluations against state-of-the-art algorithms demonstrate the promise and effectiveness of the proposed incremental tensor subspace learning algorithm, and its applications to foreground segmentation and object tracking. Weiming Hu 0004, Xi Li 0001, Xiaoqin Zhang 0002, Xinchu Shi, Stephen J. Maybank, Zhongfei Zhang |
Int. J. Comput. Vis. | 5 |
| 2011 | Visual tracking via dynamic tensor analysis with mean update
Xiaoqin Zhang 0002, Xinchu Shi, Weiming Hu 0004, Xi Li 0001, Stephen J. Maybank |
Neurocomputing | 5 |
| 2011 | A Survey on Visual Content-Based Video Indexing and RetrievalabstractVideo indexing and retrieval have a wide spectrum of promising applications, motivating the interest of researchers worldwide. This paper offers a tutorial and an overview of the landscape of general strategies in visual content-based video indexing and retrieval, focusing on methods for video structure analysis, including shot boundary detection, key frame extraction and scene segmentation, extraction of features including static key frame features, object features and motion features, video data mining, video annotation, video retrieval including query interfaces, similarity measure and relevance feedback, and video browsing. Finally, we analyze future research directions. Weiming Hu 0004, Nianhua Xie, Li Li 0010, Xianglin Zeng, Stephen J. Maybank |
IEEE Trans. Syst. Man Cybern. Part C | 5 |
| 2010 | Multiple Object Tracking Via Species-Based Particle Swarm OptimizationabstractMultiple object tracking is particularly challenging when many objects with similar appearances occlude one another. Most existing approaches concatenate the states of different objects, view the multi-object tracking as a joint motion estimation problem and search for the best state of the joint motion in a rather high dimensional space. However, this centralized framework suffers from a high computational load. We bring a new view to the tracking problem from a swarm intelligence perspective. In analogy with the foraging behavior of bird flocks, we propose a species-based particle swarm optimization algorithm for multiple object tracking, in which the global swarm is divided into many species according to the number of objects, and each species searches for its object and maintains track of it. The interaction between different objects is modeled as species competition and repulsion, and the occlusion relationship is implicitly deduced from the “power” of each species, which is a function of the image observations. Therefore, our approach decentralizes the joint tracker to a set of individual trackers, each of which tries to maximize its visual evidence. Experimental results demonstrate the efficiency and effectiveness of our method. Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | Human Action Recognition under Log-Euclidean Riemannian Metric
Chunfeng Yuan, Weiming Hu 0004, Xi Li 0001, Stephen J. Maybank, Guan Luo |
ACCV (1) | 4 |
| 2009 | A Smarter Particle Filter
Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank |
ACCV (2) | 3 |
| 2009 | Domain Transfer SVM for video concept detectionabstractCross-domain learning methods have shown promising results by leveraging labeled patterns from auxiliary domains to learn a robust classifier for target domain, which has a limited number of labeled samples. To cope with the tremendous change of feature distribution between different domains in video concept detection, we propose a new cross-domain kernel learning method. Our method, referred to as Domain Transfer SVM (DTSVM), simultaneously learns a kernel function and a robust SVM classifier by minimizing both the structural risk functional of SVM and the distribution mismatch of labeled and unlabeled samples between the auxiliary and target domains. Comprehensive experiments on the challenging TRECVID corpus demonstrate that DTSVM outperforms existing cross-domain learning and multiple kernel learning methods. Lixin Duan, Ivor W. Tsang, Dong Xu 0001, Stephen J. Maybank |
CVPR | 4 |
| 2009 | Efficient human pose estimation via parsing a tree structure based human modelabstractHuman pose estimation is the task of determining the states (location, orientation and scale) of each body part. It is important for many vision understanding applications, e.g. visual interactive gaming, immersive virtual reality, content-based image retrieval, etc. However, it remains a challenging task because of unknown image background, presence of clutter, partial occlusion and especially the high dimensional state space (usually 30+ dimensions). In this paper, we contribute to human pose estimation in two aspects. First, we design two efficient Markov Chain dynamics under the data-driven Markov Chain Monte Carlo (DDMCMC) framework to effectively explore the complex solution space. Second, we parse the tree structure state space into a lexicographic order according to the image observations and body topology, and the optimization process is conducted in this order. This realizes a much more efficient exploration than the sampling based search and exhaustive search, and thus achieves a tremendous speed-up. Experimental results demonstrate the efficiency and effectiveness of the proposed method in estimating various kinds of human poses, even with cluttered background , poor illumination or partial self-occlusion. Xiaoqin Zhang 0002, Xiaofeng Tong, Weiming Hu 0004, Stephen J. Maybank, Yimin Zhang 0002 |
ICCV | 5 |
| 2009 | Retrieval based interactive cartoon synthesis via unsupervised bi-distance metric learningabstractCartoons play important roles in many areas, but it requires a lot of labor to produce new cartoon clips. In this paper, we propose a gesture recognition method for cartoon character images with two applications, namely content-based cartoon image retrieval and cartoon clip synthesis. We first define Edge Features (EF) and Motion Direction Features (MDF) for cartoon character images. The features are classified into two different groups, namely intra-features and inter-features. An Unsupervised Bi-Distance Metric Learning (UBDML) algorithm is proposed to recognize the gestures of cartoon character images. Different from the previous research efforts on distance metric learning, UBDML learns the optimal distance metric from the heterogeneous distance metrics derived from intra-features and inter-features. Content-based cartoon character image retrieval and cartoon clip synthesis can be carried out based on the distance metric learned by UBDML. Experiments show that the cartoon character image retrieval has a high precision and that the cartoon clip synthesis can be carried out efficiently. Yi Yang 0001, Yueting Zhuang, Dong Xu 0001, Yunhe Pan, Dacheng Tao, Stephen J. Maybank |
ACM Multimedia | 6 |
| 2009 | Geometric Mean for Subspace SelectionabstractSubspace selection approaches are powerful tools in pattern classification and data visualization. One of the most important subspace approaches is the linear dimensionality reduction step in the Fisher's linear discriminant analysis (FLDA), which has been successfully employed in many fields such as biometrics, bioinformatics, and multimedia information management. However, the linear dimensionality reduction step in FLDA has a critical drawback: for a classification task with c classes, if the dimension of the projected subspace is strictly lower than c - 1, the projection to a subspace tends to merge those classes, which are close together in the original feature space. If separate classes are sampled from Gaussian distributions, all with identical covariance matrices, then the linear dimensionality reduction step in FLDA maximizes the mean value of the Kullback-Leibler (KL) divergences between different classes. Based on this viewpoint, the geometric mean for subspace selection is studied in this paper. Three criteria are analyzed: 1) maximization of the geometric mean of the KL divergences, 2) maximization of the geometric mean of the normalized KL divergences, and 3) the combination of 1 and 2. Preliminary experimental results based on synthetic data, UCI Machine Learning Repository, and handwriting digits show that the third criterion is a potential discriminative subspace selection method, which significantly reduces the class separation problem in comparing with the linear dimensionality reduction step in FLDA and its several representative extensions. Dacheng Tao, Xuelong Li 0001, Xindong Wu 0001, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2009 | Occlusion Reasoning for Tracking Multiple PeopleabstractOcclusion reasoning is one of the most challenging issues in visual surveillance. In this letter, we propose a new approach for reasoning about occlusions between multiple people. In our approach, occlusion relationships between people are explicitly defined and deduction of the occlusion relationships is integrated into the whole tracking framework. The prior knowledge is supplied by a set of models which include a 2-D elliptical shape model, a spatial-color mixture of Gaussians appearance model, and a motion model with constant velocity. An observation likelihood function is constructed based on the similarity between the observations and the object appearance models with given states. The occlusion relationships are deduced from the current states of the objects and the current observations, using the observation likelihood function. The previous occlusion relationships are not required for deducing the current occlusion relationships. The problem of tracking and occlusion reasoning for more than two people is formulated mathematically, and a solution is proposed based on particle filtering. Experimental results on several real video sequences from indoor and outdoor scenes show the effectiveness of our approach. Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | Human Behavior Analysis Based on a New Motion DescriptorabstractHuman behavior analysis is an important area of research in computer vision and is also driven by a wide spectrum of applications, such as smart video surveillance and human-computer interface. In this paper, we present a novel approach for human behavior analysis. Two research challenges, motion representation and behavior recognition, are addressed. A novel motion descriptor, which is an improved feature based on optical flow, is proposed for motion representation. Optical flow is improved with a motion filter, and feature fusion with the shape and trajectory information. To recognize the behavior, the support vector machine is employed to train the classifier where the concatenation of histograms is formed as the input features. Experimental results on the Weizmann behavior database and the Institute of Automation, Chinese Academy of Science real-world multiview behavior database demonstrate the robustness and effectiveness of our method. Kaiqi Huang, Shiquan Wang, Tieniu Tan, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | Unsupervised Active Learning Based on Hierarchical Graph-Theoretic ClusteringabstractMost existing active learning approaches are supervised. Supervised active learning has the following problems: inefficiency in dealing with the semantic gap between the distribution of samples in the feature space and their labels, lack of ability in selecting new samples that belong to new categories that have not yet appeared in the training samples, and lack of adaptability to changes in the semantic interpretation of sample categories. To tackle these problems, we propose an unsupervised active learning framework based on hierarchical graph-theoretic clustering. In the framework, two promising graph-theoretic clustering algorithms, namely, dominant-set clustering and spectral clustering, are combined in a hierarchical fashion. Our framework has some advantages, such as ease of implementation, flexibility in architecture, and adaptability to changes in the labeling. Evaluations on data sets for network intrusion detection, image classification, and video classification have demonstrated that our active learning framework can effectively reduce the workload of manual classification while maintaining a high accuracy of automatic classification. It is shown that, overall, our framework outperforms the support-vector-machine-based supervised active learning, particularly in terms of dealing much more efficiently with new samples whose categories have not yet appeared in the training samples. Weiming Hu 0004, Nianhua Xie, Stephen J. Maybank |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2008 | Sequential particle swarm optimization for visual trackingabstractVisual tracking usually involves an optimization process for estimating the motion of an object from measured images in a video sequence. In this paper, a new evolutionary approach, PSO (particle swarm optimization), is adopted for visual tracking. Since the tracking process is a dynamic optimization problem which is simultaneously influenced by the object state and the time, we propose a sequential particle swarm optimization framework by incorporating the temporal continuity information into the traditional PSO algorithm. In addition, the parameters in PSO are changed adaptively according to the fitness values of particles and the predicted motion of the tracked object, leading to a favourable performance in tracking applications. Furthermore, we show theoretically that, in a Bayesian inference view, the sequential PSO framework is in essence a multilayer importance sampling based particle filter. Experimental results demonstrate that, compared with the state-of-the-art particle filter and its variation - the unscented particle filter, the proposed tracking algorithm is more robust and effective, especially when the object has an arbitrary motion or undergoes large appearance changes. Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank, Xi Li 0001, Mingliang Zhu |
CVPR | 3 |
| 2008 | Bayesian tensor analysisabstractVector data are normally used for probabilistic graphical models with Bayesian inference. However, tensor data, i.e., multidimensional arrays, are actually natural representations of a large amount of real data, in data mining, computer vision, and many other applications. Aiming at breaking the huge gap between vectors and tensors in conventional statistical tasks, e.g., automatic model selection, this paper proposes a decoupled probabilistic algorithm, named Bayesian tensor analysis (BTA). BTA automatically selects a suitable model for tensor data, as demonstrated by empirical studies. Dacheng Tao, Jimeng Sun 0001, Jialie Shen 0001, Xindong Wu 0001, Xuelong Li 0001, Stephen J. Maybank, Christos Faloutsos |
IJCNN | 6 |
| 2008 | Visual music and musical vision
Xuelong Li 0001, Dacheng Tao, Stephen J. Maybank, Yuan Yuan 0001 |
Neurocomputing | 3 |
| 2008 | Approximation to the Fisher-Rao metric for the focus of expansion
Stephen J. Maybank |
Neurocomputing | 1 |
| 2008 | Tensor Rank One Discriminant Analysis - A convergent method for discriminative multilinear subspace selection
Dacheng Tao, Xuelong Li 0001, Xindong Wu 0001, Stephen J. Maybank |
Neurocomputing | 4 |
| 2008 | A real-time object detecting and tracking system for outdoor night surveillance
Kaiqi Huang, Liangsheng Wang, Tieniu Tan, Stephen J. Maybank |
Pattern Recognit. | 4 |
| 2008 | An Introduction to the Special Issue on Event Analysis in VideosabstractI NTEREST from industry and academia has increased dramatically over recent years in the challenging area of event analysis and recognition from various video sources including sports, surveillance, user-generated video, etc. Video event analysis and recognition is a critical task in many applications such as detection of sporting highlights, incident detection in surveillance video, indexing, retrieval and summarization of video databases, and human-computer interaction. This special issue aims to capture the latest advances by the research community working in the area of video event analysis. The call for papers was enthusiastically greeted by the research community and we received over seventy submissions. The special issue presents 16 articles which provide fundamental contributions in a wide range of topics in video event analysis: 1) human action and activity recognition; 2) motion trajectory analysis; 3) video content analysis and pattern mining; 4) audio-visual multi-modal analysis; and 5) video analysis applications. An overview of the organization and a brief summary of the articles selected for publication in the special issue are provided below. Shih-Fu Chang, Jiebo Luo 0001, Stephen J. Maybank, Dan Schonfeld, Dong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | Bayesian Tensor Approach for 3-D Face ModelingabstractEffectively modeling a collection of three-dimensional (3-D) faces is an important task in various applications, especially facial expression-driven ones, e.g., expression generation, retargeting, and synthesis. These 3-D faces naturally form a set of second-order tensors-one modality for identity and the other for expression. The number of these second-order tensors is three times of that of the vertices for 3-D face modeling. As for algorithms, Bayesian data modeling, which is a natural data analysis tool, has been widely applied with great success; however, it works only for vector data. Therefore, there is a gap between tensor-based representation and vector-based data analysis tools. Aiming at bridging this gap and generalizing conventional statistical tools over tensors, this paper proposes a decoupled probabilistic algorithm, which is named Bayesian tensor analysis (BTA). Theoretically, BTA can automatically and suitably determine dimensionality for different modalities of tensor data. With BTA, a collection of 3-D faces can be well modeled. Empirical studies on expression retargeting also justify the advantages of BTA. Dacheng Tao, Mingli Song, Xuelong Li 0001, Jialie Shen 0001, Jimeng Sun 0001, Xindong Wu 0001, Christos Faloutsos, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2008 | AdaBoost-Based Algorithm for Network Intrusion DetectionabstractNetwork intrusion detection aims at distinguishing the attacks on the Internet from normal use of the Internet. It is an indispensable part of the information security system. Due to the variety of network behaviors and the rapid development of attack fashions, it is necessary to develop fast machine-learning-based intrusion detection algorithms with high detection rates and low false-alarm rates. In this correspondence, we propose an intrusion detection algorithm based on the AdaBoost algorithm. In the algorithm, decision stumps are used as weak classifiers. The decision rules are provided for both categorical and continuous features. By combining the weak classifiers for continuous features and the weak classifiers for categorical features into a strong classifier, the relations between these two different types of features are handled naturally, without any forced conversions between continuous and categorical features. Adaptable initial weights and a simple strategy for avoiding overfitting are adopted to improve the performance of the algorithm. Experimental results show that our algorithm has low computational complexity and error rates, as compared with algorithms of higher computational complexity, as tested on the benchmark sample data. Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2008 | Gait Components and Their Application to Gender RecognitionabstractHuman gait is a promising biometrics resource. In this paper, the information about gait is obtained from the motions of the different parts of the silhouette. The human silhouette is segmented into seven components, namely head, arm, trunk, thigh, front-leg, back-leg, and feet. The leg silhouettes for the front-leg and the back-leg are considered separately because, during walking, the left leg and the right leg are in front or at the back by turns. Each of the seven components and a number of combinations of the components are then studied with regard to two useful applications: human identification (ID) recognition and gender recognition. More than 500 different experiments on human ID and gender recognition are carried out under a wide range of circumstances. The effectiveness of the seven human gait components for ID and gender recognition is analyzed. Xuelong Li 0001, Stephen J. Maybank, Shuicheng Yan, Dacheng Tao, Dong Xu 0001 |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2007 | Kernel-Bayesian Framework for Object Tracking
Xiaoqin Zhang 0002, Weiming Hu 0004, Guan Luo, Stephen J. Maybank |
ACCV (1) | 4 |
| 2007 | Graph Based Discriminative Learning for Robust and Efficient Object TrackingabstractObject tracking is viewed as a two-class 'one-versus-rest' classification problem, in which the sample distribution of the target is approximately Gausian while the background samples are often multimodal. Based on these special properties, we propose a graph embedding based discriminative learning method, in which the topology structures of graphs are carefully designed to reflect the properties of the sample distributions. This method can simultaneously learn the subspace of the target and its local discriminative structure against the background. Moreover, a heuristic negative sample selection scheme is adopted to make the classification more effective. In tracking procedure, the graph based learning is embedded into a Bayesian inference framework cascaded with hierarchical motion estimation, which significantly improves the accuracy and efficiency of the localization. Furthermore, an incremental updating technique for the graphs is developed to capture the changes in both appearance and illumination. Experimental results demonstrate that, compared with two state-of-the-art methods, the proposed tracking algorithm is more efficient and effective, especially in dynamically changing and clutter scenes. Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank, Xi Li 0001 |
ICCV | 3 |
| 2007 | General Averaged Divergence AnalysisabstractSubspace selection is a powerful tool in data mining. An important subspace method is the Fisher-Rao linear discriminant analysis (LDA), which has been successfully applied in many fields such as biometrics, bioinformatics, and multimedia retrieval. However, LDA has a critical drawback: the projection to a subspace tends to merge those classes that are close together in the original feature space. If the separated classes are sampled from Gaussian distributions, all with identical covariance matrices, then LDA maximizes the mean value of the Kullback-Leibler (KL) divergences between the different classes. We generalize this point of view to obtain a framework for choosing a subspace by 1) generalizing the KL divergence to the Bregman divergence and 2) generalizing the arithmetic mean to a general mean. The framework is named the general averaged divergence analysis (GADA). Under this GADA framework, a geometric mean divergence analysis (GMDA) method based on the geometric mean is studied. A large number of experiments based on synthetic data show that our method significantly outperforms LDA and several representative LDA extensions. Dacheng Tao, Xuelong Li 0001, Xindong Wu 0001, Stephen J. Maybank |
ICDM | 4 |
| 2007 | Probabilistic Tensor Analysis with Akaike and Bayesian Information Criteria
Dacheng Tao, Jimeng Sun 0001, Xindong Wu 0001, Xuelong Li 0001, Jialie Shen 0001, Stephen J. Maybank, Christos Faloutsos |
ICONIP (1) | 6 |
| 2007 | Gender recognition based on local body motionsabstractHuman body motions, including gait information, are a promising biometrics resource. In this paper, the human silhouette is segmented into seven components for visual surveillance applications, namely, head, arm, body, thigh, front-leg, back-leg, and feet. The legs are classified as front-leg or back-leg because of the bipedal walking style: during walking, the left-leg and the right-leg are in front or at the back in turn. The motions of the individual components and of a number of combinations of components are then studied for gender recognition. For HumanID recognition under different cases, the performances of and underlying links amongst the seven human gait components are analyzed. Xuelong Li 0001, Stephen J. Maybank, Dacheng Tao |
SMC | 2 |
| 2007 | Application of the Fisher-Rao Metric to Ellipse Detection
Stephen J. Maybank |
Int. J. Comput. Vis. | 1 |
| 2007 | The Fisher-rao Metric for Lines in a Convex ImageabstractThe Fisher–Rao metric on the parameter space for the set of lines in a two-dimensional convex image is approximated under the assumption that the errors in the measurements are small. The volume of the parameter space under the approximating metric is proportional to the area of the image under the Euclidean metric. In the case of a rectangular image, expressions for the approximating metric are obtained and an algorithm is given for sampling the parameter space. The sample points are used in an algorithm for detecting lines in a rectangular image. Experimental results are reported. In the case of a disc shaped image the parameter space for lines embeds isometrically, under the approximating metric, into three-dimensional Euclidean space. Stephen J. Maybank |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2007 | Supervised tensor learning
Dacheng Tao, Xuelong Li 0001, Xindong Wu 0001, Weiming Hu 0004, Stephen J. Maybank |
Knowl. Inf. Syst. | 5 |
| 2007 | Recognition of Pornographic Web Pages by Classifying Texts and ImagesabstractWith the rapid development of the World Wide Web, people benefit more and more from the sharing of information. However, Web pages with obscene, harmful, or illegal content can be easily accessed. It is important to recognize such unsuitable, offensive, or pornographic Web pages. In this paper, a novel framework for recognizing pornographic Web pages is described. A C4.5 decision tree is used to divide Web pages, according to content representations, into continuous text pages, discrete text pages, and image pages. These three categories of Web pages are handled, respectively, by a continuous text classifier, a discrete text classifier, and an algorithm that fuses the results from the image classifier and the discrete text classifier. In the continuous text classifier, statistical and semantic features are used to recognize pornographic texts. In the discrete text classifier, the naive Bayes rule is used to calculate the probability that a discrete text is pornographic. In the image classifier, the object's contour-based features are extracted to recognize pornographic images. In the text and image fusion algorithm, the Bayes theory is used to combine the recognition results from images and texts. Experimental results demonstrate that the continuous text classifier outperforms the traditional keyword-statistics-based classifier, the contour-based image classifier outperforms the traditional skin-region-based image classifier, the results obtained by our fusion algorithm outperform those by either of the individual classifiers, and our framework can be adapted to different categories of Web pages. Weiming Hu 0004, Ou Wu 0001, Zhouyao Chen, Zhouyu Fu, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2007 | General Tensor Discriminant Analysis and Gabor Features for Gait RecognitionabstractThe traditional image representations are not suited to conventional classification methods, such as the linear discriminant analysis (LDA), because of the under sample problem (USP): the dimensionality of the feature space is much higher than the number of training samples. Motivated by the successes of the two dimensional LDA (2DLDA) for face recognition, we develop a general tensor discriminant analysis (GTDA) as a preprocessing step for LDA. The benefits of GTDA compared with existing preprocessing methods, e.g., principal component analysis (PCA) and 2DLDA, include 1) the USP is reduced in subsequent classification by, for example, LDA; 2) the discriminative information in the training tensors is preserved; and 3) GTDA provides stable recognition rates because the alternating projection optimization algorithm to obtain a solution of GTDA converges, while that of 2DLDA does not. We use human gait recognition to validate the proposed GTDA. The averaged gait images are utilized for gait representation. Given the popularity of Gabor function based image decompositions for image understanding and object recognition, we develop three different Gabor function based image representations: 1) the GaborD representation is the sum of Gabor filter responses over directions, 2) GaborS is the sum of Gabor filter responses over scales, and 3) GaborSD is the sum of Gabor filter responses over scales and directions. The GaborD, GaborS and GaborSD representations are applied to the problem of recognizing people from their averaged gait images.A large number of experiments were carried out to evaluate the effectiveness (recognition rate) of gait recognition based on first obtaining a Gabor, GaborD, GaborS or GaborSD image representation, then using GDTA to extract features and finally using LDA for classification. The proposed methods achieved good performance for gait recognition based on image sequences from the USF HumanID Database. Experimental comparisons are made with nine state of the art classification methods in gait recognition. Dacheng Tao, Xuelong Li 0001, Xindong Wu 0001, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2007 | Semantic-Based Surveillance Video RetrievalabstractVisual surveillance produces large amounts of video data. Effective indexing and retrieval from surveillance video databases are very important. Although there are many ways to represent the content of video clips in current video retrieval algorithms, there still exists a semantic gap between users and retrieval systems. Visual surveillance systems supply a platform for investigating semantic-based video retrieval. In this paper, a semantic-based video retrieval framework for visual surveillance is proposed. A cluster-based tracking algorithm is developed to acquire motion trajectories. The trajectories are then clustered hierarchically using the spatial and temporal information, to learn activity models. A hierarchical structure of semantic indexing and retrieval of object activities, where each individual activity automatically inherits all the semantic descriptions of the activity model to which it belongs, is proposed for accessing video clips and individual objects at the semantic level. The proposed retrieval framework supports various queries including queries by keywords, multiple object queries, and queries by sketch. For multiple object queries, succession and simultaneity restrictions, together with depth and breadth first orders, are considered. For sketch-based queries, a method for matching trajectories drawn by users to spatial trajectories is proposed. The effectiveness and efficiency of our framework are tested in a crowded traffic scene. Weiming Hu 0004, Zhouyu Fu, Wenrong Zeng, Stephen J. Maybank |
IEEE Trans. Image Process. | 5 |
| 2007 | Negative Samples Analysis in Relevance FeedbackabstractRecently, relevance feedback (RF) in content-based image retrieval (CBIR) has been implemented as an online binary classifier to separate the positive samples from the negative samples, where both sets of samples are labeled by the user. In many applications, it is reasonable to assume that all the positive samples are alike and thus that the region of the feature space occupied by the positive samples can be described by a single hypersurface. However, for the negative samples, previous RF methods either treat each one of the negative samples as an isolated point or assume the whole negative set can be described by a single convex hypersurface. In this paper, we argue that these treatments of the negative samples are not sound. Our belief is all positive samples are included in a set and the negative samples split into a small number of subsets, each one of which has a simple distribution. Therefore, we first cluster the negative samples into several groups; for each such negative group, we build a marginal convex machine (MCM) subclassifier between it and the single positive group which results in a series of subclassifiers. These subclassifiers are then incorporated into a biased MCM (BMCM) for RF. Experiments were carried out to prove the advantages of BMCM-based RF over previous methods for RF Dacheng Tao, Xuelong Li 0001, Stephen J. Maybank |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2006 | Human Carrying Status in Visual SurveillanceabstractA person’s gait changes when he or she is carrying an object such as a bag, suitcase or rucksack. As a result, human identification and tracking are made more difficult because the averaged gait image is too simple to represent the carrying status. Therefore, in this paper we first introduce a set of Gabor based human gait appearance models, because Gabor functions are similar to the receptive field profiles in the mammalian cortical simple cells. The very high dimensionality of the feature space makes training difficult. In order to solve this problem we propose a general tensor discriminant analysis (GTDA), which seamlessly incorporates the object (Gabor based human gait appearance model) structure information as a natural constraint. GTDA differs from the previous tensor based discriminant analysis methods in that the training converges. Existing methods fail to converge in the training stage. This makes them unsuitable for practical tasks. Experiments are carried out on the USF baseline data set to recognize a human’s ID from the gait silhouette. The proposed Gabor gait incorporated with GTDA is demonstrated to significantly outperform the existing appearance-based methods. Dacheng Tao, Xuelong Li 0001, Stephen J. Maybank, Xindong Wu 0001 |
CVPR (2) | 3 |
| 2006 | Elapsed Time in Human Gait Recognition: A New ApproachabstractHuman gait is an effective biometric source for human identification and visual surveillance; therefore human gait recognition becomes to be a hot topic in recent research. However, the elapsed time problem, which is in its infancy, still receives poor performance. In this paper, we introduce a novel discriminant analysis method to improve the performance. The new model inherits the merits from the tensor rank one analysis, which handles the small samples size problem naturally, and the linear discriminant analysis, which is optimal for classification. Although 2DLDA and DATR also benefit from these two methods, they cannot converge during the training procedure. This means they can be hardly utilized for practical applications. Based on a lot of experiments on elapsed time problem in human gait recognition, the new method is demonstrated to significantly outperform the existing appearance-based methods, such as the principle component analysis, the linear discriminant analysis, and the tensor rank one analysis. Dacheng Tao, Xuelong Li 0001, Xindong Wu 0001, Stephen J. Maybank |
ICASSP (2) | 4 |
| 2006 | Principal Axis-Based Correspondence between Multiple Cameras for People TrackingabstractVisual surveillance using multiple cameras has attracted increasing interest in recent years. Correspondence between multiple cameras is one of the most important and basic problems which visual surveillance using multiple cameras brings. In this paper, we propose a simple and robust method, based on principal axes of people, to match people across multiple cameras. The correspondence likelihood reflecting the similarity of pairs of principal axes of people is constructed according to the relationship between "ground-points" of people detected in each camera view and the intersections of principal axes detected in different camera views and transformed to the same view. Our method has the following desirable properties: 1) Camera calibration is not needed. 2) Accurate motion detection and segmentation are less critical due to the robustness of the principal axis-based feature to noise. 3) Based on the fused data derived from correspondence results, positions of people in each camera view can be accurately located even when the people are partially occluded in all views. The experimental results on several real video sequences from outdoor environments have demonstrated the effectiveness, efficiency, and robustness of our method. Weiming Hu 0004, Tieniu Tan, Jianguang Lou, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2006 | A System for Learning Statistical Motion PatternsabstractAnalysis of motion patterns is an effective approach for anomaly detection and behavior prediction. Current approaches for the analysis of motion patterns depend on known scenes, where objects move in predefined ways. It is highly desirable to automatically construct object motion patterns which reflect the knowledge of the scene. In this paper, we present a system for automatically learning motion patterns for anomaly detection and behavior prediction based on a proposed algorithm for robustly tracking multiple objects. In the tracking algorithm, foreground pixels are clustered using a fast accurate fuzzy K-means algorithm. Growing and prediction of the cluster centroids of foreground pixels ensure that each cluster centroid is associated with a moving object in the scene. In the algorithm for learning motion patterns, trajectories are clustered hierarchically using spatial and temporal information and then each motion pattern is represented with a chain of Gaussian distributions. Based on the learned statistical motion patterns, statistical methods are used to detect anomalies and predict behaviors. Our system is tested using image sequences acquired, respectively, from a crowded real traffic scene and a model traffic scene. Experimental results show the robustness of the tracking algorithm, the efficiency of the algorithm for learning motion patterns, and the encouraging performance of algorithms for anomaly detection and behavior prediction. Weiming Hu 0004, Xuejuan Xiao, Zhouyu Fu, Tieniu Tan, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2005 | Supervised Tensor LearningabstractThis paper aims to take general tensors as inputs for supervised learning. A supervised tensor learning (STL) framework is established for convex optimization based learning techniques such as support vector machines (SVM) and minimax probability machines (MPM). Within the STL framework, many conventional learning machines can be generalized to take n/sup th/-order tensors as inputs. We also study the applications of tensors to learning machine design and feature extraction by linear discriminant analysis (LDA). Our method for tensor based feature extraction is named the tenor rank-one discriminant analysis (TR1DA). These generalized algorithms have several advantages: 1) reduce the curse of dimension problem in machine learning and data mining; 2) avoid the failure to converge; and 3) achieve better separation between the different categories of samples. As an example, we generalize MPM to its STL version, which is named the tensor MPM (TMPM). TMPM learns a series of tensor projections iteratively. It is then evaluated against the original MPM. Our experiments on a binary classification problem show that TMPM significantly outperforms the original MPM. Dacheng Tao, Xuelong Li 0001, Weiming Hu 0004, Stephen J. Maybank, Xindong Wu 0001 |
ICDM | 4 |
| 2005 | Stable Third-Order Tensor Representation for Color Image ClassificationabstractGeneral tensors can represent colour images more naturally than conventional features; however, the general tensors' stability properties are not reported and remain to be a key problem. In this paper, we use the tensor minimax probability (TMPM) to prove that the tensor representation is stable. The proof is based on the random subspace method through a large number of experiments. Dacheng Tao, Stephen J. Maybank, Weiming Hu 0004, Xuelong Li 0001 |
Web Intelligence | 2 |
| 2005 | The Fisher-Rao Metric for Projective Transformations of the Line
Stephen J. Maybank |
Int. J. Comput. Vis. | 1 |
| 2005 | 3-D Model-Based Vehicle TrackingabstractThis paper aims at tracking vehicles from monocular intensity image sequences and presents an efficient and robust approach to three-dimensional (3-D) model-based vehicle tracking. Under the weak perspective assumption and the ground-plane constraint, the movements of model projection in the two-dimensional image plane can be decomposed into two motions: translation and rotation. They are the results of the corresponding movements of 3-D translation on the ground plane (GP) and rotation around the normal of the GP, which can be determined separately. A new metric based on point-to-line segment distance is proposed to evaluate the similarity between an image region and an instantiation of a 3-D vehicle model under a given pose. Based on this, we provide an efficient pose refinement method to refine the vehicle's pose parameters. An improved EKF is also proposed to track and to predict vehicle motion with a precise kinematics model. Experimental results with both indoor and outdoor data show that the algorithm obtains desirable performance even under severe occlusion and clutter. Jianguang Lou, Tieniu Tan, Weiming Hu 0004, Hao Yang 0010, Stephen J. Maybank |
IEEE Trans. Image Process. | 5 |
| 2004 | Detection of Image Structures Using the Fisher Information and the Rao MetricabstractIn many detection problems, the structures to be detected are parameterized by the points of a parameter space. If the conditional probability density function for the measurements is known, then detection can be achieved by sampling the parameter space at a finite number of points and checking each point to see if the corresponding structure is supported by the data. The number of samples and the distances between neighboring samples are calculated using the Rao metric on the parameter space. The Rao metric is obtained from the Fisher information which is, in turn, obtained from the conditional probability density function. An upper bound is obtained for the probability of a false detection. The calculations are simplified in the low noise case by making an asymptotic approximation to the Fisher information. An application to line detection is described. Expressions are obtained for the asymptotic approximation to the Fisher information, the volume of the parameter space, and the number of samples. The time complexity for line detection is estimated. An experimental comparison is made with a Hough transform-based method for detecting lines. Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | A survey on visual surveillance of object motion and behaviorsabstractVisual surveillance in dynamic scenes, especially for humans and vehicles, is currently one of the most active research topics in computer vision. It has a wide spectrum of promising applications, including access control in special areas, human identification at a distance, crowd flux statistics and congestion analysis, detection of anomalous behaviors, and interactive surveillance using multiple cameras, etc. In general, the processing framework of visual surveillance in dynamic scenes includes the following stages: modeling of environments, detection of motion, classification of moving objects, tracking, understanding and description of behaviors, human identification, and fusion of data from multiple cameras. We review recent developments and general strategies of all these stages. Finally, we analyze possible research directions, e.g., occlusion handling, a combination of twoand three-dimensional tracking, a combination of motion analysis and biometrics, anomaly detection and behavior prediction, content-based retrieval of surveillance videos, behavior understanding and natural language description, fusion of information from multiple sensors, and remote surveillance. Weiming Hu 0004, Tieniu Tan, Liang Wang 0001, Stephen J. Maybank |
IEEE Trans. Syst. Man Cybern. Part C | 4 |
| 2004 | Learning activity patterns using fuzzy self-organizing neural networkabstractActivity understanding in visual surveillance has attracted much attention in recent years. In this paper, we present a new method for learning patterns of object activities in image sequences for anomaly detection and activity prediction. The activity patterns are constructed using unsupervised learning of motion trajectories and object features. Based on the learned activity patterns, anomaly detection and activity prediction can be achieved. Unlike existing neural network based methods, our method uses a whole trajectory as an input to the network. This makes the network structure much simpler. Furthermore, the fuzzy set theory based method and the batch learning method are introduced into the network learning process, and make the learning process much more efficient. Two sets of data acquired, respectively, from a model scene and a campus scene are both used to test the proposed algorithms. Experimental results show that the fuzzy self-organizing neural network (fuzzy SOM) is much more efficient than the Kohonen self-organizing feature map (SOFM) and vector quantization in both speed and accuracy, and the anomaly detection and activity prediction algorithms have encouraging performances. Weiming Hu 0004, Tieniu Tan, Stephen J. Maybank |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2002 | Fusion of Multiple Tracking Algorithms for Robust People Tracking
Nils T. Siebel, Stephen J. Maybank |
ECCV (4) | 2 |
| 2000 | Visual Surveillance for Moving Vehicles
James M. Ferryman, Stephen J. Maybank, Anthony D. Worrall |
Int. J. Comput. Vis. | 2 |
| 2000 | Introduction
Stephen J. Maybank, Tieniu Tan |
Int. J. Comput. Vis. | 1 |
| 2000 | Minimum Description Length Method for Facet MatchingabstractThe Minimum Description Length (MDL) criterion is used to fit a facet model of a car to an image. The best fit is achieved when the difference image between the car and the background has the greatest compression. MDL overcomes the overfitting and parameter precision problems which hamper the more usual maximum likelihood method of model fitting. Some preliminary results are shown. Stephen J. Maybank, Roberto Fraile |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1999 | MDL, Colllineations and the Fundamental MatrixabstractScene geometry can be inferred from point correspondences between two images. The inference process includes the selection of a model. Four models are considered: background (or null), collineation, affine fundamental matrix and fundamental matrix. It is shown how Minimum Description Length (MDL) can be used to compare the different models. The main result is that there is little reason for preferring the fundamental matrix model over the collineation model, even when the former the ‘true’ model. Stephen J. Maybank, Peter F. Sturm |
BMVC | 1 |
| 1999 | A Method for Interactive 3D Reconstruction of Piecewise Planar Objects from Single ImagesabstractInternational audience Peter F. Sturm, Stephen J. Maybank |
BMVC | 2 |
| 1999 | On Plane-Based Camera Calibration: A General Algorithm, Singularities, ApplicationsabstractWe present a general algorithm for plane-based calibration that can deal with arbitrary numbers of views and calibration planes. The algorithm can simultaneously calibrate different views from a camera with variable intrinsic parameters and it is easy to incorporate known values of intrinsic parameters. For some minimal cases, we describe all singularities, naming the parameters that can not be estimated. Experimental results of our method are shown that exhibit the singularities while revealing good performance in non-singular conditions. Several applications of plane-based 3D geometry inference are discussed as well. Peter F. Sturm, Stephen J. Maybank |
CVPR | 2 |
| 1998 | Learning Enhanced 3D Models for Vehicle TrackingabstractThis paper presents an enhanced hypothesis verification strategy for 3D object recognition. A new learning methodology is presented which integrates the traditional dichotomic object-centred and appearance-based representations in computer vision giving improved hypothesis verification under iconic matching. The "appearance" of a 3D object is learnt using an eigenspace representation obtained as it is tracked through a scene. The feature representation implicitly models the background and the objects observed enabling the segmentation of the objects from the background. The method is shown to enhance model-based tracking, particularly in the presence of clutter and occlusion, and to provide a basis for identification. The unified approach is discussed in the context of the traffic surveillance domain. The approach is demonstrated on real-world image sequences and compared to previous (edge-based) iconic evaluation techniques. 1 Introduction The aim of this work is to exte... James M. Ferryman, Anthony D. Worrall, Stephen J. Maybank |
BMVC | 3 |
| 1998 | Vehicle Trajectory Approximation and ClassificationabstractWe present a variational technique for finding low curvature smooth approx-imations to trajectories in the plane. The method is applied to short segments of a vehicle trajectory in a known ground plane. Estimates of the speed and steering angle are obtained for each segment and the motion during the seg-ment is assigned to one of the four classes: ahead, left, right, stop. A hidden Markov model for the motion of the car is constructed and the Viterbi algorithm is used to find the sequence of internal states for which the observed behaviour of the vehicle has the highest probability. 1 Roberto Fraile, Stephen J. Maybank |
BMVC | 2 |
| 1998 | Ambiguity in Reconstruction from Images of Six PointsabstractLet S be a set of six points in space, let /spl psi/ be any hyperboloid of one sheet containing S, and let I be a sequence of images of S taken by an uncalibrated camera moving over /spl psi/. Then reconstruction from I is subject to a three way ambiguity which is unbroken as long as the optical centre of the camera remains on /spl psi/. Let p be an image of S taken from a point on /spl psi/. The images 'near' p define a tangent space which splits into a direct sum W/sub p//spl oplus/N/sub p//spl oplus/F/sub p/, where W/sub p/ corresponds to images near p for which the ambiguity is maintained, N/sub p/ corresponds to images for which the ambiguity is broken and F/sub p/ corresponds to images which are physically impossible. Stephen J. Maybank, Amnon Shashua |
ICCV | 1 |
| 1998 | Robust Detection of Degenerate Configurations while Estimating the Fundamental Matrix
Philip Torr 0001, Andrew Zisserman, Stephen J. Maybank |
Comput. Vis. Image Underst. | 3 |
| 1998 | Relation between 3D invariants and 2D invariants
Stephen J. Maybank |
Image Vis. Comput. | 1 |
| 1997 | Vehicle Tracking with Applications to Collision Alert
James M. Ferryman, Stephen J. Maybank, Anthony D. Worrall |
BMVC | 2 |
| 1997 | Reply to Pizlo, Rosenfeld, and Weiss
Stephen J. Maybank |
Comput. Vis. Image Underst. | 1 |
| 1996 | Filter for Car Tracking Based on Acceleration and Steering AngleabstractThe motion of a car is described using a stochastic model in which the driving processes are the steering angle and the tangential acceleration. The model incorporates exactly the kinematic constraint that the wheels do not slip sideways. Two filters based on this model have been implemented, namely the standard EKF, and a new filter (the CUF) in which the expectation and the covariance of the system state are propagated accurately. Experiments show that i) the CUF is better than the EKF at predicting future positions of the car; and ii) the filter outputs can be used to control the measurement process, leading to improved ability to recover from errors in predictive tracking. 1 Introduction In systems for monitoring road traffic it is an advantage to have filters which can model accurately the motion of a car and predict its future positions. It becomes easier to track individual cars and to analyse their behaviour. Many current filters use over-simplified models based on general mot... Stephen J. Maybank, Anthony D. Worrall, Geoffrey D. Sullivan |
BMVC | 1 |
| 1996 | A Filter for Visual Tracking Based on a Stochastic Model for Driver Behaviour
Stephen J. Maybank, Anthony D. Worrall, Geoffrey D. Sullivan |
ECCV (2) | 1 |
| 1996 | Stochastic properties of the cross ratio
Stephen J. Maybank |
Pattern Recognit. Lett. | 1 |
| 1995 | Robust Detection of Degenerate Configurations for the Fundamental MatrixabstractNew methods are reported for the detection of multiple solutions (degeneracy) when estimating the fundamental matrix, with specific emphasis on robustness in the presence of data contamination (outliers). The fundamental matrix can be used as a first step in the recovery of structure from motion. If the set of correspondences is degenerate then this structure cannot be accurately recovered and many solutions will explain the data equally well. It is essential that we are alerted to such eventualities. However, current feature matchers are very prone to mismatching, giving a high rate of contamination within the data. Such contamination can make a degenerate data set appear non degenerate, thus the need for robust methods becomes apparent. The paper presents such methods with a particular emphasis on providing a method that will work on real imagery and with an automated (non perfect) feature detector and matcher. It is demonstrated that proper modelling of degeneracy in the presence of outliers enables the detection of outliers which would otherwise be missed. Results using real image sequences are presented. All processing, point matching, degeneracy detection and outlier detection is automatic.> Philip Torr 0001, Andrew Zisserman, Stephen J. Maybank |
ICCV | 3 |
| 1995 | Probabilistic analysis of the application of the cross ratio to model based vision: Misclassification
Stephen J. Maybank |
Int. J. Comput. Vis. | 1 |
| 1995 | Probabilistic analysis of the application of the cross ratio to model based vision
Stephen J. Maybank |
Int. J. Comput. Vis. | 1 |
| 1992 | Camera Self-Calibration: Theory and Experiments
Olivier D. Faugeras, Quang-Tuan Luong, Stephen J. Maybank |
ECCV | 3 |
| 1992 | A theory of self-calibration of a moving camera
Stephen J. Maybank, Olivier D. Faugeras |
Int. J. Comput. Vis. | 1 |
| 1991 | Ambiguity in reconstruction from image correspondences
Stephen J. Maybank |
Image Vis. Comput. | 1 |
| 1990 | Filter based estimates of depth
Stephen J. Maybank |
BMVC | 1 |
| 1990 | Ambiguity In Reconstruction From Image Correspondences
Stephen J. Maybank |
ECCV | 1 |
| 1990 | Motion from point matches: Multiplicity of solutions
Olivier D. Faugeras, Stephen J. Maybank |
Int. J. Comput. Vis. | 2 |
| 1990 | Rigid velocities compatible with five image velocity vectors
Stephen J. Maybank |
Image Vis. Comput. | 1 |
| 1987 | Apparent area of a rigid moving body
Stephen J. Maybank |
Image Vis. Comput. | 1 |
| 1987 | The Nearest Neighbor and the Bayes Error RatesabstractThe (k, l) nearest neighbor method of pattern classification is compared to the Bayes method. If the two acceptance rates are equal then the asymptotic error rates satisfy the inequalities Ek,l + 1 ¿ E*(¿) ¿ Ek,l dE*(¿), where d is a function of k, l, and the number of pattern classes, and ¿ is the reject threshold for the Bayes method. An explicit expression for d is given which is optimal in the sense that for some probability distributions Ek,l and dE* (¿) are equal. George Loizou, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1986 | Algorithm for analysing optical flow based on the least-squares method
Stephen J. Maybank |
Image Vis. Comput. | 1 |
| 1986 | Optical flow and the Taylor expansion
Stephen J. Maybank |
Pattern Recognit. Lett. | 1 |
| 1983 | A note on nearest neighbour error rates
Stephen J. Maybank |
Pattern Recognit. Lett. | 1 |