VLDB 2026 Research / reviewers in the wild / expert
Shaogang Gong
dblp:17/4389
· DBLP profile ↗
307ranked-venue papers
12as first author
54since 2021 · last 2026
0000-0001-8156-2299ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 272 · 11 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 218 · 8 first-author · 35 since 2021Databases, data management, data science and information retrieval · 5Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Large-Scale Pre-Trained Models Empowering Phrase Generalization in Temporal Sentence Localization
Yang Liu 0105, Minghang Zheng, Qingchao Chen, Shaogang Gong, Yuxin Peng 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | Leverage cross-domain variations for generalizable person ReID representation learning
Qilei Li, Shitong Sun, Weitong Cai, Shaogang Gong |
Pattern Recognit. | 4 |
| 2026 | Boosting Multimodal Chain of Thought Reasoning by Selective Mixture of Experts
Qilei Li, Shitong Sun, Da Li 0001, Timothy M. Hospedales, Shaogang Gong |
Pattern Recognit. | 5 |
| 2026 | GridCLIP: One-stage object detection by grid-level CLIP representation learningabstract• We exploit CLIP to supplement the missing knowledge of undersampled and unseen categories in training a one-stage detector, mitigating the poor performance due to the long-tail data distribution in most existing detection training data. • We propose a simple yet effective visual-to-visual knowledge distillation method for learning undersampled and unseen categories for constructing a one-stage CLIP-based detector, providing 2.4 AP gains on unseen categories compared to the baseline. • GridCLIP is capable of handling Open-Vocabulary Object Detection with considerable scalability and generalizability, reaching comparable performance to two-stage detectors with much higher training and inference speed, without using extra pretraining processes or additional fine-tuning datasets. CLIP provides a shared image-text representation space with rich and diverse vocabulary, enabling object detection in undersampled and unseen categories. Recent CLIP-based object detection works show two-stage detectors typically outperform one-stage designs, but with significantly higher computational costs. A fundamental limitation of a two-stage detector is region-level alignment (distillation), which requires hundreds of image encoder forward passes from both the detector and CLIP in each image. In this work, we propose GridCLIP, a one-stage detector that requires only a single image encoder inference per input image, achieving up to 43 × faster training and 5 × faster inference compared to its two-stage counterpart ViLD, while substantially narrowing the accuracy gap. GridCLIP introduces a dual alignment strategy to learn fine-grained, grid-level representations: (1) grid-level alignment: learning grid-level features aligned with CLIP text encoder using annotated category labels, and (2) image-level alignment: aggregating grid-level features into an image-level representation aligned with the CLIP image encoder, which allows GridCLIP to learn grid-level representations of a broad range of categories, especially undersampled and unseen categories. Experiments on the LVIS benchmark show that GridCLIP achieves competitive results, with strong generalization to COCO and VOC, demonstrating its efficiency and effectiveness as a CLIP-based detector. Jiayi Lin 0002, Shitong Sun, Shaogang Gong |
Pattern Recognit. | 3 |
| 2026 | BenchCIR: Benchmarking robustness in composed image retrieval across modalitiesabstractComposed image retrieval aims to retrieve images based on a query that consists of a reference image and text describing desired modifications to that image. It has recently attracted attention for its ability to tailor image retrieval to user intentions by combining information-rich reference images with concise natural language instructions. Despite its current success, the robustness of composed image retrieval methods to either (1) common corruptions or (2) variations of the textual descriptions have never been systematically evaluated. In this paper, we perform the first robustness study of composed image retrieval, establishing three new benchmarks for a systematic evaluation of robustness to common corruption (in both the textual and visual domains) and robustness in text understanding. For analysis of natural image corruption, we introduce two new large-scale benchmark datasets, CIRR-C and FashionIQ-C, for the open domains and fashion domains respectively–both of which feature 75 visual corruptions and 35 textual corruptions. To facilitate robust evaluation of text understanding, we introduce a new diagnostic dataset CIRR-D by expanding the CIRR dataset with synthetic data, specifically probing text understanding across variations in: numerical, attribute, object removal, and background. We introduce BenchCIR, a testbed for evaluating composed image retrieval model robustness with standardized evaluation protocols. Through benchmarking ten published models in the testbed, we reveal insights into how the composition of visual and textual modalities affects model robustness. The code is in https://suntongtongtong.github.io/BenchCIR/ Shitong Sun, Qilei Li, Shaogang Gong, Weitong Cai, Philip Torr 0001, Jindong Gu |
Pattern Recognit. | 3 |
| 2026 | Class-Aware Diversified Augmentation for Open-Set Single Domain GeneralizationabstractIn Open-Set Single Domain Generalization (OS-SDG), one only has access to a single labeled source domain for training. It assumes that the learned model generalizes well to target samples belonging to the source label space whilst classifies target samples outside the source label space into a single “unknown” class. The current method synthesizes new samples that are semantically unrelated to known classes to simulate target unknown classes. This ignores that unknown classes actually may semantically correlated to known classes, making it difficult to discriminate samples at the margins of class decision boundaries as “unknown”. In this work, we introduce a Class-Aware Diversified Augmentation (CADA) method to overcome this problem. Our key idea is to synthesize explicitly new multiple unknown target classes with diversified semantic and learn the inherent correlation among the known and unknown classes, so to both increase the coverage of multiple target unknown classes and to optimize class margin separation. CADA is optimized by enhanced diversity maximization and class-aware minimization. The former synthesizes more novel classes by considering both semantic relationships to known classes and domain shift between the source and target domains. The latter employs class-agnostic clustering with synthesized samples to simulate class correlations among target classes, maximizing class margin separation. Theoretical analysis and experiments on five benchmarks show the efficacy of our CADA. Jian Hu 0002, Shaogang Gong, Weitong Cai, Junchi Yan |
IEEE Trans. Multim. | 2 |
| 2025 | InvSeg: Test-Time Prompt Inversion for Semantic SegmentationabstractVisual-textual correlations in the attention maps derived from text-to-image diffusion models are proven beneficial to dense visual prediction tasks, e.g., semantic segmentation. However, a significant challenge arises due to the input distributional discrepancy between the context-rich sentences used for image generation and the isolated class names typically used in semantic segmentation. This discrepancy hinders diffusion models from capturing accurate visual-textual correlations. To solve this, we propose InvSeg, a test-time prompt inversion method that tackles open-vocabulary semantic segmentation by inverting image-specific visual context into text prompt embedding space, leveraging structure information derived from the diffusion model's reconstruction process to enrich text prompts so as to associate each class with a structure-consistent mask. Specifically, we introduce Contrastive Soft Clustering (CSC) to align derived masks with the image's structure information, softly selecting anchors for each class and calculating weighted distances to push inner-class pixels closer while separating inter-class pixels, thereby ensuring mask distinction and internal consistency. By incorporating sample-specific context, InvSeg learns context-rich text prompts in embedding space and achieves accurate semantic alignment across modalities. Experiments show that InvSeg achieves state-of-the-art performance on the PASCAL VOC, PASCAL Context and COCO Object datasets. Jiayi Lin 0002, Jiabo Huang, Jian Hu 0002, Shaogang Gong |
AAAI | 4 |
| 2025 | Generative Video Diffusion for Unseen Novel Semantic Video Moment RetrievalabstractVideo moment retrieval (VMR) aims to locate the most likely video moment(s) corresponding to a text query in untrimmed videos. Training of existing methods is limited by the lack of diverse and generalisable VMR datasets, hindering their ability to generalise moment-text associations to queries containing novel semantic concepts (unseen both visually and textually in a training source domain). For model generalisation to novel semantics, existing methods rely heavily on assuming to have access to both video and text sentence pairs from a target domain in addition to the source domain pair-wise training data. This is neither practical nor scalable. In this work, we introduce a more generalisable approach by assuming only text sentences describing new semantics are available in model training without having seen any videos from a target domain. To that end, we propose a Fine-grained Video Editing framework, termed FVE, that explores generative video diffusion to facilitate fine-grained video editing from the seen source concepts to the unseen target sentences consisting of new concepts. This enables generative hypotheses of unseen video moments corresponding to the novel concepts in the target domain. This fine-grained generative video diffusion retains the original video structure and subject specifics from the source domain while introducing semantic distinctions of unseen novel vocabularies in the target domain. A critical challenge is how to enable this generative fine-grained diffusion process to be meaningful in optimising VMR, more than just synthesising visually pleasing videos. We solve this problem by introducing a hybrid selection mechanism that integrates three quantitative metrics to selectively incorporate synthetic video moments (novel video hypotheses) as enlarged additions to the original source training data, whilst minimising potential detrimental noise or unnecessary repetitions in the novel synthetic videos harmful to VMR learning. Experiments on three datasets demonstrate the effectiveness of FVE to unseen novel semantic video moment retrieval tasks Dezhao Luo, Shaogang Gong, Jiabo Huang, Hailin Jin, Yang Liu 0105 |
AAAI | 2 |
| 2025 | Temporal Score Analysis for Understanding and Correcting Diffusion ArtifactsabstractVisual artifacts remain a persistent challenge in diffusion models, even with training on massive datasets. Current solutions primarily rely on supervised detectors, yet lack understanding of why these artifacts occur in the first place. In our analysis, we identify three distinct phases in the diffusion generative process: Profiling, Mutation, and Refinement. Artifacts typically emerge during the Mutation phase, where certain regions exhibit anomalous score dynamics over time, causing abrupt disruptions in the normal evolution pattern. This temporal nature explains why existing methods focusing only on spatial uncertainty of the final output fail at effective artifact localization. Based on these insights, we propose ASCED (Abnormal Score Correction for Enhancing Diffusion), that detects artifacts by monitoring abnormal score dynamics during the diffusion process, with a trajectory-aware on-the-fly mitigation strategy that appropriate generation of noise in the detected areas. Unlike most existing methods that apply post hoc corrections, e.g., by applying a noising-denoising scheme after generation, our mitigation strategy operates seamlessly within the existing diffusion process. Extensive experiments demonstrate that our proposed approach effectively reduces artifacts across diverse domains, matching or surpassing existing supervised methods without additional training. Project page: YuCao16.github.io/ASCED. Zengqun Zhao, Ioannis Patras, Shaogang Gong |
CVPR | 4 |
| 2025 | AIM-Fair: Advancing Algorithmic Fairness via Selectively Fine-Tuning Biased Models with Contextual Synthetic DataabstractRecent advances in generative models have sparked research on improving model fairness with AI-generated data. However, existing methods often face limitations in the diversity and quality of synthetic data, leading to compromised fairness and overall model accuracy. Moreover, many approaches rely on the availability of demographic group labels, which are often costly to annotate. This paper proposes AIM-Fair, aiming to overcome these limitations and harness the potential of cutting-edge generative models in promoting algorithmic fairness. We investigate a fine-tuning paradigm starting from a biased model initially trained on real-world data without demographic annotations. This model is then fine-tuned using unbiased synthetic data generated by a state-of-the-art diffusion model to improve its fairness. Two key challenges are identified in this fine-tuning paradigm, 1) the low quality of synthetic data, which can still happen even with advanced generative models, and 2) the domain and bias gap between real and synthetic data. To address the limitation of synthetic data quality, we propose Contextual Synthetic Data Generation (CSDG) to generate data using a text-to-image diffusion model (T2I) with prompts generated by a context-aware LLM, ensuring both data diversity and control of bias in synthetic data. To resolve domain and bias shifts, we introduce a novel selective fine-tuning scheme in which only model parameters more sensitive to bias and less sensitive to domain shift are updated. Experiments on CelebA and UTKFace datasets show that our AIM-Fair improves model fairness while maintaining utility, outperforming both fully and partially fine-tuned approaches to model fairness. The code is available at https://github.com/zengqunzhao/AIM-Fair. Zengqun Zhao, Ziquan Liu, Shaogang Gong, Ioannis Patras |
CVPR | 4 |
| 2025 | Multi-Modal Multi-Platform Person Re-Identification: Benchmark and Method
Ruiyang Ha, Songyi Jiang, Bikang Pan, Yihang Zhu, Junjie Zhang 0002, Xiatian Zhu, Shaogang Gong, Jingya Wang 0001 |
ICCV | 8 |
| 2025 | INT: Instance-Specific Negative Mining for Task-Generic Promptable SegmentationabstractTask-generic promptable image segmentation aims to achieve segmentation of diverse samples under a single task description by utilizing only one task-generic prompt. Current methods leverage the generalization capabilities of Vision-Language Models (VLMs) to infer instance-specific prompts from these task-generic prompts in order to guide the segmentation process. However, when VLMs struggle to generalise to some image instances, predicting instance-specific prompts becomes poor. To solve this problem, we introduce Instance-specific Negative Mining for Task-Generic Promptable Segmentation (INT). The key idea of INT is to adaptively reduce the influence of irrelevant (negative) prior knowledge whilst to increase the use the most plausible prior knowledge, selected by negative mining with higher contrast, in order to optimise instance-specific prompts generation. Specifically, INT consists of two components: (1) instance-specific prompt generation, which progressively fliters out incorrect information in prompt generation; (2) semantic mask generation, which ensures each image instance segmentation matches correctly the semantics of the instance-specific prompts. INT is validated on six datasets, including camouflaged objects and medical images, demonstrating its effectiveness, robustness and scalability. Zixu Cheng, Shaogang Gong |
IJCAI | 3 |
| 2025 | XFMamba: Cross-Fusion Mamba for Multi-view Medical Image Classification
Xiaoyu Zheng 0001, Xu Chen 0030, Shaogang Gong, Xavier Griffin, Gregory Slabaugh |
MICCAI (1) | 3 |
| 2025 | Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Video Temporal GroundingabstractVideo Temporal Grounding (TG) aims to temporally locate video segments matching a natural language description (a query) in a long video. While Vision-Language Models (VLMs) are effective at holistic semantic matching, they often struggle with fine-grained temporal
localisation. Recently, Group Relative Policy Optimisation (GRPO) reformulates the inference process as a reinforcement learning task, enabling fine-grained grounding and achieving strong in-domain performance. However, GRPO relies on labelled data, making it unsuitable in unlabelled domains. Moreover, because videos are large and expensive to store and process, performing full-scale adaptation introduces prohibitive latency and computational overhead, making it impractical for real-time deployment. To overcome both problems, we introduce a Data-Efficient Unlabelled Cross-domain Temporal Grounding method, from which a model is first trained on a labelled source domain, then adapted to a target domain using only a small number of {\em unlabelled videos from the target domain}. This approach eliminates the need for target annotation and keeps both computational and storage overhead low enough to run in real time. Specifically, we introduce \textbf{U}ncertainty-quantified \textbf{R}ollout \textbf{P}olicy \textbf{A}daptation (\textbf{URPA}) for cross-domain knowledge transfer in learning video temporal grounding without target labels. URPA generates multiple candidate predictions using GRPO rollouts, averages them to form a pseudo label, and estimates confidence from the variance across these rollouts. This confidence then weights the training rewards, guiding the model to focus on reliable supervision. Experiments on three datasets across six cross-domain settings show that URPA generalises well using only a few unlabelled target videos. Codes are given in supplemental materials. Zixu Cheng, Shaogang Gong, Isabel Guan, Jianye Hao, Jun Wang 0012, Kun Shao |
NeurIPS | 3 |
| 2025 | Enhancing Zero-Shot Facial Expression Recognition by LLM Knowledge TransferabstractCurrent facial expression recognition (FER) models are often designed in a supervised learning manner and thus are constrained by the lack of large-scale facial expression images with high-quality annotations. Consequently, these models often fail to generalize well, performing poorly on unseen images in inference. Vision-language-based zero-shot models demonstrate a promising potential for addressing such challenges. However, these models lack task-specific knowledge and therefore are not optimized for the nuances of recognizing facial expressions. To bridge this gap, this work proposes a novel method, Exp-CLIP, to enhance zero-shot FER by transferring the task knowledge from large language models (LLMs). Specifically, based on the pre-trained vision-language encoders, we incorporate a projection head designed to map the initial joint vision-language space into a space that captures representations of facial actions. To train this projection head for subsequent zero-shot predictions, we propose to align the projected visual representations with task-specific semantic meanings derived from the LLM encoder, and the text instruction-based strategy is employed to customize the LLM knowledge. Given unlabelled facial data and efficient training of the projection head, Exp-CLIP achieves superior zero-shot results to the CLIP models and several other large vision-language models (LVLMs) on seven in-the-wild FER datasets. The code is available at https://github.com/zengqunzhao/Exp-CLIP. Zengqun Zhao, Shaogang Gong, Ioannis Patras |
WACV | 3 |
| 2025 | MLLM as video narrator: Mitigating modality imbalance in video moment retrieval
Weitong Cai, Jiabo Huang, Shaogang Gong, Hailin Jin, Yang Liu 0105 |
Pattern Recognit. | 3 |
| 2024 | Relax Image-Specific Prompt Requirement in SAM: A Single Generic Prompt for Segmenting Camouflaged ObjectsabstractCamouflaged object detection (COD) approaches heavily rely on pixel-level annotated datasets. Weakly-supervised COD (WSCOD) approaches use sparse annotations like scribbles or points to reduce annotation efforts, but this can lead to decreased accuracy. The Segment Anything Model (SAM) shows remarkable segmentation ability with sparse prompts like points. However, manual prompt is not always feasible, as it may not be accessible in real-world application. Additionally, it only provides localization information instead of semantic one, which can intrinsically cause ambiguity in interpreting targets. In this work, we aim to eliminate the need for manual prompt. The key idea is to employ Cross-modal Chains of Thought Prompting (CCTP) to reason visual prompts using the semantic information given by a generic text prompt. To that end, we introduce a test-time instance-wise adaptation mechanism called Generalizable SAM (GenSAM) to automatically generate and optimize visual prompts from the generic task prompt for WSCOD. In particular, CCTP maps a single generic text prompt onto image-specific consensus foreground and background heatmaps using vision-language models, acquiring reliable visual prompts. Moreover, to test-time adapt the visual prompts, we further propose Progressive Mask Generation (PMG) to iteratively reweight the input image, guiding the model to focus on the targeted region in a coarse-to-fine manner. Crucially, all network parameters are fixed, avoiding the need for additional training. Experiments on three benchmarks demonstrate that GenSAM outperforms point supervision approaches and achieves comparable results to scribble supervision ones, solely relying on general task descriptions. Our codes is in https://github.com/jyLin8100/GenSAM. Jian Hu 0002, Jiayi Lin 0002, Shaogang Gong, Weitong Cai |
AAAI | 3 |
| 2024 | Few-Shot Image Generation by Conditional Relaxing Diffusion Inversion
Shaogang Gong |
ECCV (84) | 2 |
| 2024 | SHINE: Saliency-Aware Hierarchical Negative Ranking for Compositional Temporal Grounding
Zixu Cheng, Yujiang Pu, Shaogang Gong, Parisa Kordjamshidi, Yu Kong 0001 |
ECCV (19) | 3 |
| 2024 | Feature-Distribution Perturbation and Calibration for Generalized ReidabstractPerson Re-identification (ReID) has been advanced remarkably over the last 10 years. However, the i.i.d. (independent and identically distributed) assumption is somewhat non-applicable to ReID considering its objective to identify images of the same pedestrian across cameras at different locations. In this work, we propose a Feature-Distribution Perturbation and Calibration (PECA) method to derive generic feature representations for person ReID. Specifically, we perform per-domain feature-distribution perturbation to refrain the model from overfitting to the domain-biased distribution of each source (seen) domain by enforcing feature invariance to distribution shifts caused by perturbation. Furthermore, we design a global calibration mechanism to align feature distributions across all the source domains to improve the model’s generalization capacity by eliminating domain bias. These local perturbation and global calibration are conducted simultaneously, which share the same principle to avoid models overfitting by regularization respectively on the perturbed and the original distributions. Extensive experiments were conducted and the proposed PECA model outperformed the state-of-the-art competitors by significant margins. Qilei Li, Jiabo Huang, Jian Hu 0002, Shaogang Gong |
ICASSP | 4 |
| 2024 | Leveraging Hallucinations to Reduce Manual Prompt Dependency in Promptable SegmentationabstractPromptable segmentation typically requires instance-specific manual prompts to guide the segmentation of each desired object. To minimize such a need, task-generic promptable segmentation has been introduced, which employs a single task-generic prompt to segment various images of different objects in the same task. Current methods use Multimodal Large Language Models (MLLMs) to reason detailed instance-specific prompts from a task-generic prompt for improving segmentation accuracy. The effectiveness of this segmentation heavily depends on the precision of these derived prompts. However, MLLMs often suffer hallucinations during reasoning, resulting in inaccurate prompting. While existing methods focus on eliminating hallucinations to improve a model, we argue that MLLM hallucinations can reveal valuable contextual insights when leveraged correctly, as they represent pre-trained large-scale knowledge beyond individual images. In this paper, we first utilize hallucinations to mine task-related information from images and verify its accuracy to enhance precision of the generated prompts. Specifically, we introduce an iterative \textbf{Pro}mpt-\textbf{Ma}sk \textbf{C}ycle generation framework (ProMaC) with a prompt generator and a mask generator. The prompt generator uses a multi-scale chain of thought prompting, initially leveraging hallucinations to extract extended contextual prompts on a test image. These hallucinations are then minimized to formulate precise instance-specific prompts, directing the mask generator to produce masks that are consistent with task semantics by mask semantic alignment. Iteratively the generated masks induce the prompt generator to focus more on task-relevant image areas and reduce irrelevant hallucinations, resulting jointly in better prompts and masks. Experiments on 5 benchmarks demonstrate the effectiveness of ProMaC. Code is in https://lwpyh.github.io/ProMaC/. Jian Hu 0002, Jiayi Lin 0002, Junchi Yan, Shaogang Gong |
NeurIPS | 4 |
| 2024 | Mitigate Domain Shift by Primary-Auxiliary Objectives Association for Generalizing Person ReIDabstractWhile deep learning has significantly improved ReID model accuracy under the independent and identical distribution (IID) assumption, it has also become clear that such models degrade notably when applied to an unseen novel domain due to unpredictable/unknown domain shift. Contemporary domain generalization (DG) ReID models struggle in learning domain-invariant representation solely through training on an instance classification objective. We consider that a deep learning model is heavily influenced and therefore biased towards domain-specific characteristics, e.g., background clutter, scale and viewpoint variations, limiting the generalizability of the learned model, and hypothesize that the pedestrians are domain invariant owning they share the same structural characteristics. To enable the ReID model to be less domain-specific from these pure pedestrians, we introduce a method that guides model learning of the primary ReID instance classification objective by a concurrent auxiliary learning objective on weakly la-beled pedestrian saliency detection. To solve the problem of conflicting optimization criteria in the model parameter space between the two learning objectives, we introduce a Primary-Auxiliary Objectives Association (PAOA) mechanism to calibrate the loss gradients of the auxiliary task towards the primary learning task gradients. Benefiting from the harmonious multitask learning design, our model can be extended with the recent test-time diagram to form the PAOA+, which performs on-the-fly optimization against the auxiliary objective in order to maximize the model’s generative capacity in the test target domain. Experiments demonstrate the superiority of the proposed PAOA model. Qilei Li, Shaogang Gong |
WACV | 2 |
| 2024 | Zero-Shot Video Moment Retrieval from Frozen Vision-Language ModelsabstractAccurate video moment retrieval (VMR) requires universal visual-textual correlations that can handle unknown vocabulary and unseen scenes. However, the learned correlations are likely either biased when derived from a limited amount of moment-text data which is hard to scale up because of the prohibitive annotation cost (fully-supervised), or unreliable when only the video-text pairwise relationships are available without fine-grained temporal annotations (weakly-supervised). Recently, the vision-language models (VLM) demonstrate a new transfer learning paradigm to benefit different vision tasks through the universal visual-textual correlations derived from large-scale vision-language pairwise web data, which has also shown benefits to VMR by fine-tuning in the target domains.In this work, we propose a zero-shot method for adapting generalisable visual-textual priors from arbitrary VLM to facilitate moment-text alignment, without the need for accessing the VMR data. To this end, we devise a conditional feature refinement module to generate boundary-aware visual features conditioned on text queries to enable better moment boundary understanding. Additionally, we design a bottom-up proposal generation strategy that mitigates the impact of domain discrepancies and breaks down complex-query retrieval tasks into individual action retrievals, thereby maximizing the benefits of VLM. Extensive experiments conducted on three VMR benchmark datasets demonstrate the notable performance advantages of our zero-shot algorithm, especially in the novel-word and novel-location out-of-distribution setups. Dezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin, Yang Liu 0105 |
WACV | 3 |
| 2024 | Faster Person Re-Identification: One-Shot-Filter and Coarse-to-Fine SearchabstractFast person re-identification (ReID) aims to search person images quickly and accurately. The main idea of recent fast ReID methods is the hashing algorithm, which learns compact binary codes and performs fast Hamming distance and counting sort. However, a very long code is needed for high accuracy (e.g.2048), which compromises search speed. In this work, we introduce a new solution for fast ReID by formulating a novel Coarse-to-Fine (CtF) hashing code search strategy, which complementarily uses short and long codes, achieving both faster speed and better accuracy. It uses shorter codes to coarsely rank broad matching similarities and longer codes to refine only a few top candidates for more accurate instance ReID. Specifically, we design an All-in-One (AiO) module together with a Distance Threshold Optimization (DTO) algorithm. In AiO, we simultaneously learn and enhance multiple codes of different lengths in a single model. It learns multiple codes in a pyramid structure, and encourage shorter codes to mimic longer codes by self-distillation. DTO solves a complex threshold search problem by a simple optimization process, and the balance between accuracy and speed is easily controlled by a single parameter. It formulates the optimization target as a$F_{\beta }$score that can be optimised by Gaussian cumulative distribution functions. Besides, we find even short code (e.g.32) still takes a long time under large-scale gallery due to the$O(n)$time complexity. To solve the problem, we propose a gallery-size-free latent-attributes-based One-Shot-Filter (OSF) strategy, that is always$O(1)$time complexity, to quickly filter major easy negative gallery images, Specifically, we design a Latent-Attribute-Learning (LAL) module supervised a Single-Direction-Metric (SDM) Loss. LAL is derived from principal component analysis (PCA) that keeps largest variance using shortest feature vector, meanwhile enabling batch and end-to-end learning. Every logit of a feature vector represents a meaningful attribute. SDM is carefully designed for fine-grained attribute supervision, outperforming common metrics such as Euclidean and Cosine metrics. Experimental results on 2 datasets show that CtF+OSF is not only$2\%$more accurate but also$5\times$faster than contemporary hashing ReID methods. Compared with non-hashing ReID methods, CtF is$50\times$faster with comparable accuracy. OSF further speeds CtF by$2\times$again and upto$10\times$in total with almost no accuracy drop. Guan'an Wang, Xiaowen Huang 0001, Shaogang Gong, Jian Zhang 0018, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Federated zero-shot learning with mid-level semantic knowledge transferabstractConventional centralized deep learning paradigms are not feasible when data from different sources cannot be shared due to data privacy or transmission limitation. To resolve this problem, federated learning has been introduced to transfer knowledge across multiple sources (clients) with non-shared data while optimizing a globally generalized central model (server). Existing federated learning paradigms mostly focus on transmitting image encoders that take instance-sensitive images as input, making them less generalizable and vulnerable to privacy inference attacks. In contrast, in this work, we consider transferring mid-level semantic knowledge (such as attribute) which is not sensitive to specific objects of interest and therefore is more privacy-preserving and general. To this end, we formulate a new Federated Zero-Shot Learning (FZSL) paradigm to learn mid-level semantic knowledge at multiple local clients with non-shared local data and cumulatively aggregate a globally generalized central model for deployment. To improve model discriminative ability, we explore semantic knowledge available from either a language or a vision-language foundation model in order to enrich the mid-level semantic space in FZSL. Extensive experiments on five zero-shot learning benchmark datasets validate the effectiveness of our approach for optimizing a generalizable federated learning model with mid-level semantic knowledge transfer. Shitong Sun, Chenyang Si, Guile Wu, Shaogang Gong |
Pattern Recognit. | 4 |
| 2023 | Generating Structured Pseudo Labels for Noise-resistant Zero-shot Video Sentence LocalizationabstractVideo sentence localization aims to locate moments in an unstructured video according to a given natural language query.A main challenge is the expensive annotation costs and the annotation bias.In this work, we study video sentence localization in a zero-shot setting, which learns with only video data without any annotation.Existing zero-shot pipelines usually generate event proposals and then generate a pseudo query for each event proposal.However, their event proposals are obtained via visual feature clustering, which is query-independent and inaccurate; and the pseudo-queries are short or less interpretable.Moreover, existing approaches ignores the risk of pseudo-label noise when leveraging them in training.To address the above problems, we propose a Structurebased Pseudo Label generation (SPL), which first generate free-form interpretable pseudo queries before constructing query-dependent event proposals by modeling the event temporal structure.To mitigate the effect of pseudolabel noise, we propose a noise-resistant iterative method that repeatedly re-weight the training sample based on noise estimation to train a grounding model and correct pseudo labels.Experiments on the ActivityNet Captions and Charades-STA datasets demonstrate the advantages of our approach.Code can be found at https://github.com/minghangz/SPL. Minghang Zheng, Shaogang Gong, Hailin Jin, Yuxin Peng 0001, Yang Liu 0105 |
ACL (1) | 2 |
| 2023 | Towards Generalisable Video Moment Retrieval: Visual-Dynamic Injection to Image-Text Pre-TrainingabstractThe correlation between the vision and text is essential for video moment retrieval (VMR), however, existing methods heavily rely on separate pre-training feature extractors for visual and textual understanding. Without sufficient temporal boundary annotations, it is non-trivial to learn universal video-text alignments. In this work, we explore multi-modal correlations derived from large-scale image-text data to facilitate generalisable VMR. To address the limitations of image-text pre-training models on capturing the video changes, we propose a generic method, referred to as Visual-Dynamic Injection (VDI), to empower the model's understanding of video moments. Whilst existing VMR methods are focusing on building temporalaware video features, being aware of the text descriptions about the temporal changes is also critical but originally overlooked in pre-training by matching static images with sentences. Therefore, we extract visual context and spatial dynamic information from video frames and explicitly enforce their alignments with the phrases describing video changes (e.g. verb). By doing so, the potentially relevant visual and motion patterns in videos are encoded in the corresponding text embeddings (injected) so to enable more accurate video-text alignments. We conduct extensive experiments on two VMR benchmark datasets (Charades-STA and ActivityNet-Captions) and achieve state-of-the-art performances. Especially, VDI yields notable advantages when being tested on the out-of-distribution splits where the testing samples involve novel scenes and vocabulary. Dezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin, Yang Liu 0105 |
CVPR | 3 |
| 2023 | Rapid Person Re-Identification via Sub-space Consistency Regularization
Qingze Yin, Guan'an Wang, Guodong Ding, Qilei Li, Shaogang Gong, Zhenmin Tang |
Neural Process. Lett. | 5 |
| 2023 | Neural operator searchabstractExisting neural architecture search (NAS) methods usually explore a limited feature-transformation-only search space, ignoring other advanced feature operations such as feature self-calibration by attention and dynamic convolutions. This disables the NAS algorithms to discover more advanced network architectures. We address this limitation by additionally exploiting feature self-calibration operations, resulting in a heterogeneous search space. To solve the challenges of operation heterogeneity and significantly larger search space, we formulate a neural operator search (NOS) method. NOS presents a novel heterogeneous residual block for integrating the heterogeneous operations in a unified structure, and an attention guided search strategy for facilitating the search process over a vast space. Extensive experiments show that NOS can search novel cell architectures with highly competitive performance on the CIFAR and ImageNet benchmarks. Wei Li 0022, Shaogang Gong, Xiatian Zhu |
Pattern Recognit. | 2 |
| 2022 | 3D Shape Temporal Aggregation for Video-Based Clothing-Change Person Re-identification
Yan Huang 0008, Shaogang Gong, Liang Wang 0001, Tieniu Tan |
ACCV (5) | 3 |
| 2022 | Hybrid-Learning Video Moment Retrieval across Multi-Domain Labels
Weitong Cai, Jiabo Huang, Shaogang Gong |
BMVC | 3 |
| 2022 | Deep Clustering by Semantic Contrastive Learning
Jiabo Huang, Shaogang Gong |
BMVC | 2 |
| 2022 | Ranking Distance Calibration for Cross-Domain Few-Shot LearningabstractRecent progress in few-shot learning promotes a more realistic cross-domain setting, where the source and target datasets are in different domains. Due to the domain gap and disjoint label spaces between source and target datasets, their shared knowledge is extremely limited. This encourages us to explore more information in the target domain rather than to overly elaborate training strategies on the source domain as in many existing methods. Hence, we start from a generic representation pre-trained by a cross-entropy loss and a conventional distance-based classifier, along with an image retrieval view, to employ a re-ranking process to calibrate a target distance matrix by discovering the k-reciprocal neighbours within the task. Assuming the pre-trained representation is biased towards the source, we construct a non-linear subspace to minimise task-irrelevant features therewithin while keep more transferrable discriminative information by a hyperbolic tangent transformation. The calibrated distance in this target-aware non-linear sub-space is complementary to that in the pre-trained representation. To impose such distance calibration information onto the pre-trained representation, a Kullback-Leibler divergence loss is employed to gradually guide the model towards the calibrated distance-based distribution. Extensive evaluations on eight target domains show that this target ranking calibration process can improve conventional distance-based classifiers in few-shot learning. Shaogang Gong, Chengjie Wang 0001, Yanwei Fu 0001 |
CVPR | 2 |
| 2022 | Learning Unbiased Transferability for Domain Adaptation by Uncertainty Modeling
Jian Hu 0002, Haowen Zhong, Shaogang Gong, Guile Wu, Junchi Yan |
ECCV (31) | 4 |
| 2022 | Video Activity Localisation with Uncertainties in Temporal Boundary
Jiabo Huang, Hailin Jin, Shaogang Gong, Yang Liu 0105 |
ECCV (34) | 3 |
| 2022 | Joint Bilateral-Resolution Identity Modeling for Cross-Resolution Person Re-Identification
Wei-Shi Zheng 0001, Jincheng Hong, Jiening Jiao, Ancong Wu, Xiatian Zhu, Shaogang Gong, Jiayin Qin, Jian-Huang Lai |
Int. J. Comput. Vis. | 6 |
| 2022 | Unsupervised cross-domain person re-identification by instance and distribution alignmentabstractMost existing person re-identification (re-id) methods assume supervised model training on a separate large set of training samples from the target domain. While performing well in the training domain, such trained models are seldom generalisable to a new independent unsupervised target domain without further labelled training data from the target domain. To solve this scalability limitation, we develop a novel Hierarchical Unsupervised Domain Adaptation (HUDA) method. It can transfer labelled information of an existing dataset (a source domain) to an unlabelled target domain for unsupervised person re-id. Specifically, HUDA is designed to model jointly global distribution alignment and local instance alignment in a two-level hierarchy for discovering transferable source knowledge in unsupervised domain adaptation. Crucially, this approach aims to overcome the under-constrained learning problem of existing unsupervised domain adaptation methods. Extensive evaluations show the superiority of HUDA for unsupervised cross-domain person re-id over a wide variety of state-of-the-art methods on four re-id benchmarks: Market-1501, DukeMTMC, MSMT17 and CUHK03. Xu Lan, Xiatian Zhu, Shaogang Gong |
Pattern Recognit. | 3 |
| 2022 | Learning hybrid ranking representation for person re-identification
Guile Wu, Xiatian Zhu, Shaogang Gong |
Pattern Recognit. | 3 |
| 2021 | Generalising without Forgetting for Lifelong Person Re-IdentificationabstractExisting person re-identification (Re-ID) methods mostly prepare all training data in advance, while real-world Re-ID data are inherently captured over time or from different locations, which requires a model to be incrementally generalised from sequential learning of piecemeal new data without forgetting what is already learned. In this work, we call this lifelong person Re-ID, characterised by solving a problem of unseen class identification subject to continuous new domain generalisation and adaptation with class imbalanced learning. We formulate a new Generalising without Forgetting method (GwFReID) for lifelong Re-ID and design a comprehensive learning objective that accounts for classification coherence, distribution coherence and representation coherence in a unified framework. This design helps to simultaneously learn new information, distil old knowledge and solve class imbalance, which enables GwFReID to incrementally improve model generalisation without catastrophic forgetting of what is already learned. Extensive experiments on eight Re-ID benchmarks, CIFAR-100 and ImageNet show the superiority of GwFReID over the state-of-the-art methods. Guile Wu, Shaogang Gong |
AAAI | 2 |
| 2021 | Decentralised Learning from Independent Multi-Domain Labels for Person Re-IdentificationabstractDeep learning has been successful for many computer vision tasks due to the availability of shared and centralised large-scale training data. However, increasing awareness of privacy concerns poses new challenges to deep learning, especially for human subject related recognition such as person re-identification (Re-ID). In this work, we solve the Re-ID problem by decentralised learning from non-shared private training data distributed at multiple user sites of independent multi-domain label spaces. We propose a novel paradigm called Federated Person Re-Identification (FedReID) to construct a generalisable global model (a central server) by simultaneously learning with multiple privacy-preserved local models (local clients). Specifically, each local client receives global model updates from the server and trains a local model using its local data independent from all the other clients. Then, the central server aggregates transferrable local model updates to construct a generalisable global feature embedding model without accessing local data so to preserve local privacy. This client-server collaborative learning process is iteratively performed under privacy control, enabling FedReID to realise decentralised learning without sharing distributed data nor collecting any centralised data. Extensive experiments on ten Re-ID benchmarks show that FedReID achieves compelling generalisation performance beyond any locally trained models without using shared training data, whilst inherently protects the privacy of each local client. This is uniquely advantageous over contemporary Re-ID methods. Guile Wu, Shaogang Gong |
AAAI | 2 |
| 2021 | Peer Collaborative Learning for Online Knowledge DistillationabstractTraditional knowledge distillation uses a two-stage training strategy to transfer knowledge from a high-capacity teacher model to a compact student model, which relies heavily on the pre-trained teacher. Recent online knowledge distillation alleviates this limitation by collaborative learning, mutual learning and online ensembling, following a one-stage end-to-end training fashion. However, collaborative learning and mutual learning fail to construct an online high-capacity teacher, whilst online ensembling ignores the collaboration among branches and its logit summation impedes the further optimisation of the ensemble teacher. In this work, we propose a novel Peer Collaborative Learning method for online knowledge distillation, which integrates online ensembling and network collaboration into a unified framework. Specifically, given a target network, we construct a multi-branch network for training, in which each branch is called a peer. We perform random augmentation multiple times on the inputs to peers and assemble feature representations outputted from peers with an additional classifier as the peer ensemble teacher. This helps to transfer knowledge from a high-capacity teacher to peers, and in turn further optimises the ensemble teacher. Meanwhile, we employ the temporal mean model of each peer as the peer mean teacher to collaboratively transfer knowledge among peers, which helps each peer to learn richer knowledge and facilitates to optimise a more stable model with better generalisation. Extensive experiments on CIFAR-10, CIFAR-100 and ImageNet show that the proposed method significantly improves the generalisation of various backbone networks and outperforms the state-of-the-art methods. Guile Wu, Shaogang Gong |
AAAI | 2 |
| 2021 | Local-Global Associative Frame Assemble in Video Re-ID
Qilei Li, Jiabo Huang, Shaogang Gong |
BMVC | 3 |
| 2021 | Decentralised Person Re-Identification with Selective Knowledge Aggregation
Shitong Sun, Guile Wu, Shaogang Gong |
BMVC | 3 |
| 2021 | Cross-Sentence Temporal and Semantic Relations in Video Activity LocalisationabstractVideo activity localisation has recently attained increasing attention due to its practical values in automatically localising the most salient visual segments corresponding to their language descriptions (sentences) from untrimmed and unstructured videos. For supervised model training, a temporal annotation of both the start and end time index of each video segment for a sentence (a video moment) must be given. This is not only very expensive but also sensitive to ambiguity and subjective annotation bias, a much harder task than image labelling. In this work, we develop a more accurate weakly-supervised solution by introducing Cross-Sentence Relations Mining (CRM) in video moment proposal generation and matching when only a paragraph description of activities without per-sentence temporal annotation is available. Specifically, we explore two cross-sentence relational constraints: (1) Temporal ordering and (2) semantic consistency among sentences in a paragraph description of video activities. Existing weakly-supervised techniques only consider within-sentence video segment correlations in training without considering cross-sentence paragraph context. This can mislead due to ambiguous expressions of individual sentences with visually indiscriminate video moment proposals in isolation. Experiments on two publicly available activity localisation datasets show the advantages of our approach over the state-of-the-art weakly supervised methods, especially so when the video activity descriptions become more complex. Jiabo Huang, Yang Liu 0105, Shaogang Gong, Hailin Jin |
ICCV | 3 |
| 2021 | A Simple Feature Augmentation for Domain GeneralizationabstractThe topical domain generalization (DG) problem asks trained models to perform well on an unseen target domain with different data statistics from the source training domains. In computer vision, data augmentation has proven one of the most effective ways of better exploiting the source data to improve domain generalization. However, existing approaches primarily rely on image-space data augmentation, which requires careful augmentation design, and provides limited diversity of augmented data. We argue that feature augmentation is a more promising direction for DG. We find that an extremely simple technique of perturbing the feature embedding with Gaussian noise during training leads to a classifier with domain-generalization performance comparable to existing state of the art. To model more meaningful statistics reflective of cross-domain variability, we further estimate the full class-conditional feature covariance matrix iteratively during training. Subsequent joint stochastic feature augmentation provides an effective domain randomization method, perturbing features in the directions of intra-class/cross-domain variability. We verify our proposed method on three standard domain generalization benchmarks, Digit-DG, VLCS and PACS, and show it is outperforming or comparable to the state of the art in all setups, together with experimental analysis to illustrate how our method works towards training a robust generalisable model. Da Li 0001, Wei Li 0132, Shaogang Gong, Yanwei Fu 0001, Timothy M. Hospedales |
ICCV | 4 |
| 2021 | Collaborative Optimization and Aggregation for Decentralized Domain Generalization and AdaptationabstractContemporary domain generalization (DG) and multisource unsupervised domain adaptation (UDA) methods mostly collect data from multiple domains together for joint optimization. However, this centralized training paradigm poses a threat to data privacy and is not applicable when data are non-shared across domains. In this work, we propose a new approach called Collaborative Optimization and Aggregation (COPA), which aims at optimizing a generalized target model for decentralized DG and UDA, where data from different domains are non-shared and private. Our base model consists of a domain-invariant feature extractor and an ensemble of domain-specific classifiers. In an iterative learning process, we optimize a local model for each domain, and then centrally aggregate local feature extractors and assemble domain-specific classifiers to construct a generalized global model, without sharing data from different domains. To improve generalization of feature extractors, we employ hybrid batch-instance normalization and collaboration of frozen classifiers. For better decentralized UDA, we further introduce a prediction agreement mechanism to overcome local disparities towards central model aggregation. Extensive experiments on five DG and UDA benchmark datasets show that COPA is capable of achieving comparable performance against the state-of-the-art DG and UDA methods without the need for centralized data collection in model training. Guile Wu, Shaogang Gong |
ICCV | 2 |
| 2021 | Striking a Balance between Stability and Plasticity for Class-Incremental LearningabstractClass-incremental learning (CIL) aims at continuously updating a trained model with new classes (plasticity) without forgetting previously learned old ones (stability). Contemporary studies resort to storing representative exemplars for rehearsal or preventing consolidated model parameters from drifting, but the former requires an additional space for storing exemplars at every incremental phase while the latter usually shows poor model generalization. In this paper, we focus on resolving the stability-plasticity dilemma in class-incremental learning where no exemplars from old classes are stored. To make a trade-off between learning new information and maintaining old knowledge, we reformulate a simple yet effective baseline method based on a cosine classifier framework and reciprocal adaptive weights. With the reformulated baseline, we present two new approaches to CIL by learning class-independent knowledge and multi-perspective knowledge, respectively. The former exploits class-independent knowledge to bridge learning new and old classes, while the latter learns knowledge from different perspectives to facilitate CIL. Extensive experiments on several widely used CIL benchmark datasets show the superiority of our approaches over the state-of-the-art methods. Guile Wu, Shaogang Gong |
ICCV | 2 |
| 2021 | Semi-Supervised Few-Shot Learning with Pseudo Label RefinementabstractFew-shot classification aims at recognising novel categories with very limited labelled samples. Although substantial achievements have been obtained, few-shot classification remains challenging due to the scarcity of labelled examples. Recent studies resort to leveraging unlabelled data to expand the training set using pseudo labelling, but this strategy often yields significant label noise. In this work, we introduce a new baseline method for semi-supervised few-shot learning by iterative pseudo label refinement to reduce noise. Then, we investigate the label noise propagation problem and improve the baseline with a denoising network to learn distributions of clean and noisy pseudo-labelled examples via a mixture model. This helps to estimate confidence values of pseudo labelled examples and to select the reliable ones with less noise for iteratively refining a few-shot classifier. Extensive experiments on three widely used benchmarks, minilma- genet, tieredImagenet and CIFAR-FS, show the superiority of the proposed methods over the state-of-the-art methods. Guile Wu, Shaogang Gong, Xu Lan |
ICME | 3 |
| 2021 | Regularising Knowledge Transfer by Meta Functional LearningabstractMachine learning classifiers’ capability is largely dependent on the scale of available training data and limited by the model overfitting in data-scarce learning tasks. To address this problem, this work proposes a novel Meta Functional Learning (MFL) by meta-learning a generalisable functional model from data-rich tasks whilst simultaneously regularising knowledge transfer to data-scarce tasks. The MFL computes meta-knowledge on functional regularisation generalisable to different learning tasks by which functional training on limited labelled data promotes more discriminative functions to be learned. Moreover, we adopt an Iterative Update strategy on MFL (MFL-IU). This improves knowledge transfer regularisation from MFL by progressively learning the functional regularisation in knowledge transfer. Experiments on three Few-Shot Learning (FSL) benchmarks (miniImageNet, CIFAR-FS and CUB) show that meta functional learning for regularisation knowledge transfer can benefit improving FSL classifiers. Yanwei Fu 0001, Shaogang Gong |
IJCAI | 3 |
| 2021 | Multi-perspective cross-class domain adaptation for open logo detectionabstractExisting logo detection methods mostly rely on supervised learning with a large quantity of labelled training data in limited classes. This restricts their scalability to a large number of logo classes subject to limited labelling budget. In this work, we consider a more scalable open logo detection problem where only a fraction of logo classes are fully labelled whilst the remaining classes are only annotated with a clean icon image (e.g. 1-shot icon supervised). To generalise and transfer knowledge of fully supervised logo classes to other 1-shot icon supervised classes, we propose a Multi-Perspective Cross-Class (MPCC) domain adaptation method. In a data augmentation principle, MPCC conducts feature distribution alignment in two perspectives. Specifically, we align the feature distribution between synthetic logo images of 1-shot icon supervised classes and genuine logo images of fully supervised classes, and that between logo images and non-logo images, concurrently. This allows for mitigating the domain shift problem between model training and testing on 1-shot icon supervised logo classes, simultaneously reducing the model overfitting towards fully labelled logo classes. Extensive comparative experiments show the advantage of MPCC over existing state-of-the-art competitors on the challenging QMUL-OpenLogo benchmark (Su et al., 2018). Hang Su 0004, Shaogang Gong, Xiatian Zhu |
Comput. Vis. Image Underst. | 2 |
| 2021 | Intra-Camera Supervised Person Re-IdentificationabstractAbstract Existing person re-identification (re-id) methods mostly exploit a large set of cross-camera identity labelled training data. This requires a tedious data collection and annotation process, leading to poor scalability in practical re-id applications. On the other hand unsupervised re-id methods do not need identity label information, but they usually suffer from much inferior and insufficient model performance. To overcome these fundamental limitations, we propose a novel person re-identification paradigm based on an idea ofindependentper-camera identity annotation. This eliminates the most time-consuming and tedious inter-camera identity labelling process, significantly reducing the amount of human annotation efforts. Consequently, it gives rise to a more scalable and more feasible setting, which we callIntra-Camera Supervised (ICS)person re-id, for which we formulate a Multi-tAsk mulTi-labEl (MATE) deep learning method. Specifically, MATE is designed for self-discovering the cross-camera identity correspondence in a per-camera multi-task inference framework. Extensive experiments demonstrate the cost-effectiveness superiority of our method over the alternative approaches on three large person re-id datasets. For example, MATE yields 88.7% rank-1 score on Market-1501 in the proposed ICS person re-id setting, significantly outperforming unsupervised learning models and closely approaching conventional fully supervised learning competitors. Xiangping Zhu, Xiatian Zhu, Minxian Li, Pietro Morerio, Vittorio Murino, Shaogang Gong |
Int. J. Comput. Vis. | 6 |
| 2021 | Hierarchical distillation learning for scalable person search
Wei Li 0132, Shaogang Gong, Xiatian Zhu |
Pattern Recognit. | 2 |
| 2021 | Multi-View Label Prediction for Unsupervised Learning Person Re-IdentificationabstractPerson re-identification (ReID) aims to match pedestrian images across disjoint cameras. Existing supervised ReID methods utilize deep networks and train them with identity-labeled images, which suffer from limited annotations. Recently, clustering-based unsupervised ReID attracts more and more attention. It first clusters unlabeled images and assigns cluster index to the pseudo-identity-labels, then trains a ReID model with the pseudo-identity-labels. However, considering the slight inter-class variations and significant intra-class variations, pseudo-identity-labels learned from clustering algorithms are usually noisy and coarse. To alleviate the problems above, besides clustering pseudo-identity-labels, we propose to learn pseudo-patch-labels, which brings two advantages: (1) Patch naturally alleviates the effect of backgrounds, occlusions, and carryings since they usually occupy small parts in images, thus overcome noisy labels. (2) It is plausible that patches from different pedestrians belong to the same pseudo-identity-label. For example, pedestrians have a high probability of wearing either the same shoes or pants but a low possibility of wearing both. The experiments demonstrate our proposed method achieves the best performance by a large margin on both image- and video-based datasets. Qingze Yin, Guan'an Wang, Guodong Ding, Shaogang Gong, Zhenmin Tang |
IEEE Signal Process. Lett. | 4 |
| 2021 | Guest Editorial Introduction to the Special Issue on Large-Scale Visual Sensor Networks: Architectures and ApplicationsabstractLarge–scale visual sensor networks have become progressively an essential part of our daily lives underpinning many technological, financial, and social advancements today, with applications in smart cities, traffic monitoring, environmental pollution control, public safety, and crime prevention. Paolo Spagnolo, Hamid K. Aghajan, George Bebis, Shaogang Gong, Amy Loutfi, Leonid Sigal, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Semi-Supervised Learning under Class Distribution MismatchabstractSemi-supervised learning (SSL) aims to avoid the need for collecting prohibitively expensive labelled training data. Whilst demonstrating impressive performance boost, existing SSL methods artificially assume that small labelled data and large unlabelled data are drawn from the same class distribution. In a more realistic scenario with class distribution mismatch between the two sets, they often suffer severe performance degradation due to error propagation introduced by irrelevant unlabelled samples. Our work addresses this under-studied and realistic SSL problem by a novel algorithm named Uncertainty-Aware Self-Distillation (UASD). Specifically, UASD produces soft targets that avoid catastrophic error propagation, and empower learning effectively from unconstrained unlabelled data with out-of-distribution (OOD) samples. This is based on joint Self-Distillation and OOD filtering in a unified formulation. Without bells and whistles, UASD significantly outperforms six state-of-the-art methods in more realistic SSL under class distribution mismatch on three popular image classification datasets: CIFAR10, CIFAR100, and TinyImageNet. Yanbei Chen, Xiatian Zhu, Wei Li 0132, Shaogang Gong |
AAAI | 4 |
| 2020 | Unsupervised Deep Learning via Affinity DiffusionabstractConvolutional neural networks (CNNs) have achieved unprecedented success in a variety of computer vision tasks. However, they usually rely on supervised model learning with the need for massive labelled training data, limiting dramatically their usability and deployability in real-world scenarios without any labelling budget. In this work, we introduce a general-purpose unsupervised deep learning approach to deriving discriminative feature representations. It is based on self-discovering semantically consistent groups of unlabelled training samples with the same class concepts through a progressive affinity diffusion process. Extensive experiments on object image classification and clustering show the performance superiority of the proposed method over the state-of-the-art unsupervised learning models using six common image recognition benchmarks including MNIST, SVHN, STL10, CIFAR10, CIFAR100 and ImageNet. Jiabo Huang, Qi Dong 0004, Shaogang Gong, Xiatian Zhu |
AAAI | 3 |
| 2020 | Neural Graph Embedding for Neural Architecture SearchabstractExisting neural architecture search (NAS) methods often operate in discrete or continuous spaces directly, which ignores the graphical topology knowledge of neural networks. This leads to suboptimal search performance and efficiency, given the factor that neural networks are essentially directed acyclic graphs (DAG). In this work, we address this limitation by introducing a novel idea of neural graph embedding (NGE). Specifically, we represent the building block (i.e. the cell) of neural networks with a neural DAG, and learn it by leveraging a Graph Convolutional Network to propagate and model the intrinsic topology information of network architectures. This results in a generic neural network representation integrable with different existing NAS frameworks. Extensive experiments show the superiority of NGE over the state-of-the-art methods on image classification and semantic segmentation. Wei Li 0132, Shaogang Gong, Xiatian Zhu |
AAAI | 2 |
| 2020 | Tracklet Self-Supervised Learning for Unsupervised Person Re-IdentificationabstractExisting unsupervised person re-identification (re-id) methods mainly focus on cross-domain adaptation or one-shot learning. Although they are more scalable than the supervised learning counterparts, relying on a relevant labelled source domain or one labelled tracklet per person initialisation still restricts their scalability in real-world deployments. To alleviate these problems, some recent studies develop unsupervised tracklet association and bottom-up image clustering methods, but they still rely on explicit camera annotation or merely utilise suboptimal global clustering. In this work, we formulate a novel tracklet self-supervised learning (TSSL) method, which is capable of capitalising directly from abundant unlabelled tracklet data, to optimise a feature embedding space for both video and image unsupervised re-id. This is achieved by designing a comprehensive unsupervised learning objective that accounts for tracklet frame coherence, tracklet neighbourhood compactness, and tracklet cluster structure in a unified formulation. As a pure unsupervised learning re-id model, TSSL is end-to-end trainable at the absence of source data annotation, person identity labels, and camera prior knowledge. Extensive experiments demonstrate the superiority of TSSL over a wide variety of the state-of-the-art alternative methods on four large-scale person re-id benchmarks, including Market-1501, DukeMTMC-ReID, MARS and DukeMTMC-VideoReID. Guile Wu, Xiatian Zhu, Shaogang Gong |
AAAI | 3 |
| 2020 | Image Search With Text Feedback by Visiolinguistic Attention LearningabstractImage search with text feedback has promising impacts in various real-world applications, such as e-commerce and internet search. Given a reference image and text feedback from user, the goal is to retrieve images that not only resemble the input image, but also change certain aspects in accordance with the given text. This is a challenging task as it requires the synergistic understanding of both image and text. In this work, we tackle this task by a novel Visiolin-guistic Attention Learning (VAL) framework. Specifically, we propose a composite transformer that can be seamlessly plugged in a CNN to selectively preserve and transform the visual features conditioned on language semantics. By inserting multiple composite transformers at varying depths, VAL is incentive to encapsulate the multi-granular visiolinguistic information, thus yielding an expressive representation for effective image search. We conduct comprehensive evaluation on three datasets: Fashion200k, Shoes and FashionIQ. Extensive experiments show our model exceeds existing approaches on all datasets, demonstrating consistent superiority in coping with various text feedbacks, including attribute-like and natural language descriptions. Yanbei Chen, Shaogang Gong, Loris Bazzani |
CVPR | 2 |
| 2020 | Inter-Task Association Critic for Cross-Resolution Person Re-IdentificationabstractPerson images captured by unconstrained surveillance cameras often have low resolutions (LR). This causes the resolution mismatch problem when matched against the high-resolution (HR) gallery images, negatively affecting the performance of person re-identification (re-id). An effective approach is to leverage image super-resolution (SR) along with person re-id in a joint learning manner. However, this scheme is limited due to dramatically more difficult gradients backpropagation during training. In this paper, we introduce a novel model training regularisation method, called Inter-Task Association Critic (INTACT), to address this fundamental problem. Specifically, INTACT discovers the underlying association knowledge between image SR and person re-id, and leverages it as an extra learning constraint for enhancing the compatibility of SR model with person re-id in HR image space. This is realised by parameterising the association constraint which enables it to be automatically learned from the training data. Extensive experiments validate the superiority of INTACT over the state-of-the-art approaches on the cross-resolution re-id task using five standard person re-id datasets. Zhiyi Cheng, Qi Dong 0004, Shaogang Gong, Xiatian Zhu |
CVPR | 3 |
| 2020 | Deep Semantic Clustering by Partition Confidence MaximisationabstractBy simultaneously learning visual features and data grouping, deep clustering has shown impressive ability to deal with unsupervised learning for structure analysis of high-dimensional visual data. Existing deep clustering methods typically rely on local learning constraints based on inter-sample relations and/or self-estimated pseudo labels. This is susceptible to the inevitable errors distributed in the neighbourhoods and suffers from error-propagation during training. In this work, we propose to solve this problem by learning the most confident clustering solution from all the possible separations, based on the observation that assigning samples from the same semantic categories into different clusters will reduce both the intra-cluster compactness and inter-cluster diversity, i.e. lower partition confidence. Specifically, we introduce a novel deep clustering method named PartItion Confidence mAximisation (PICA). It is established on the idea of learning the most semantically plausible data separation, in which all clusters can be mapped to the ground-truth classes one-to-one, by maximising the "global" partition confidence of clustering solution. This is realised by introducing a differentiable partition uncertainty index and its stochastic approximation as well as a principled objective loss function that minimises such index, all of which together enables a direct adoption of the conventional deep networks and mini-batch based model training. Extensive experiments on six widely-adopted clustering benchmarks demonstrate our model's performance superiority over a wide range of the state-of-the-art approaches. The code is available online. Jiabo Huang, Shaogang Gong, Xiatian Zhu |
CVPR | 2 |
| 2020 | Faster Person Re-identification
Guan'an Wang, Shaogang Gong, Jian Cheng 0001, Zeng-Guang Hou |
ECCV (8) | 2 |
| 2020 | Characteristic Regularisation for Super-Resolving Face ImagesabstractExisting facial image super-resolution (SR) methods focus mostly on improving "artificially down-sampled" lowresolution (LR) imagery. Such SR models, although strong at handling artificial LR images, often suffer from significant performance drop on genuine LR test data. Previous unsupervised domain adaptation (UDA) methods address this issue by training a model using unpaired genuine LR and HR data as well as cycle consistency loss formulation. However, this renders the model overstretched with two tasks: consistifying the visual characteristics and enhancing the image resolution. Importantly, this makes the end-to-end model training ineffective due to the difficulty of back-propagating gradients through two concatenated CNNs. To solve this problem, we formulate a method that joins the advantages of conventional SR and UDA models. Specifically, we separate and control the optimisations for characteristics consistifying and image super-resolving by introducing Characteristic Regularisation (CR) between them. This task split makes the model training more effective and computationally tractable. Extensive evaluations demonstrate the performance superiority of our method over state-of-the-art SR and UDA models on both genuine and artificial LR facial imagery data. Zhiyi Cheng, Xiatian Zhu, Shaogang Gong |
WACV | 3 |
| 2020 | Scalable Person Re-Identification by Harmonious AttentionabstractAbstract Existing person re-identification (re-id) deep learning methods rely heavily on the utilisation of large and computationally expensive convolutional neural networks. They are thereforenot scalableto large scale re-id deployment scenarios with the need of processing a large amount of surveillance video data, due to the lengthy inference process with high computing costs. In this work, we address this limitation via jointly learning re-id attention selection. Specifically, we formulate a novelharmonious attention network(HAN) framework to jointly learn soft pixel attention and hard region attention alongside simultaneous deep feature representation learning, particularly enabling more discriminative re-id matching byefficientnetworks with more scalable model inference and feature matching. Extensive evaluations validate the cost-effectiveness superiority of the proposed HAN approach for person re-id against a wide variety of state-of-the-art methods on four large benchmark datasets: CUHK03, Market-1501, DukeMTMC, and MSMT17. Wei Li 0132, Xiatian Zhu, Shaogang Gong |
Int. J. Comput. Vis. | 3 |
| 2020 | RGB-IR Person Re-identification by Cross-Modality Similarity Preservation
Ancong Wu, Wei-Shi Zheng 0001, Shaogang Gong, Jian-Huang Lai |
Int. J. Comput. Vis. | 3 |
| 2020 | Unsupervised Tracklet Person Re-IdentificationabstractMost existing person re-identification (re-id) methods rely on supervised model learning on per-camera-pair manually labelled pairwise training data. This leads to poor scalability in a practical re-id deployment, due to the lack of exhaustive identity labelling of positive and negative image pairs for every camera-pair. In this work, we present an unsupervised re-id deep learning approach. It is capable of incrementally discovering and exploiting the underlying re-id discriminative information from automatically generated person tracklet data end-to-end. We formulate an Unsupervised Tracklet Association Learning (UTAL) framework. This is by jointly learning within-camera tracklet discrimination and cross-camera tracklet association in order to maximise the discovery of tracklet identity matching both within and across camera views. Extensive experiments demonstrate the superiority of the proposed model over the state-of-the-art unsupervised learning and domain adaptation person re-id methods on eight benchmarking datasets. Minxian Li, Xiatian Zhu, Shaogang Gong |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Face re-identification challenge: Are face recognition models good enough?
Zhiyi Cheng, Xiatian Zhu, Shaogang Gong |
Pattern Recognit. | 3 |
| 2020 | Scalable logo detection by self co-learning
Hang Su 0004, Shaogang Gong, Xiatian Zhu |
Pattern Recognit. | 2 |
| 2019 | Single-Label Multi-Class Image Classification by Deep Logistic RegressionabstractThe objective learning formulation is essential for the success of convolutional neural networks. In this work, we analyse thoroughly the standard learning objective functions for multiclass classification CNNs: softmax regression (SR) for singlelabel scenario and logistic regression (LR) for multi-label scenario. Our analyses lead to an inspiration of exploiting LR for single-label classification learning, and then the disclosing of the negative class distraction problem in LR. To address this problem, we develop two novel LR based objective functions that not only generalise the conventional LR but importantly turn out to be competitive alternatives to SR in single label classification. Extensive comparative evaluations demonstrate the model learning advantages of the proposed LR functions over the commonly adopted SR in single-label coarse-grained object categorisation and cross-class fine-grained person instance identification tasks. We also show the performance superiority of our method on clothing attribute classification in comparison to the vanilla LR function. The code had been made publicly available. Qi Dong 0004, Xiatian Zhu, Shaogang Gong |
AAAI | 3 |
| 2019 | Spatio-Temporal Associative Representation for Video Person Re-Identification
Guile Wu, Xiatian Zhu, Shaogang Gong |
BMVC | 3 |
| 2019 | Unsupervised Person Re-Identification by Soft Multilabel LearningabstractAlthough unsupervised person re-identification (RE-ID) has drawn increasing research attentions due to its potential to address the scalability problem of supervised RE-ID models, it is very challenging to learn discriminative information in the absence of pairwise labels across disjoint camera views. To overcome this problem, we propose a deep model for the soft multilabel learning for unsupervised RE-ID. The idea is to learn a soft multilabel (real-valued label likelihood vector) for each unlabeled person by comparing the unlabeled person with a set of known reference persons from an auxiliary domain. We propose the soft multilabel-guided hard negative mining to learn a discriminative embedding for the unlabeled target domain by exploring the similarity consistency of the visual features and the soft multilabels of unlabeled target pairs. Since most target pairs are cross-view pairs, we develop the cross-view consistent soft multilabel learning to achieve the learning goal that the soft multilabels are consistently good across different camera views. To enable effecient soft multilabel learning, we introduce the reference agent learning to represent each reference person by a reference agent in a joint embedding. We evaluate our unified deep model on Market-1501 and DukeMTMC-reID. Our model outperforms the state-of-the-art unsupervised RE-ID methods by clear margins. Code is available at https://github.com/KovenYu/MAR. Hong-Xing Yu, Wei-Shi Zheng 0001, Ancong Wu, Shaogang Gong, Jian-Huang Lai |
CVPR | 5 |
| 2019 | Person Search by Text Attribute Query As Zero-Shot LearningabstractExisting person search methods predominantly assume the availability of at least one-shot imagery sample of the queried person. This assumption is limited in circumstances where only a brief textual (or verbal) description of the target person is available. In this work, we present a deep learning method for attribute text description based person search without any query imagery. Whilst conventional cross-modality matching methods, such as global visual-textual embedding based zero-shot learning and local individual attribute recognition, are functionally applicable, they are limited by several assumptions invalid to person search in deployment scale, data quality, and/or category name semantics. We overcome these issues by formulating an Attribute-Image Hierarchical Matching (AIHM) model. It is able to more reliably match text attribute descriptions with noisy surveillance person images by jointly learning global category-level and local attribute-level textual-visual embedding as well as matching. Extensive evaluations demonstrate the superiority of our AIHM model over a wide variety of state-of-the-art methods on three publicly available attribute labelled surveillance person search benchmarks: Market-1501, DukeMTMC, and PA100K. Qi Dong 0004, Xiatian Zhu, Shaogang Gong |
ICCV | 3 |
| 2019 | Instance-Guided Context Rendering for Cross-Domain Person Re-IdentificationabstractExisting person re-identification (re-id) methods mostly assume the availability of large-scale identity labels for model learning in any target domain deployment. This greatly limits their scalability in practice. To tackle this limitation, we propose a novel Instance-Guided Context Rendering scheme, which transfers the source person identities into diverse target domain contexts to enable supervised re-id model learning in the unlabelled target domain. Unlike previous image synthesis methods that transform the source person images into limited fixed target styles, our approach produces more visually plausible, and diverse synthetic training data. Specifically, we formulate a dual conditional generative adversarial network that augments each source person image with rich contextual variations. To explicitly achieve diverse rendering effects, we leverage abundant unlabelled target instances as contextual guidance for image generation. Extensive experiments on Market-1501, DukeMTMC-reID and CUHK03 benchmarks show that the re-id performance can be significantly improved when using our synthetic data in cross-domain re-id model learning. Yanbei Chen, Xiatian Zhu, Shaogang Gong |
ICCV | 3 |
| 2019 | Deep Reinforcement Active Learning for Human-in-the-Loop Person Re-IdentificationabstractMost existing person re-identification(Re-ID) approaches achieve superior results based on the assumption that a large amount of pre-labelled data is usually available and can be put into training phrase all at once. However, this assumption is not applicable to most real-world deployment of the Re-ID task. In this work, we propose an alternative reinforcement learning based human-in-the-loop model which releases the restriction of pre-labelling and keeps model upgrading with progressively collected data. The goal is to minimize human annotation efforts while maximizing Re-ID performance. It works in an iteratively updating framework by refining the RL policy and CNN parameters alternately. In particular, we formulate a Deep Reinforcement Active Learning (DRAL) method to guide an agent (a model in a reinforcement learning process) in selecting training samples on-the-fly by a human user/annotator. The reinforcement learning reward is the uncertainty value of each human selected sample. A binary feedback (positive or negative) labelled by the human annotator is used to select the samples of which are used to fine-tune a pre-trained CNN Re-ID model. Extensive experiments demonstrate the superiority of our DRAL method for deep reinforcement learning based human-in-the-loop person Re-ID when compared to existing unsupervised and transfer learning models as well as active learning models. Zimo Liu, Jingya Wang 0001, Shaogang Gong, Dacheng Tao, Huchuan Lu |
ICCV | 3 |
| 2019 | Person Re-Identification by Ranking Ensemble RepresentationsabstractExisting deep learning algorithms for person re-identification (re-id) typically rely on single-sample classification or pairwise matching constraints. This indicates a breach of deployment due to ignoring the probe-specific matching information against the gallery set encoded in ranking lists. In this work, we address this problem by exploring the idea of RANkinG Ensembles (RANGE) that learns such information from the ranking lists. Specifically, given an off-the-self deep re-id feature representation model, we construct per-probe ranking lists and exploit them to learn inter ranking ensemble representation. To mitigate the harm of inevitable false gallery positives, we further introduce a complementary intra ranking ensemble representation. Extensive experiments show that both supervised and unsupervised re-id benefit from the proposed RANGE method on four challenging benchmarks: MSMT17, Market-1501, DukeMTMC-ReID, and CUHK03. Guile Wu, Xiatian Zhu, Shaogang Gong |
ICIP | 3 |
| 2019 | Unsupervised Deep Learning by Neighbourhood DiscoveryabstractDeep convolutional neural networks (CNNs) have demonstrated remarkable success in computer vision by supervisedly learning strong visual feature representations. However, training CNNs relies heavily on the availability of exhaustive training data annotations, limiting significantly their deployment and scalability in many application scenarios. In this work, we introduce a generic unsupervised deep learning approach to training deep models without the need for any manual label supervision. Specifically, we progressively discover sample anchored/centred neighbourhoods to reason and learn the underlying class decision boundaries iteratively and accumulatively. Every single neighbourhood is specially formulated so that all the member samples can share the same unseen class labels at high probability for facilitating the extraction of class discriminative feature representations during training. Experiments on image classification show the performance advantages of the proposed method over the state-of-the-art unsupervised learning models on six benchmarks including both coarse-grained and fine-grained object image categorisation. Jiabo Huang, Qi Dong 0004, Shaogang Gong, Xiatian Zhu |
ICML | 3 |
| 2019 | TC-Net for iSBIR: Triplet Classification Network for Instance-level Sketch Based Image RetrievalabstractSketch has been employed as an effective communication tool to express the abstract and intuitive meaning of object. While content-based sketch recognition has been studied for several decades, the instance-level Sketch Based Image Retrieval (iSBIR) task has attracted significant research attention recently. In many previous iSBIR works -- TripletSN, and DSSA, edge maps were employed as intermediate representations in bridging the cross-domain discrepancy between photos and sketches. However, it is nontrivial to efficiently train and effectively use the edge maps in an iSBIR system. Particularly, we find that such an edge map based iSBIR system has several major limitations. First, the system has to be pre-trained on a significant amount of edge maps, either from large-scale sketch datasets, e.g., TU-Berlin~\citeeitz2012hdhso, or converted from other large-scale image datasets, e.g., ImageNet-1K\citedeng2009imagenet dataset. Second, the performance of such an iSBIR system is very sensitive to the quality of edge maps. Third and empirically, the multi-cropping strategy is essentially very important in improving the performance of previous iSBIR systems. To address these limitations, this paper advocates an end-to-end iSBIR system without using the edge maps. Specifically, we present a Triplet Classification Network (TC-Net) for iSBIR which is composed of two major components: triplet Siamese network, and auxiliary classification loss. Our TC-Net can break the limitations existed in previous works. Extensive experiments on several datasets validate the efficacy of the proposed network and system. Yanwei Fu 0001, Shaogang Gong, Xiangyang Xue 0001, Yu-Gang Jiang 0001 |
ACM Multimedia | 4 |
| 2019 | Imbalanced Deep Learning by Minority Class Incremental RectificationabstractModel learning from class imbalanced training data is a long-standing and significant challenge for machine learning. In particular, existing deep learning methods consider mostly either class balanced data or moderately imbalanced data in model training, and ignore the challenge of learning from significantly imbalanced training data. To address this problem, we formulate a class imbalanced deep learning model based on batch-wise incremental minority (sparsely sampled) class rectification by hard sample mining in majority (frequently sampled) classes during model training. This model is designed to minimise the dominant effect of majority classes by discovering sparsely sampled boundaries of minority classes in an iterative batch-wise learning process. To that end, we introduce a Class Rectification Loss (CRL) function that can be deployed readily in deep network architectures. Extensive experimental evaluations are conducted on three imbalanced person attribute benchmark datasets (CelebA, X-Domain, DeepFashion) and one balanced object category benchmark dataset (CIFAR-100). These experimental results demonstrate the performance advantages and model scalability of the proposed batch-wise incremental minority class rectification model over the existing state-of-the-art models for addressing the problem of imbalanced data learning. Qi Dong 0004, Shaogang Gong, Xiatian Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Deep Low-Resolution Person Re-IdentificationabstractPerson images captured by public surveillance cameras often have low resolutions (LR) in addition to uncontrolled pose variations, background clutters and occlusions. This gives rise to the resolution mismatch problem when matched against the high resolution (HR) gallery images (typically available in enrolment), which adversely affects the performance of person re-identification (re-id) that aims to associate images of the same person captured at different locations and different time. Most existing re-id methods either ignore this problem or simply upscale LR images. In this work, we address this problem by developing a novel approach called Super-resolution and Identity joiNt learninG (SING) to simultaneously optimise image super-resolution and person re-id matching. This approach is instantiated by designing a hybrid deep Convolutional Neural Network for improving cross-resolution re-id performance. We further introduce an adaptive fusion algorithm for accommodating multi-resolution LR images. Extensive evaluations show the advantages of our method over related state-of-the-art re-id and super-resolution methods on cross-resolution re-id benchmarks. Jiening Jiao, Wei-Shi Zheng 0001, Ancong Wu, Xiatian Zhu, Shaogang Gong |
AAAI | 5 |
| 2018 | Low-Resolution Face Recognition
Zhiyi Cheng, Xiatian Zhu, Shaogang Gong |
ACCV (3) | 3 |
| 2018 | Self-Referenced Deep Learning
Xu Lan, Xiatian Zhu, Shaogang Gong |
ACCV (2) | 3 |
| 2018 | Deep Association Learning for Unsupervised Video Person Re-identification
Yanbei Chen, Xiatian Zhu, Shaogang Gong |
BMVC | 3 |
| 2018 | Open Logo Detection Challenge
Hang Su 0004, Xiatian Zhu, Shaogang Gong |
BMVC | 3 |
| 2018 | Harmonious Attention Network for Person Re-IdentificationabstractExisting person re-identification (re-id) methods either assume the availability of well-aligned person bounding box images as model input or rely on constrained attention selection mechanisms to calibrate misaligned images. They are therefore sub-optimal for re-id matching in arbitrarily aligned person images potentially with large human pose variations and unconstrained auto-detection errors. In this work, we show the advantages of jointly learning attention selection and feature representation in a Convolutional Neural Network (CNN) by maximising the complementary information of different levels of visual attention subject to re-id discriminative learning constraints. Specifically, we formulate a novel Harmonious Attention CNN (HA-CNN) model for joint learning of soft pixel attention and hard regional attention along with simultaneous optimisation of feature representations, dedicated to optimise person re-id in uncontrolled (misaligned) images. Extensive comparative evaluations validate the superiority of this new HA-CNN model for person re-id over a wide variety of state-of-the-art methods on three large-scale benchmarks including CUHK03, Market-1501, and DukeMTMC-ReID. Wei Li 0132, Xiatian Zhu, Shaogang Gong |
CVPR | 3 |
| 2018 | Transferable Joint Attribute-Identity Deep Learning for Unsupervised Person Re-IdentificationabstractMost existing person re-identification (re-id) methods require supervised model learning from a separate large set of pairwise labelled training data for every single camera pair. This significantly limits their scalability and usability in real-world large scale deployments with the need for performing re-id across many camera views. To address this scalability problem, we develop a novel deep learning method for transferring the labelled information of an existing dataset to a new unseen (unlabelled) target domain for person re-id without any supervised learning in the target domain. Specifically, we introduce an Transferable Joint Attribute-Identity Deep Learning (TJ-AIDL) for simultaneously learning an attribute-semantic and identity-discriminative feature representation space transferrable to any new (unseen) target domain for re-id tasks without the need for collecting new labelled training data from the target domain (i.e. unsupervised learning in the target domain). Extensive comparative evaluations validate the superiority of this new TJ-AIDL model for unsupervised person re-id over a wide range of state-of-the-art methods on four challenging benchmarks including VIPeR, PRID, Market-1501, and DukeMTMC-ReID. Jingya Wang 0001, Xiatian Zhu, Shaogang Gong, Wei Li 0132 |
CVPR | 3 |
| 2018 | Semi-supervised Deep Learning with Memory
Yanbei Chen, Xiatian Zhu, Shaogang Gong |
ECCV (1) | 3 |
| 2018 | Person Search by Multi-Scale Matching
Xu Lan, Xiatian Zhu, Shaogang Gong |
ECCV (1) | 3 |
| 2018 | Unsupervised Person Re-identification by Deep Learning Tracklet Association
Minxian Li, Xiatian Zhu, Shaogang Gong |
ECCV (4) | 3 |
| 2018 | Knowledge Distillation by On-the-Fly Native EnsembleabstractKnowledge distillation is effective to train the small and generalisable network models for meeting the low-memory and fast running requirements. Existing offline distillation methods rely on a strong pre-trained teacher, which enables favourable knowledge discovery and transfer but requires a complex two-phase training procedure. Online counterparts address this limitation at the price of lacking a high-capacity teacher. In this work, we present an On-the-fly Native Ensemble (ONE) learning strategy for one-stage online distillation. Specifically, ONE only trains a single multi-branch network while simultaneously establishing a strong teacher on-the-fly to enhance the learning of target network. Extensive evaluations show that ONE improves the generalisation performance of a variety of deep neural networks more significantly than alternative methods on four image classification dataset: CIFAR10, CIFAR100, SVHN, and ImageNet, whilst having the computational efficiency advantages. Xu Lan, Xiatian Zhu, Shaogang Gong |
NeurIPS | 3 |
| 2018 | Person Re-identification in Identity Regression SpaceabstractMost existing person re-identification (re-id) methods are unsuitable for real-world deployment due to two reasons: Unscalability to large population size , and Inadaptability over time . In this work, we present a unified solution to address both problems. Specifically, we propose to construct an identity regression space (IRS) based on embedding different training person identities (classes) and formulate re-id as a regression problem solved by identity regression in the IRS. The IRS approach is characterised by a closed-form solution with high learning efficiency and an inherent incremental learning capability with human-in-the-loop. Extensive experiments on four benchmarking datasets (VIPeR, CUHK01, CUHK03 and Market-1501) show that the IRS model not only outperforms state-of-the-art re-id methods, but also is more scalable to large re-id population size by rapidly updating model and actively selecting informative samples with reduced human labelling effort. Hanxiao Wang 0001, Xiatian Zhu, Shaogang Gong, Tao Xiang 0002 |
Int. J. Comput. Vis. | 3 |
| 2018 | Zero-Shot Learning on Semantic Class Prototype GraphabstractZero-Shot Learning (ZSL) for visual recognition is typically achieved by exploiting a semantic embedding space. In such a space, both seen and unseen class labels as well as image features can be embedded so that the similarity among them can be measured directly. In this work, we consider that the key to effective ZSL is to compute an optimal distance metric in the semantic embedding space. Existing ZSL works employ either euclidean or cosine distances. However, in a high-dimensional space where the projected class labels (prototypes) are sparse, these distances are suboptimal, resulting in a number of problems including hubness and domain shift. To overcome these problems, a novel manifold distance computed on a semantic class prototype graph is proposed which takes into account the rich intrinsic semantic structure, i.e., semantic manifold, of the class prototype distribution. To further alleviate the domain shift problem, a new regularisation term is introduced into a ranking loss based embedding model. Specifically, the ranking loss objective is regularised by unseen class prototypes to prevent the projected object features from being biased towards the seen prototypes. Extensive experiments on four benchmarks show that our method significantly outperforms the state-of-the-art. Zhenyong Fu, Tao Xiang 0002, Elyor Kodirov, Shaogang Gong |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Deep Reinforcement Learning Attention Selection For Person Re-Identification
Xu Lan, Hangxiao Wang, Shaogang Gong, Xiatian Zhu |
BMVC | 3 |
| 2017 | Semantic Autoencoder for Zero-Shot LearningabstractExisting zero-shot learning (ZSL) models typically learn a projection function from a feature space to a semantic embedding space (e.g. attribute space). However, such a projection function is only concerned with predicting the training seen class semantic representation (e.g. attribute prediction) or classification. When applied to test data, which in the context of ZSL contains different (unseen) classes without training data, a ZSL model typically suffers from the project domain shift problem. In this work, we present a novel solution to ZSL based on learning a Semantic AutoEncoder (SAE). Taking the encoder-decoder paradigm, an encoder aims to project a visual feature vector into the semantic space as in the existing ZSL models. However, the decoder exerts an additional constraint, that is, the projection/code must be able to reconstruct the original visual feature. We show that with this additional reconstruction constraint, the learned projection function from the seen classes is able to generalise better to the new unseen classes. Importantly, the encoder and decoder are linear and symmetric which enable us to develop an extremely efficient learning algorithm. Extensive experiments on six benchmark datasets demonstrate that the proposed SAE outperforms significantly the existing ZSL models with the additional benefit of lower computational cost. Furthermore, when the SAE is applied to supervised clustering problem, it also beats the state-of-the-art. Elyor Kodirov, Tao Xiang 0002, Shaogang Gong |
CVPR | 3 |
| 2017 | Learning a Deep Embedding Model for Zero-Shot LearningabstractZero-shot learning (ZSL) models rely on learning a joint embedding space where both textual/semantic description of object classes and visual representation of object images can be projected to for nearest neighbour search. Despite the success of deep neural networks that learn an end-to-end model between text and images in other vision problems such as image captioning, very few deep ZSL model exists and they show little advantage over ZSL models that utilise deep feature representations but do not learn an end-to-end embedding. In this paper we argue that the key to make deep ZSL models succeed is to choose the right embedding space. Instead of embedding into a semantic space or an intermediate space, we propose to use the visual space as the embedding space. This is because that in this space, the subsequent nearest neighbour search would suffer much less from the hubness problem and thus become more effective. This model design also provides a natural mechanism for multiple semantic modalities (e.g.,~attributes and sentence descriptions) to be fused and optimised jointly in an end-to-end manner. Extensive experiments on four benchmarks show that our model significantly outperforms the existing models. Li Zhang 0040, Tao Xiang 0002, Shaogang Gong |
CVPR | 3 |
| 2017 | Class Rectification Hard Mining for Imbalanced Deep LearningabstractRecognising detailed facial or clothing attributes in images of people is a challenging task for computer vision, especially when the training data are both in very large scale and extremely imbalanced among different attribute classes. To address this problem, we formulate a novel scheme for batch incremental hard sample mining of minority attribute classes from imbalanced large scale training data. We develop an end-to-end deep learning framework capable of avoiding the dominant effect of majority classes by discovering sparsely sampled boundaries of minority classes. This is made possible by introducing a Class Rectification Loss (CRL) regularising algorithm. We demonstrate the advantages and scalability of CRL over existing state-of-the-art attribute recognition and imbalanced data learning models on two large scale imbalanced benchmark datasets, the CelebA facial attribute dataset and the X-Domain clothing attribute dataset. Qi Dong 0004, Shaogang Gong, Xiatian Zhu |
ICCV | 2 |
| 2017 | Attribute Recognition by Joint Recurrent Learning of Context and CorrelationabstractRecognising semantic pedestrian attributes in surveillance images is a challenging task for computer vision, particularly when the imaging quality is poor with complex background clutter and uncontrolled viewing conditions, and the number of labelled training data is small. In this work, we formulate a Joint Recurrent Learning (JRL) model for exploring attribute context and correlation in order to improve attribute recognition given small sized training data with poor quality images. The JRL model learns jointly pedestrian attribute correlations in a pedestrian image and in particular their sequential ordering dependencies (latent high-order correlation) in an end-to-end encoder/ decoder recurrent network. We demonstrate the performance advantage and robustness of the JRL model over a wide range of state-of-the-art deep models for pedestrian attribute recognition, multi-label image classification, and multi-person image annotation on two largest pedestrian attribute benchmarks PETA and RAP. Jingya Wang 0001, Xiatian Zhu, Shaogang Gong, Wei Li 0132 |
ICCV | 3 |
| 2017 | RGB-Infrared Cross-Modality Person Re-identificationabstractPerson re-identification (Re-ID) is an important problem in video surveillance, aiming to match pedestrian images across camera views. Currently, most works focus on RGB-based Re-ID. However, in some applications, RGB images are not suitable, e.g. in a dark environment or at night. Infrared (IR) imaging becomes necessary in many visual systems. To that end, matching RGB images with infrared images is required, which are heterogeneous with very different visual characteristics. For person Re-ID, this is a very challenging cross-modality problem that has not been studied so far. In this work, we address the RGB-IR cross-modality Re-ID problem and contribute a new multiple modality Re-ID dataset named SYSU-MM01, including RGB and IR images of 491 identities from 6 cameras, giving in total 287,628 RGB images and 15,792 IR images. To explore the RGB-IR Re-ID problem, we evaluate existing popular cross-domain models, including three commonly used neural network structures (one-stream, two-stream and asymmetric FC layer) and analyse the relation between them. We further propose deep zero-padding for training one-stream network towards automatically evolving domain-specific nodes in the network for cross-modality matching. Our experiments show that RGB-IR cross-modality matching is very challenging but still feasible using the proposed model with deep zero-padding, giving the best performance. Our dataset is available at http:// isee.sysu.edu.cn/project/RGBIRReID.htm. Ancong Wu, Wei-Shi Zheng 0001, Hong-Xing Yu, Shaogang Gong, Jian-Huang Lai |
ICCV | 4 |
| 2017 | Deep learning prototype domains for person re-identificationabstractPerson re-identification (re-id) is the task of matching multiple occurrences of the same person from different cameras, poses, lighting conditions, and a multitude of other factors which alter the visual appearance. Typically, this is achieved by learning either optimal features or distance metrics which are adapted to specific pairs of camera views dictated by the pairwise labelled training datasets. In this work, we formulate a deep learning based novel approach to automatic prototype-domain discovery for domain perceptive person re-id. The approach scales to new and unseen scenes without requiring new training data. We learn a separate re-id model for each of the discovered prototype-domains and during model deployment, use the person probe image to automatically select the model of the closest prototype-domain. Our approach requires neither supervised nor unsupervised transfer learning, i.e. no data available from target domains. Extensive evaluations are carried out using automatically detected bounding boxes with low-resolution and partial occlusion on two large scale re-id benchmarks, CUHK-SYSU and PRW. Our approach outperforms state-of-the-art unsupervised methods significantly and is competitive against supervised methods which use labelled test domain data. Arne Schumann, Shaogang Gong, Tobias Schuchert |
ICIP | 2 |
| 2017 | Person Re-Identification by Deep Joint Learning of Multi-Loss ClassificationabstractExisting person re-identification (re-id) methods rely mostly on either localised or global feature representation. This ignores their joint benefit and mutual complementary effects. In this work, we show the advantages of jointly learning local and global features in a Convolutional Neural Network (CNN) by aiming to discover correlated local and global features in different context. Specifically, we formulate a method for joint learning of local and global feature selection losses designed to optimise person re-id when using generic matching metrics such as the L2 distance. We design a novel CNN architecture for Jointly Learning Multi-Loss (JLML) of local and global discriminative feature optimisation subject concurrently to the same re-id labelled information. Extensive comparative evaluations demonstrate the advantages of this new JLML model for person re-id over a wide range of state-of-the-art re-id methods on five benchmarks (VIPeR, GRID, CUHK01, CUHK03, Market-1501). Wei Li 0132, Xiatian Zhu, Shaogang Gong |
IJCAI | 3 |
| 2017 | Multi-task Curriculum Transfer Deep Learning of Clothing AttributesabstractRecognising detailed clothing characteristics (finegrained attributes) in unconstrained images of people inthe-wild is a challenging task for computer vision, especially when there is only limited training data from the wild whilst most data available for model learning are captured in well-controlled environments using fashion models (well lit, no background clutter, frontal view, high-resolution). In this work, we develop a deep learning framework capable of model transfer learning from well-controlled shop clothing images collected from web retailers to in-the-wild images from the street. Specifically, we formulate a novel Multi-Task Curriculum Transfer (MTCT) deep learning method to explore multiple sources of different types of web annotations with multi-labelled fine-grained attributes. Our multi-task loss function is designed to extract more discriminative representations in training by jointly learning all attributes, and our curriculum strategy exploits the staged easy-to-hard transfer learning motivated by cognitive studies. We demonstrate the advantages of the MTCT model over the state-of-the-art methods on the X-Domain benchmark, a large scale clothing attribute dataset. Moreover, we show that the MTCT model has a notable advantage over contemporary models when the training data size is small. Qi Dong 0004, Shaogang Gong, Xiatian Zhu |
WACV | 2 |
| 2017 | Deep Learning Logo Detection with Data Expansion by Synthesising ContextabstractLogo detection in unconstrained images is challenging, particularly when only very sparse labelled training images are accessible due to high labelling costs. In this work, we describe a model training image synthesising method capable of improving significantly logo detection performance when only a handful of (e.g., 10) labelled training images captured in realistic context are available, avoiding extensive manual labelling costs. Specifically, we design a novel algorithm for generating Synthetic Context Logo (SCL) training images to increase model robustness against unknown background clutters, resulting in superior logo detection performance. For benchmarking model performance, we introduce a new logo detection dataset TopLogo-10 collected from top 10 most popular clothing/wearable brandname logos captured in rich visual context. Extensive comparisons show the advantages of our proposed SCL model over the state-of-the-art alternatives for logo detection using two real-world logo benchmark datasets: FlickrLogo-32 and our new TopLogo-101. Hang Su 0004, Xiatian Zhu, Shaogang Gong |
WACV | 3 |
| 2017 | Discovering visual concept structure with sparse and incomplete tags
Jingya Wang 0001, Xiatian Zhu, Shaogang Gong |
Artif. Intell. | 3 |
| 2017 | Image and Video Understanding in Big DataabstractAn active object recognition system has the advantage of acting in the environment to capture images that are more suited for training and lead to better performance at test time. In this paper, we utilize deep convolutional neural networks for active object recognition by simultaneously predicting the object label and the next action to be performed on the object with the aim of improving recognition performance. We treat active object recognition as a reinforcement learning problem and derive the cost function to train the network for joint prediction of the object label and the action. A generative model of object similarities based on the Dirichlet distribution is proposed and embedded in the network for encoding the state of the system. The training is carried out by simultaneously minimizing the label and action prediction errors using gradient descent. We empirically show that the proposed network is able to predict both the object label and the actions on GERMS, a dataset for active object recognition. We compare the test label prediction accuracy of the proposed model with Dirichlet and Naive Bayes state encoding. The results of experiments suggest that the proposed model equipped with Dirichlet state encoding is superior in performance, and selects images that lead to better training and higher accuracy of label prediction at test time. Vittorio Murino, Shaogang Gong, Chen Change Loy, Loris Bazzani |
Comput. Vis. Image Underst. | 2 |
| 2017 | Free-Hand Sketch Synthesis with Deformable Stroke ModelsabstractWe present a generative model which can automatically summarize the stroke composition of free-hand sketches of a given category. When our model is fit to a collection of sketches with similar poses, it discovers and learns the structure and appearance of a set of coherent parts, with each part represented by a group of strokes. It represents both consistent (topology) as well as diverse aspects (structure and appearance variations) of each sketch category. Key to the success of our model are important insights learned from a comprehensive study performed on human stroke data. By fitting this model to images, we are able to synthesize visually similar and pleasant free-hand sketches. Yi Li 0004, Yi-Zhe Song, Timothy M. Hospedales, Shaogang Gong |
Int. J. Comput. Vis. | 4 |
| 2017 | Transductive Zero-Shot Action Recognition by Word-Vector Embedding
Xun Xu 0002, Timothy M. Hospedales, Shaogang Gong |
Int. J. Comput. Vis. | 3 |
| 2017 | Person re-identification by unsupervised video matching
Xiatian Zhu, Shaogang Gong, Xudong Xie, Jianming Hu, Kin-Man Lam 0001, Yisheng Zhong |
Pattern Recognit. | 3 |
| 2017 | Discovery of Shared Semantic Spaces for Multiscene Video Query and SummarizationabstractThe growing rate of public space closed-circuit television (CCTV) installations has generated a need for automated methods for exploiting video surveillance data, including scene understanding, query, behavior annotation, and summarization. For this reason, extensive research has been performed on surveillance scene understanding and analysis. However, most studies have considered single scenes or groups of adjacent scenes. The semantic similarity between different but related scenes (e.g., many different traffic scenes of a similar layout) is not generally exploited to improve any automated surveillance tasks and reduce manual effort. Exploiting commonality and sharing any supervised annotations between different scenes is, however, challenging due to the following reason: some scenes are totally unrelated and thus any information sharing between them would be detrimental, whereas others may share only a subset of common activities and thus information sharing is only useful if it is selective. Moreover, semantically similar activities that should be modeled together and shared across scenes may have quite different pixel-level appearances in each scene. To address these issues, we develop a new framework for distributed multiple-scene global understanding that clusters surveillance scenes by their ability to explain each other's behaviors and further discovers which subset of activities are shared versus scene specific within each cluster. We show how to use this structured representation of multiple scenes to improve common surveillance tasks, including scene activity understanding, cross-scene query-by-example, behavior classification with reduced supervised labeling requirements, and video summarization. In each case, we demonstrate how our multiscene model improves on a collection of standard single-scene models and a flat model of all scenes. Xun Xu 0002, Timothy M. Hospedales, Shaogang Gong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Video Semantic Clustering with Sparse and Incomplete TagsabstractClustering tagged videos into semantic groups is importantbut challenging due to the need for jointly learning correlations between heterogeneous visual and tag data. The taskis made more difficult by inherently sparse and incompletetag labels. In this work, we develop a method for accuratelyclustering tagged videos based on a novel Hierarchical-MultiLabel Random Forest model capable of correlating structured visual and tag information. Specifically, our model exploits hierarchically structured tags of different abstractnessof semantics and multiple tag statistical correlations, thus discovers more accurate semantic correlations among differentvideo data, even with highly sparse/incomplete tags. Jingya Wang 0001, Xiatian Zhu, Shaogang Gong |
AAAI | 3 |
| 2016 | Learning Robust Graph Regularisation for Subspace Clustering
Elyor Kodirov, Tao Xiang 0002, Zhenyong Fu, Shaogang Gong |
BMVC | 4 |
| 2016 | Highly Efficient Regression for Scalable Person Re-Identification
Hanxiao Wang 0001, Shaogang Gong, Tao Xiang 0002 |
BMVC | 2 |
| 2016 | Unsupervised Cross-Dataset Transfer Learning for Person Re-identificationabstractMost existing person re-identification (Re-ID) approaches follow a supervised learning framework, in which a large number of labelled matching pairs are required for training. This severely limits their scalability in realworld applications. To overcome this limitation, we develop a novel cross-dataset transfer learning approach to learn a discriminative representation. It is unsupervised in the sense that the target dataset is completely unlabelled. Specifically, we present an multi-task dictionary learning method which is able to learn a dataset-shared but target-data-biased representation. Experimental results on five benchmark datasets demonstrate that the method significantly outperforms the state-of-the-art. Peixi Peng, Tao Xiang 0002, Yaowei Wang 0001, Massimiliano Pontil, Shaogang Gong, Tiejun Huang 0001, Yonghong Tian 0001 |
CVPR | 5 |
| 2016 | Learning a Discriminative Null Space for Person Re-identificationabstractMost existing person re-identification (re-id) methods focus on learning the optimal distance metrics across camera views. Typically a person's appearance is represented using features of thousands of dimensions, whilst only hundreds of training samples are available due to the difficulties in collecting matched training images. With the number of training samples much smaller than the feature dimension, the existing methods thus face the classic small sample size (SSS) problem and have to resort to dimensionality reduction techniques and/or matrix regularisation, which lead to loss of discriminative power. In this work, we propose to overcome the SSS problem in re-id distance metric learning by matching people in a discriminative null space of the training data. In this null space, images of the same person are collapsed into a single point thus minimising the within-class scatter to the extreme and maximising the relative between-class separation simultaneously. Importantly, it has a fixed dimension, a closed-form solution and is very efficient to compute. Extensive experiments carried out on five person re-identification benchmarks including VIPeR, PRID2011, CUHK01, CUHK03 and Market1501 show that such a simple approach beats the state-of-the-art alternatives, often by a big margin. Li Zhang 0040, Tao Xiang 0002, Shaogang Gong |
CVPR | 3 |
| 2016 | Person Re-Identification by Unsupervised \ell _1 ℓ 1 Graph Learning
Elyor Kodirov, Tao Xiang 0002, Zhenyong Fu, Shaogang Gong |
ECCV (1) | 4 |
| 2016 | Human-in-the-Loop Person Re-identification
Hanxiao Wang 0001, Shaogang Gong, Xiatian Zhu, Tao Xiang 0002 |
ECCV (4) | 2 |
| 2016 | Multi-Task Zero-Shot Action Recognition with Prioritised Data Augmentation
Xun Xu 0002, Timothy M. Hospedales, Shaogang Gong |
ECCV (2) | 3 |
| 2016 | Exploring synonyms as context in zero-shot action recognitionabstractZero shot learning (ZSL) provides a solution to recognising unseen classes without class labelled data for model learning. Most ZSL methods aim to learn a mapping from a visual feature space to a semantic embedding space, e.g. attribute or word vector spaces. The use of word vector space is particularly attractive as compared to attribute, it offers vast auxiliary classes with free parts embedding without human annotation. However, using the word vector embedding often provides weaker discriminative power than manually labelled attributes of the auxiliary classes. This is compounded further in zero-shot action recognition due to richer content variations among action classes. In this work we propose to explore a broader semantic contextual information in the text domain to enrich the word vector representation of action classes. We show through extensive experiments that this method improves significantly the performance of a number of existing word vector embedding ZSL methods. Moreover, it also outperforms attribute embedding ZSL with human annotation. Ioannis Alexiou, Tao Xiang 0002, Shaogang Gong |
ICIP | 3 |
| 2016 | Towards unsupervised open-set person re-identificationabstractMost existing person re-identification (ReID) methods assume the availability of extensively labelled cross-view person pairs and a closed-set scenario (i.e. all the probe people exist in the gallery set). These two assumptions significantly limit their usefulness and scalability in real-world applications, particularly with large scale camera networks. To overcome the limitations, we introduce a more challenging yet realistic ReID setting termed OneShot-OpenSet-RelD, and propose a novel Regularised Kernel Subspace Learning model for ReID under this setting. Our model differs significantly from existing ReID methods due to its ability of effectively learning cross-view identity-specific information from unlabelled data alone, and its flexibility of naturally accommodating pairwise labels if available. Hanxiao Wang 0001, Xiatian Zhu, Tao Xiang 0002, Shaogang Gong |
ICIP | 4 |
| 2016 | Learning from Multiple Sources for Video Summarisation
Xiatian Zhu, Chen Change Loy, Shaogang Gong |
Int. J. Comput. Vis. | 3 |
| 2016 | Robust Subjective Visual Property Prediction from Crowdsourced Pairwise LabelsabstractThe problem of estimating subjective visual properties from image and video has attracted increasing interest. A subjective visual property is useful either on its own (e.g. image and video interestingness) or as an intermediate representation for visual recognition (e.g. a relative attribute). Due to its ambiguous nature, annotating the value of a subjective visual property for learning a prediction model is challenging. To make the annotation more reliable, recent studies employ crowdsourcing tools to collect pairwise comparison labels. However, using crowdsourced data also introduces outliers. Existing methods rely on majority voting to prune the annotation outliers/errors. They thus require a large amount of pairwise labels to be collected. More importantly as a local outlier detection method, majority voting is ineffective in identifying outliers that can cause global ranking inconsistencies. In this paper, we propose a more principled way to identify annotation outliers by formulating the subjective visual property prediction task as a unified robust learning to rank problem, tackling both the outlier detection and learning to rank jointly. This differs from existing methods in that (1) the proposed method integrates local pairwise comparison labels together to minimise a cost that corresponds to global inconsistency of ranking order, and (2) the outlier detection and learning to rank problems are solved jointly. This not only leads to better detection of annotation outliers but also enables learning with extremely sparse annotations. Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Jiechao Xiong, Shaogang Gong, Yizhou Wang 0001, Yuan Yao 0011 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2016 | Person Re-Identification by Discriminative Selection in Video RankingabstractCurrent person re-identification (ReID) methods typically rely on single-frame imagery features, whilst ignoring space-time information from image sequences often available in the practical surveillance scenarios. Single-frame (single-shot) based visual appearance matching is inherently limited for person ReID in public spaces due to the challenging visual ambiguity and uncertainty arising from non-overlapping camera views where viewing condition changes can cause significant people appearance variations. In this work, we present a novel model to automatically select the most discriminative video fragments from noisy/incomplete image sequences of people from which reliable space-time and appearance features can be computed, whilst simultaneously learning a video ranking function for person ReID. Using the PRID 2011, iLIDS-VID, and HDA+ image sequence datasets, we extensively conducted comparative evaluations to demonstrate the advantages of the proposed model over contemporary gait recognition, holistic image sequence matching and state-of-the-art single-/multi-shot ReID methods. Taiqing Wang, Shaogang Gong, Xiatian Zhu, Shengjin Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Towards Open-World Person Re-Identification by One-Shot Group-Based VerificationabstractSolving the problem of matching people across non-overlapping multi-camera views, known as person re-identification (re-id), has received increasing interests in computer vision. In a real-world application scenario, a watch-list (gallery set) of a handful of known target people are provided with very few (in many cases only a single) image(s) (shots) per target. Existing re-id methods are largely unsuitable to address this open-world re-id challenge because they are designed for (1) a closed-world scenario where the gallery and probe sets are assumed to contain exactly the same people, (2) person-wise identification whereby the model attempts to verify exhaustively against each individual in the gallery set, and (3) learning a matching model using multi-shots. In this paper, a novel transfer local relative distance comparison (t-LRDC) model is formulated to address the open-world person re-identification problem by one-shot group-based verification. The model is designed to mine and transfer useful information from a labelled open-world non-target dataset. Extensive experiments demonstrate that the proposed approach outperforms both non-transfer learning and existing transfer learning based re-id methods. Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Exemplar-Based Recognition of Human-Object InteractionsabstractHuman action can be recognized from a single still image by modeling human-object interactions (HOIs), which infers the mutual spatial structure information between human and the manipulated object as well as their appearance. Existing approaches rely heavily on accurate detection of human and object and estimation of human pose; they are thus sensitive to large variations of human poses, occlusion, and unsatisfactory detection of small size objects. To overcome this limitation, a novel exemplar-based approach is proposed in this paper. Our approach learns a set of spatial pose-object interaction exemplars, which are probabilistic density functions describing spatially how a person is interacting with a manipulated object for different activities. Specifically, a new framework consisting of an exemplar-based HOI descriptor and an associated matching model is formulated for robust human action recognition in still images. In addition, the framework is extended to perform HOI recognition in videos, where the proposed exemplar representation is used for implicit frame selection to negate irrelevant or noisy frames by temporal structured HOI modeling. Extensive experiments are carried out on two image action datasets and two video action datasets. The results demonstrate the effectiveness of our proposed methods and show that our approach is able to achieve state-of-the-art performance, compared with several recently proposed competitors. Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Shaogang Gong, Tao Xiang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Constrained Clustering With Imperfect OraclesabstractWhile clustering is usually an unsupervised operation, there are circumstances where we have access to prior belief that pairs of samples should (or should not) be assigned with the same cluster. Constrained clustering aims to exploit this prior belief as constraint (or weak supervision) to influence the cluster formation so as to obtain a data structure more closely resembling human perception. Two important issues remain open: 1) how to exploit sparse constraints effectively and 2) how to handle ill-conditioned/noisy constraints generated by imperfect oracles. In this paper, we present a novel pairwise similarity measure framework to address the above issues. Specifically, in contrast to existing constrained clustering approaches that blindly rely on all features for constraint propagation, our approach searches for neighborhoods driven by discriminative feature selection for more effective constraint diffusion. Crucially, we formulate a novel approach to handling the noisy constraint problem, which has been unrealistically ignored in the constrained clustering literature. Extensive comparative results show that our method is superior to the state-of-the-art constrained clustering approaches and can generally benefit existing pairwise similarity-based data clustering algorithms, such as spectral clustering and affinity propagation. Xiatian Zhu, Chen Change Loy, Shaogang Gong |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Dictionary Learning with Iterative Laplacian Regularisation for Unsupervised Person Re-identificationabstractMany existing approaches to person re-identification (Re-ID) are based on supervised learning, which requires hundreds of matching pairs to be labelled for each pair of cameras. This severely limits their scalability for real-world applications. This work aims to overcome this limitation by developing a novel unsupervised Re-ID approach. The approach is based on a new dictionary learning for sparse coding formulation with a graph Laplacian regularisation term whose value is set iteratively. As an unsupervised model, the dictionary learning model is well-suited to the unsupervised task, whilst the regularisation term enables the exploitation of cross-view identity-discriminative information ignored by existing unsupervised Re-ID methods. Importantly this model is also flexible in utilising any labelled data if available. Experiments on two benchmark datasets demonstrate that the proposed approach significantly outperforms the state-of-the-arts. Elyor Kodirov, Tao Xiang 0002, Shaogang Gong |
BMVC | 3 |
| 2015 | Zero-shot object recognition by semantic manifold distanceabstractObject recognition by zero-shot learning (ZSL) aims to recognise objects without seeing any visual examples by learning knowledge transfer between seen and unseen object classes. This is typically achieved by exploring a semantic embedding space such as attribute space or semantic word vector space. In such a space, both seen and unseen class labels, as well as image features can be embedded (projected), and the similarity between them can thus be measured directly. Existing works differ in what embedding space is used and how to project the visual data into the semantic embedding space. Yet, they all measure the similarity in the space using a conventional distance metric (e.g. cosine) that does not consider the rich intrinsic structure, i.e. semantic manifold, of the semantic categories in the embedding space. In this paper we propose to model the semantic manifold in an embedding space using a semantic class label graph. The semantic manifold structure is used to redefine the distance metric in the semantic embedding space for more effective ZSL. The proposed semantic manifold distance is computed using a novel absorbing Markov chain process (AMP), which has a very efficient closed-form solution. The proposed new model improves upon and seamlessly unifies various existing ZSL algorithms. Extensive experiments on both the large scale ImageNet dataset and the widely used Animal with Attribute (AwA) dataset show that our model outperforms significantly the state-of-the-arts. Zhenyong Fu, Tao A. Xiang, Elyor Kodirov, Shaogang Gong |
CVPR | 4 |
| 2015 | Unsupervised Domain Adaptation for Zero-Shot LearningabstractZero-shot learning (ZSL) can be considered as a special case of transfer learning where the source and target domains have different tasks/label spaces and the target domain is unlabelled, providing little guidance for the knowledge transfer. A ZSL method typically assumes that the two domains share a common semantic representation space, where a visual feature vector extracted from an image/video can be projected/embedded using a projection function. Existing approaches learn the projection function from the source domain and apply it without adaptation to the target domain. They are thus based on naive knowledge transfer and the learned projections are prone to the domain shift problem. In this paper a novel ZSL method is proposed based on unsupervised domain adaptation. Specifically, we formulate a novel regularised sparse coding framework which uses the target domain class labels' projections in the semantic space to regularise the learned target domain projection thus effectively overcoming the projection domain shift problem. Extensive experiments on four object and action recognition benchmark datasets show that the proposed ZSL method significantly outperforms the state-of-the-arts. Elyor Kodirov, Tao Xiang 0002, Zhenyong Fu, Shaogang Gong |
ICCV | 4 |
| 2015 | Multi-Scale Learning for Low-Resolution Person Re-IdentificationabstractIn real world person re-identification (re-id), images of people captured at very different resolutions from different locations need be matched. Existing re-id models typically normalise all person images to the same size. However, a low-resolution (LR) image contains much less information about a person, and direct image scaling and simple size normalisation as done in conventional re-id methods cannot compensate for the loss of information. To solve this LR person re-id problem, we propose a novel joint multi-scale learning framework, termed joint multi-scale discriminant component analysis (JUDEA). The key component of this framework is a heterogeneous class mean discrepancy (HCMD) criterion for cross-scale image domain alignment, which is optimised simultaneously with discriminant modelling across multiple scales in the joint learning framework. Our experiments show that the proposed JUDEA framework outperforms existing representative re-id methods as well as other related LR visual matching models applied for the LR person re-id problem. Xiang Li 0032, Wei-Shi Zheng 0001, Tao Xiang 0002, Shaogang Gong |
ICCV | 5 |
| 2015 | Partial Person Re-IdentificationabstractWe address a new partial person re-identification (re-id) problem, where only a partial observation of a person is available for matching across different non-overlapping camera views. This differs significantly from the conventional person re-id setting where it is assumed that the full body of a person is detected and aligned. To solve this more challenging and realistic re-id problem without the implicit assumption of manual body-parts alignment, we propose a matching framework consisting of 1) a local patch-level matching model based on a novel sparse representation classification formulation with explicit patch ambiguity modelling, and 2) a global part-based matching model providing complementary spatial layout information. Our framework is evaluated on a new partial person re-id dataset as well as two existing datasets modified to include partial person images. The results show that the proposed method outperforms significantly existing re-id methods as well as other partial visual matching methods. Wei-Shi Zheng 0001, Xiang Li 0032, Tao Xiang 0002, Shengcai Liao, Jian-Huang Lai, Shaogang Gong |
ICCV | 6 |
| 2015 | Semantic embedding space for zero-shot action recognitionabstractThe number of categories for action recognition is growing rapidly. It is thus becoming increasingly hard to collect sufficient training data to learn conventional models for each category. This issue may be ameliorated by the increasingly popular “zero-shot learning” (ZSL) paradigm. In this framework a mapping is constructed between visual features and a human interpretable semantic description of each category, allowing categories to be recognised in the absence of any training data. Existing ZSL studies focus primarily on image data, and attribute-based semantic representations. In this paper, we address zero-shot recognition in contemporary video action recognition tasks, using semantic word vector space as the common space to embed videos and category labels. This is more challenging because the mapping between the semantic space and space-time features of videos containing complex actions is more complex and harder to learn. We demonstrate that a simple self-training and data augmentation strategy can significantly improve the efficacy of this mapping. Experiments on human action datasets including HMDB51 and UCF101 demonstrate that our approach achieves the state-of-the-art zero-shot action recognition performance. Xun Xu 0002, Timothy M. Hospedales, Shaogang Gong |
ICIP | 3 |
| 2015 | Free-hand sketch recognition by multi-kernel feature learning
Yi Li 0004, Timothy M. Hospedales, Yi-Zhe Song, Shaogang Gong |
Comput. Vis. Image Underst. | 4 |
| 2015 | Transductive Multi-View Zero-Shot LearningabstractMost existing zero-shot learning approaches exploit transfer learning via an intermediate semantic representation shared between an annotated auxiliary dataset and a target dataset with different classes and no annotation. A projection from a low-level feature space to the semantic representation space is learned from the auxiliary dataset and applied without adaptation to the target dataset. In this paper we identify two inherent limitations with these approaches. First, due to having disjoint and potentially unrelated classes, the projection functions learned from the auxiliary dataset/domain are biased when applied directly to the target dataset/domain. We call this problem the projection domain shift problem and propose a novel framework, transductive multi-view embedding, to solve it. The second limitation is the prototype sparsity problem which refers to the fact that for each target class, only a single prototype is available for zero-shot learning given a semantic representation. To overcome this problem, a novel heterogeneous multi-view hypergraph label propagation method is formulated for zero-shot learning in the transductive embedding space. It effectively exploits the complementary information offered by different semantic representations and takes advantage of the manifold structures of multiple representation spaces in a coherent manner. We demonstrate through extensive experiments that the proposed approach (1) rectifies the projection shift between the auxiliary and target domains, (2) exploits the complementarity of multiple semantic representations, (3) significantly outperforms existing methods for both zero-shot and N-shot recognition on three image and video benchmark datasets, and (4) enables novel cross-view annotation tasks. Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Shaogang Gong |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | Open-world Person Re-Identification by Multi-Label Assignment Inference
Brais Cancela, Timothy M. Hospedales, Shaogang Gong |
BMVC | 3 |
| 2014 | Transductive Multi-label Zero-shot Learning
Yanwei Fu 0001, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002, Shaogang Gong |
BMVC | 5 |
| 2014 | Re-id: Hunting Attributes in the Wild
Ryan Layne, Timothy M. Hospedales, Shaogang Gong |
BMVC | 3 |
| 2014 | Intra-category sketch-based image retrieval by matching deformable part models
Yi Li 0004, Timothy M. Hospedales, Yi-Zhe Song, Shaogang Gong |
BMVC | 4 |
| 2014 | Unsupervised Learning of Generative Topic Saliency for Person Re-identification
Hanxiao Wang 0001, Shaogang Gong, Tao Xiang 0002 |
BMVC | 2 |
| 2014 | Constructing Robust Affinity Graphs for Spectral ClusteringabstractSpectral clustering requires robust and meaningful affinity graphs as input in order to form clusters with desired structures that can well support human intuition. To construct such affinity graphs is non-trivial due to the ambiguity and uncertainty inherent in the raw data. In contrast to most existing clustering methods that typically employ all available features to construct affinity matrices with the Euclidean distance, which is often not an accurate representation of the underlying data structures, we propose a novel unsupervised approach to generating more robust affinity graphs via identifying and exploiting discriminative features for improving spectral clustering. Specifically, our model is capable of capturing and combining subtle similarity information distributed over discriminative feature subspaces for more accurately revealing the latent data distribution and thereby leading to improved data clustering, especially with heterogeneous data sources. We demonstrate the efficacy of the proposed approach on challenging image and video datasets. Xiatian Zhu, Chen Change Loy, Shaogang Gong |
CVPR | 3 |
| 2014 | Transductive Multi-view Embedding for Zero-Shot Recognition and Annotation
Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Zhenyong Fu, Shaogang Gong |
ECCV (2) | 5 |
| 2014 | Interestingness Prediction by Robust Learning to Rank
Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Shaogang Gong, Yuan Yao 0011 |
ECCV (2) | 4 |
| 2014 | Person Re-identification by Video Ranking
Taiqing Wang, Shaogang Gong, Xiatian Zhu, Shengjin Wang |
ECCV (4) | 2 |
| 2014 | Learning Multimodal Latent AttributesabstractThe rapid development of social media sharing has created a huge demand for automatic media classification and annotation techniques. Attribute learning has emerged as a promising paradigm for bridging the semantic gap and addressing data sparsity via transferring attribute knowledge in object recognition and relatively simple action classification. In this paper, we address the task of attribute learning for understanding multimedia data with sparse and incomplete labels. In particular, we focus on videos of social group activities, which are particularly challenging and topical examples of this task because of their multimodal content and complex and unstructured nature relative to the density of annotations. To solve this problem, we 1) introduce a concept of semilatent attribute space, expressing user-defined and latent attributes in a unified framework, and 2) propose a novel scalable probabilistic topic model for learning multimodal semilatent attributes, which dramatically reduces requirements for an exhaustive accurate attribute ontology and expensive annotation effort. We show that our framework is able to exploit latent attributes to outperform contemporary approaches for addressing a variety of realistic multimedia sparse data learning tasks including: multitask learning, learning with label noise, N-shot transfer learning, and importantly zero-shot learning. Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Shaogang Gong |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | On-the-fly feature importance mining for person re-identification
Shaogang Gong, Chen Change Loy |
Pattern Recognit. | 2 |
| 2013 | Sketch Recognition by Ensemble Matching of Structured FeaturesabstractSketch recognition aims to automatically classify human hand sketches of objects into known categories. This has become increasingly a desirable capability due to recent advances in human computer interaction on portable devices. The problem is nontrivial because of the sparse and abstract nature of hand drawings as compared to photographic images of objects, compounded by a highly variable degree of details in human sketches. To this end, we present a method for the representation and matching of sketches by exploiting not only local features but also global structures of sketches, through a star graph based ensemble matching strategy. Different local feature representations were evaluated using the star graph model to demonstrate the effectiveness of the ensemble matching of structured features. We further show that by encapsulating holistic structure matching and learned bag-of-features models into a single framework, notable recognition performance improvement over the state-of-the-art can be observed. Extensive comparative experiments were carried out using the currently largest sketch dataset released by Eitz et al. [15], with over 20,000 sketches of 250 object categories generated by AMT (Amazon Mechanical Turk) crowd-sourcing. Yi Li 0004, Yi-Zhe Song, Shaogang Gong |
BMVC | 3 |
| 2013 | Cumulative Attribute Space for Age and Crowd Density EstimationabstractA number of computer vision problems such as human age estimation, crowd density estimation and body/face pose (view angle) estimation can be formulated as a regression problem by learning a mapping function between a high dimensional vector-formed feature input and a scalar-valued output. Such a learning problem is made difficult due to sparse and imbalanced training data and large feature variations caused by both uncertain viewing conditions and intrinsic ambiguities between observable visual features and the scalar values to be estimated. Encouraged by the recent success in using attributes for solving classification problems with sparse training data, this paper introduces a novel cumulative attribute concept for learning a regression model when only sparse and imbalanced data are available. More precisely, low-level visual features extracted from sparse and imbalanced image samples are mapped onto a cumulative attribute space where each dimension has clearly defined semantic interpretation (a label) that captures how the scalar output value (e.g. age, people count) changes continuously and cumulatively. Extensive experiments show that our cumulative attribute framework gains notable advantage on accuracy for both age estimation and crowd counting when compared against conventional regression models, especially when the labelled training data is sparse with imbalanced sampling. Ke Chen 0004, Shaogang Gong, Tao Xiang 0002, Chen Change Loy |
CVPR | 2 |
| 2013 | Recognising Human-Object Interaction via Exemplar Based ModellingabstractHuman action can be recognised from a single still image by modelling Human-object interaction (HOI), which infers the mutual spatial structure information between human and object as well as their appearance. Existing approaches rely heavily on accurate detection of human and object, and estimation of human pose. They are thus sensitive to large variations of human poses, occlusion and unsatisfactory detection of small size objects. To overcome this limitation, a novel exemplar based approach is proposed in this work. Our approach learns a set of spatial pose-object interaction exemplars, which are density functions describing how a person is interacting with a manipulated object for different activities spatially in a probabilistic way. A representation based on our HOI exemplar thus has great potential for being robust to the errors in human/object detection and pose estimation. A new framework consists of a proposed exemplar based HOI descriptor and an activity specific matching model that learns the parameters is formulated for robust human activity recognition. Experiments on two benchmark activity datasets demonstrate that the proposed approach obtains state-of-the-art performance. Jianfang Hu, Wei-Shi Zheng 0001, Jian-Huang Lai, Shaogang Gong, Tao Xiang 0002 |
ICCV | 4 |
| 2013 | POP: Person Re-identification Post-rank OptimisationabstractOwing to visual ambiguities and disparities, person re-identification methods inevitably produce sub optimal rank-list, which still requires exhaustive human eyeballing to identify the correct target from hundreds of different likely-candidates. Existing re-identification studies focus on improving the ranking performance, but rarely look into the critical problem of optimising the time-consuming and error-prone post-rank visual search at the user end. In this study, we present a novel one-shot Post-rank Optimization (POP) method, which allows a user to quickly refine their search by either "one-shot" or a couple of sparse negative selections during a re-identification process. We conduct systematic behavioural studies to understand user's searching behaviour and show that the proposed method allows correct re-identification to converge 2.6 times faster than the conventional exhaustive search. Importantly, through extensive evaluations we demonstrate that the method is capable of achieving significant improvement over the state-of-the-art distance metric learning based ranking models, even with just "one shot" feedback optimisation, by as much as over 30% performance improvement for rank 1 re-identification on the VIPeR and i-LIDS datasets. Chen Change Loy, Shaogang Gong, Guijin Wang |
ICCV | 3 |
| 2013 | From Semi-supervised to Transfer Counting of CrowdsabstractRegression-based techniques have shown promising results for people counting in crowded scenes. However, most existing techniques require expensive and laborious data annotation for model training. In this study, we propose to address this problem from three perspectives: (1) Instead of exhaustively annotating every single frame, the most informative frames are selected for annotation automatically and actively. (2) Rather than learning from only labelled data, the abundant unlabelled data are exploited. (3) Labelled data from other scenes are employed to further alleviate the burden for data annotation. All three ideas are implemented in a unified active and semi-supervised regression framework with ability to perform transfer learning, by exploiting the underlying geometric structure of crowd patterns via manifold analysis. Extensive experiments validate the effectiveness of our approach. Chen Change Loy, Shaogang Gong, Tao Xiang 0002 |
ICCV | 2 |
| 2013 | Video Synopsis by Heterogeneous Multi-source CorrelationabstractGenerating coherent synopsis for surveillance video stream remains a formidable challenge due to the ambiguity and uncertainty inherent to visual observations. In contrast to existing video synopsis approaches that rely on visual cues alone, we propose a novel multi-source synopsis framework capable of correlating visual data and independent non-visual auxiliary information to better describe and summarise subtle physical events in complex scenes. Specifically, our unsupervised framework is capable of seamlessly uncovering latent correlations among heterogeneous types of data sources, despite the non-trivial heteroscedasticity and dimensionality discrepancy problems. Additionally, the proposed model is robust to partial or missing non-visual information. We demonstrate the effectiveness of our framework on two crowded public surveillance datasets. Xiatian Zhu, Chen Change Loy, Shaogang Gong |
ICCV | 3 |
| 2013 | Constrained Clustering: Effective Constraint Propagation with Imperfect OraclesabstractWhile spectral clustering is usually an unsupervised operation, there are circumstances in which we have prior belief that pairs of samples should (or should not) be assigned with the same cluster. Constrained spectral clustering aims to exploit this prior belief as constraint (or weak supervision) to influence the cluster formation so as to obtain a structure more closely resembling human perception. Two important issues remain open: (1) how to propagate sparse constraints effectively, (2) how to handle ill-conditioned/noisy constraints generated by imperfect oracles. In this paper we present a unified framework to address the above issues. Specifically, in contrast to existing constrained spectral clustering approaches that blindly rely on all features for constructing the spectral, our approach searches for neighbours driven by discriminative feature selection for more effective constraint diffusion. Crucially, we formulate a novel data-driven filtering approach to handle the noisy constraint problem, which has been unrealistically ignored in constrained spectral clustering literature. Xiatian Zhu, Chen Change Loy, Shaogang Gong |
ICDM | 3 |
| 2013 | Person re-identification by manifold rankingabstractExisting person re-identification methods conventionally rely on labelled pairwise data to learn a task-specific distance metric for ranking. The value of unlabelled gallery instances is generally overlooked. In this study, we show that it is possible to propagate the query information along the unlabelled data manifold in an unsupervised way to obtain robust ranking results. In addition, we demonstrate that the performance of existing supervised metric learning methods can be significantly boosted once integrated into the proposed manifold ranking-based framework. Extensive evaluation is conducted on three benchmark datasets. Chen Change Loy, Shaogang Gong |
ICIP | 3 |
| 2013 | Reidentification by Relative Distance ComparisonabstractMatching people across nonoverlapping camera views at different locations and different times, known as person reidentification, is both a hard and important problem for associating behavior of people observed in a large distributed space over a prolonged period of time. Person reidentification is fundamentally challenging because of the large visual appearance changes caused by variations in view angle, lighting, background clutter, and occlusion. To address these challenges, most previous approaches aim to model and extract distinctive and reliable visual features. However, seeking an optimal and robust similarity measure that quantifies a wide range of features against realistic viewing conditions from a distance is still an open and unsolved problem for person reidentification. In this paper, we formulate person reidentification as a relative distance comparison (RDC) learning problem in order to learn the optimal similarity measure between a pair of person images. This approach avoids treating all features indiscriminately and does not assume the existence of some universally distinctive and reliable features. To that end, a novel relative distance comparison model is introduced. The model is formulated to maximize the likelihood of a pair of true matches having a relatively smaller distance than that of a wrong match pair in a soft discriminant manner. Moreover, in order to maintain the tractability of the model in large scale learning, we further develop an ensemble RDC model. Extensive experiments on three publicly available benchmarking datasets are carried out to demonstrate the clear superiority of the proposed RDC models over related popular person reidentification techniques. The results also show that the new RDC models are more robust against visual appearance changes and less susceptible to model overfitting compared to other related existing models. Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Finding Rare Classes: Active Learning with Generative and Discriminative ModelsabstractDiscovering rare categories and classifying new instances of them are important data mining issues in many fields, but fully supervised learning of a rare class classifier is prohibitively costly in labeling effort. There has therefore been increasing interest both in active discovery: to identify new classes quickly, and active learning: to train classifiers with minimal supervision. These goals occur together in practice and are intrinsically related because examples of each class are required to train a classifier. Nevertheless, very few studies have tried to optimise them together, meaning that data mining for rare classes in new domains makes inefficient use of human supervision. Developing active learning algorithms to optimise both rare class discovery and classification simultaneously is challenging because discovery and classification have conflicting requirements in query criteria. In this paper, we address these issues with two contributions: a unified active learning model to jointly discover new categories and learn to classify them by adapting query criteria online; and a classifier combination algorithm that switches generative and discriminative classifiers as learning progresses. Extensive evaluation on a batch of standard UCI and vision data sets demonstrates the superiority of this approach over existing methods. Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | Feature Mining for Localised Crowd CountingabstractThis paper presents a multi-output regression model for crowd counting in public scenes. Existing counting by regression methods either learn a single model for global counting, or train a large number of separate regressors for localised density estimation. In contrast, our single regression model based approach is able to estimate people count in spatially localised regions and is more scalable without the need for training a large number of regressors proportional to the number of local regions. In particular, the proposed model automatically learns the functional mapping between interdependent low-level features and multi-dimensional structured outputs. The model is able to discover the inherent importance of different features for people counting at different spatial locations. Extensive evaluations on an existing crowd analysis benchmark dataset and a new more challenging dataset demonstrate the effectiveness of our approach. 1 Ke Chen 0004, Chen Change Loy, Shaogang Gong, Tony Xiang |
BMVC | 3 |
| 2012 | Person Re-identification by AttributesabstractVisually identifying a target individual reliably in a crowded environment observed by a distributed camera network is critical to a variety of tasks in managing business information, border control, and crime prevention. Automatic re-identification of a human candidate from public space CCTV video is challenging due to spatiotemporal visual feature variations and strong visual similarity between different people, compounded by low-resolution and poor quality video data. In this work, we propose a novel method for re-identification that learns a selection and weighting of mid-level semantic attributes to describe people. Specifically, the model learns an attribute-centric, parts-based feature representation. This differs from and complements existing low-level features for re-identification that rely purely on bottom-up statistics for feature selection, which are limited in discriminating and identifying reliably visual appearances of target people appearing in different camera views under certain degrees of occlusion due to crowdedness. Our experiments demonstrate the effectiveness of our approach compared to existing feature representations when applied to benchmarking datasets. 1 Ryan Layne, Timothy M. Hospedales, Shaogang Gong |
BMVC | 3 |
| 2012 | Comparing Visual Feature Coding for Learning Disjoint Camera DependenciesabstractProblem: This work systematically investigates the effectiveness of various visual feature coding schemes for facilitating the learning of timedelayed dependencies among disjoint multi-camera views. Related work: Quite a few studies [3, 4, 6] have been proposed to model inter-camera dependency across non-overlapping camera views. Learning time-delayed correlations among disjoint cameras in crowded public scenarios is a non-trivial task: (1) the time gaps between camera views are unknown therefore activities in two related views may occur at arbitrary time delays with high uncertainty; (2) the features are inevitably noisy, ambiguous, and may vary drastically across views owning to illumination condition, camera angles, and changes in object pose. Most state-ofthe-art methods typically hand pick a few features tailored to the target environment, with the hope that those chosen features contain robust and sufficient statistics for correlating the time-delayed activity patterns across disjoint views. These manual approaches to hard selection of features are neither principled nor generalisable to different scene context. Our solution: In this study, we wish to examine the concept that visual features should be coded and selected automatically for robust and accurate time-delayed dependency learning. The contributions of this study are two-fold: (1) We present a systematic study and evaluation to investigate the effectiveness of supervised and unsupervised feature coding methods to facilitate the learning of inter-camera activity pattern dependencies. (2) We systematically evaluate the sensitivity of inter-camera time delayed dependency learning given different training video sizes and region decomposition qualities. These factors are critical for accurate dependency learning but have been largely ignored by the published existing work in the literature. Approach overview: We employ the Random Forest [2] as the supervised feature coding approach. In particular, given a set of localised features extracted from a region, together with people count training label over time, we first train a regression forest to learn the non-linear mapping between the crowd density and the corresponding low-level features. Given unseen data, we then construct a time series based on the predicted crowd density ŷ obtained from the regression forest (RF pred), the treestructured code (tree code) [5], or the combination of the two. As for unsupervised coding scheme, we use the Latent Dirichlet Allocation (LDA) [1] to map the low-level features into codewords that capture the topic distribution, whereby an image region patch (document) d is treated as a collection of j = 1 . . .Ni features (words). To form the unsupervised feature codes, given a sequence of localised feature vectors detected from a region, we first perform quantisation on each feature to generate a bag-of-word representation for all image patches. Similar to text documents, these bag-of-word represented image patches are fed into the LDA, which gives us a topic-based representation. Once having the topic-based code (topic code), we perform k-means quantisation on them, producing the final compact topic-based code, and concatenate them over time to form a time series. To solve the problem of using the feature codes for learning intercamera dependencies, we adopt the Time Delayed Mutual Information (TDMI) proposed in [3] due to its reported effectiveness and simplicity. The input to TDMI are time series generated from either the supervised or the unsupervised coding scheme. In addition to measuring deviation error in transition time, we propose a new metrics to evaluate the effectiveness of different coding methods, called Mutual Information Margin (MIM): Xiatian Zhu, Shaogang Gong, Chen Change Loy |
BMVC | 2 |
| 2012 | Stream-based joint exploration-exploitation active learningabstractLearning from streams of evolving and unbounded data is an important problem, for example in visual surveillance or internet scale data. For such large and evolving real-world data, exhaustive supervision is impractical, particularly so when the full space of classes is not known in advance therefore joint class discovery (exploration) and boundary learning (exploitation) becomes critical. Active learning has shown promise in jointly optimising exploration-exploitation with minimal human supervision. However, existing active learning methods either rely on heuristic multi-criteria weighting or are limited to batch processing. In this paper, we present a new unified framework for joint exploration-exploitation active learning in streams without any heuristic weighting. Extensive evaluation on classification of various image and surveillance video datasets demonstrates the superiority of our framework over existing methods. Chen Change Loy, Timothy M. Hospedales, Tao Xiang 0002, Shaogang Gong |
CVPR | 4 |
| 2012 | Transfer re-identification: From person to set-based verificationabstractSolving the person re-identification problem has become important for understanding people's behaviours in a multicamera network of non-overlapping views. In this work, we address the problem of re-identification from a set-based verification perspective. More specifically, we have a small set of target people on a watch list (a set) and we aim to verify whether a query image of a person is on this watch list. This differs from the existing person re-identification problem in that the probe is verified against a small set of known people but requires much higher degree of verification accuracy with very limited sampling data for each candidate in the set. That is, rather than recognising everybody in the scene, we consider identifying a small set of target people against non-target people when there is only a limited number of target training samples and a large number of unlabelled (unknown) non-target samples available. To this end, we formulate a transfer learning framework for mining discriminant information from non-target people data to solve the watch list set verification problem. Based on the proposed approach, we introduce the concepts of multi-shot and one-shot verifications. We also design new criteria for evaluating the performance of the proposed transfer learning method against the i-LIDS and ETHZ data sets. Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
CVPR | 2 |
| 2012 | Attribute Learning for Understanding Unstructured Social Activity
Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Shaogang Gong |
ECCV (4) | 4 |
| 2012 | A Unifying Theory of Active Discovery and Learning
Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
ECCV (5) | 2 |
| 2012 | Video Behaviour Mining Using a Dynamic Topic Model
Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
Int. J. Comput. Vis. | 2 |
| 2012 | Learning Behavioural Context
Shaogang Gong, Tao Xiang 0002 |
Int. J. Comput. Vis. | 2 |
| 2012 | Incremental Activity Modeling in Multiple Disjoint CamerasabstractActivity modeling and unusual event detection in a network of cameras is challenging, particularly when the camera views are not overlapped. We show that it is possible to detect unusual events in multiple disjoint cameras as context-incoherent patterns through incremental learning of time delayed dependencies between distributed local activities observed within and across camera views. Specifically, we model multicamera activities using a Time Delayed Probabilistic Graphical Model (TD-PGM) with different nodes representing activities in different decomposed regions from different views and the directed links between nodes encoding their time delayed dependencies. To deal with visual context changes, we formulate a novel incremental learning method for modeling time delayed dependencies that change over time. We validate the effectiveness of the proposed approach using a synthetic data set and videos captured from a camera network installed at a busy underground station. Chen Change Loy, Tao Xiang 0002, Shaogang Gong |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Quantifying and Transferring Contextual Information in Object DetectionabstractContext is critical for reducing the uncertainty in object detection. However, context modeling is challenging because there are often many different types of contextual information coexisting with different degrees of relevance to the detection of target object(s) in different images. It is therefore crucial to devise a context model to automatically quantify and select the most effective contextual information for assisting in detecting the target object. Nevertheless, the diversity of contextual information means that learning a robust context model requires a larger training set than learning the target object appearance model, which may not be available in practice. In this work, a novel context modeling framework is proposed without the need for any prior scene segmentation or context annotation. We formulate a polar geometric context descriptor for representing multiple types of contextual information. In order to quantify context, we propose a new maximum margin context (MMC) model to evaluate and measure the usefulness of contextual information directly and explicitly through a discriminant context inference method. Furthermore, to address the problem of context learning with limited data, we exploit the idea of transfer learning based on the observation that although two categories of objects can have very different visual appearance, there can be similarity in their context and/or the way contextual information helps to distinguish target objects from nontarget objects. To that end, two novel context transfer learning models are proposed which utilize training samples from source object classes to improve the learning of the context model for a target object class based on a joint maximum margin learning framework. Experiments are carried out on PASCAL VOC2005 and VOC2007 data sets, a luggage detection data set extracted from the i-LIDS data set, and a vehicle detection data set extracted from outdoor surveillance footage. Our results validate the effectiveness of the proposed models for quantifying and transferring contextual information, and demonstrate that they outperform related alternative context models. Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Fusing appearance and distribution information of interest points for action recognition
Matteo Bregonzio, Tao Xiang 0002, Shaogang Gong |
Pattern Recognit. | 3 |
| 2011 | Person re-identification by probabilistic relative distance comparisonabstractMatching people across non-overlapping camera views, known as person re-identification, is challenging due to the lack of spatial and temporal constraints and large visual appearance changes caused by variations in view angle, lighting, background clutter and occlusion. To address these challenges, most previous approaches aim to extract visual features that are both distinctive and stable under appearance changes. However, most visual features and their combinations under realistic conditions are neither stable nor distinctive thus should not be used indiscriminately. In this paper, we propose to formulate person re-identification as a distance learning problem, which aims to learn the optimal distance that can maximises matching accuracy regardless the choice of representation. To that end, we introduce a novel Probabilistic Relative Distance Comparison (PRDC) model, which differs from most existing distance learning methods in that, rather than minimising intra-class variation whilst maximising intra-class variation, it aims to maximise the probability of a pair of true match having a smaller distance than that of a wrong match pair. This makes our model more tolerant to appearance changes and less susceptible to model over-fitting. Extensive experiments are carried out to demonstrate that 1) by formulating the person re-identification problem as a distance learning problem, notable improvement on matching accuracy can be obtained against conventional person re-identification techniques, which is particularly significant when the training sample size is small; and 2) our PRDC outperforms not only existing distance learning methods but also alternative learning methods based on boosting and learning to rank. Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
CVPR | 2 |
| 2011 | Learning Tags from Unsegmented Videos of Multiple Human ActionsabstractProviding methods to support semantic interaction with growing volumes of video data is an increasingly important challenge for data mining. To this end, there has been some success in recognition of simple objects and actions in video, however most of this work requires strongly supervised training data. The supervision cost of these approaches therefore renders them economically non-scalable for real world applications. In this paper we address the problem of learning to annotate and retrieve semantic tags of human actions in realistic video data with sparsely provided tags of semantically salient activities. This is challenging because of (1) the multi-label nature of the learning problem and (2) realistic videos are often dominated by (semantically uninteresting) background activity un-supported by any tags of interest, leading to a strong irrelevant data problem. To address these challenges, we introduce a new topic model based approach to video tag annotation. Our model simultaneously learns a low dimensional representation of the video data, which dimensions are semantically relevant (supported by tags), and how to annotate videos with tags. Experimental evaluation on three different video action/activity datasets demonstrate the challenge of this problem, and value of our contribution. Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
ICDM | 2 |
| 2011 | Finding Rare Classes: Adapting Generative and Discriminative Models in Active Learning
Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
PAKDD (2) | 2 |
| 2011 | Identifying Rare and Subtle Behaviors: A Weakly Supervised Joint Topic ModelabstractOne of the most interesting and desired capabilities for automated video behavior analysis is the identification of rarely occurring and subtle behaviors. This is of practical value because dangerous or illegal activities often have few or possibly only one prior example to learn from and are often subtle. Rare and subtle behavior learning is challenging for two reasons: (1) Contemporary modeling approaches require more data and supervision than may be available and (2) the most interesting and potentially critical rare behaviors are often visually subtle-occurring among more obvious typical behaviors or being defined by only small spatio-temporal deviations from typical behaviors. In this paper, we introduce a novel weakly supervised joint topic model which addresses these issues. Specifically, we introduce a multiclass topic model with partially shared latent structure and associated learning and inference algorithms. These contributions will permit modeling of behaviors from as few as one example, even without localization by the user and when occurring in clutter, and subsequent classification and localization of such behaviors online and in real time. We extensively validate our approach on two standard public-space data sets, where it clearly outperforms a batch of contemporary alternatives. Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Detecting and discriminating behavioural anomalies
Chen Change Loy, Tao Xiang 0002, Shaogang Gong |
Pattern Recognit. | 3 |
| 2010 | Learning Rare Behaviours
Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
ACCV (2) | 3 |
| 2010 | Stream-Based Active Unusual Event Detection
Chen Change Loy, Tao Xiang 0002, Shaogang Gong |
ACCV (1) | 3 |
| 2010 | Unsupervised Selective Transfer Learning for Object Recognition
Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
ACCV (2) | 2 |
| 2010 | Cross View Gait Recognition Using Correlation StrengthabstractAmong various factors that can affect the performance of gait recognition, changes in viewpoint pose the biggest problem. In this work, we develop a novel approach to cross-view gait recognition with the view angle of a probe gait sequence unknown. We formulate a Gaussian Process (GP) classification framework to estimate the view angle of each probe gait sequence. To measure the similarity of gait sequences captured at different view angles, we model the correlation of gait sequences from different views using Canonical Correlation Analysis (CCA) and use the correlation strength as similarity measure. This differs significantly from existing approaches, which reconstruct gait features in different views either through 2D view transformation or 3D calibration. Without explicit reconstruction, our approach can cope with feature mis-match across view and is more robust against feature noise. Our experiments validate that the proposed method significantly outperforms the existing state-of-the-art methods. Khalid Bashir, Tao Xiang 0002, Shaogang Gong |
BMVC | 3 |
| 2010 | Discriminative Topics Modelling for Action Feature Selection and RecognitionabstractProblem This paper addresses the problem of recognising realistic human actions captured in unconstrained environments (Fig. 1). Existing approaches for action recognition have been focused on improving visual feature representation using either spatio-temporal interest points or key-points trajectories. However, these methods are insufficient to handle the situations when action videos are recorded in unconstrained environments because: (1) Reliable visual features are hard to be extracted due to occlusions, illumination change, scale variation and background clutters. (2) Effectiveness of visual features are strongly dependent on the unpredictable characteristics of camera movements. (3) Complicated visual actions result in unequal discriminativeness of visual features. Our Solutions In this paper, we present a novel framework for recognising realistic human actions in unconstrained environments. The novelties of our work lie in three aspects: First, we propose a new action representation based on computing a rich set of descriptors from key point trajectories. Second, in order to cope with drastic changes in motion characteristics with and without camera movements, we develop an adaptive feature fusion method to combine different local motion descriptors for improving model robustness against feature noise and background clutters. Finally, we propose a novel Multi-Class Delta Latent Dirichlet Allocation (MC-∆LDA) model for feature selection. The most informative features in a high dimensional feature space are selected collaboratively rather than independently. Motion Descriptors We first compute trajectories of key-points using KLT tracker and SIFT matching. After trajectory pruning by identifying the Region of Interest (ROI), we compute three types of motion descriptors from the survived trajectories. First, Orientation-Magnitude Descriptor is extracted by quantising orientation and magnitude of motion between two consecutive points in the same trajectory. Second, Trajectory Shape Descriptor is extracted by computing Fourier coefficients of a single trajectory. Finally, Appearance Descriptor is extracted by computing the SIFT features at all points of a trajectory. Interest Point Features We also detect spatio-temporal interest points as they contain complementary information to trajectory features. At an interest point, a surrounding 3D cuboid is extracted. We use gradient vectors to describe these cuboids and PCA to reduce descriptor’s dimensionality. Adaptive Feature Fusion We wish to fuse adaptively trajectory based descriptors with 3D interest point based descriptors according the presence of camera movement. The presence of moving camera is detected by computing the global optical flow over all frames in a clip. If the majority of the frames contain global motion, we regard the clip as being recorded by a moving camera. For clips without camera movement, both interest point and trajectory based descriptors can be computed reliably and thus both types of descriptors are used for recognition. In contrast, when camera motion can be detected, interest point based descriptors are less meaningful so only trajectory descriptors are employed. Collaborative Feature Selection We propose a MC-∆LDA model (Fig. 2) for collaboratively selecting dominant features for classification. We consider each video clip x j is a mixture of Nt topics Φ = {φt}t t=1 (to be discovered), each of which φt is a multinomial distribution over Nw words (visual features). The MC-∆LDA model aims to constrain topic proportion non-uniformly and on a per-clip basis. For each video clip belonging to action category Ac, we model it as a mixture of: (1) Ns t topics which are shared by all Nc category of actions, and (2) Nt,c topics which are uniquely associated with action category Ac. In MC-∆LDA, the nonuniform proportion of topic mixture for a single clip x j is enforced by its action class label c j and the hyperparameter αc for the corresponding action class c. Given the total number of topics Nt = Ns t +∑ Nc c=1 Nt,c, the structure of the MC-∆LDA model, and the observable variables (clips x j and action labels c j), we can learn the Ns t shared topics as well as all ∑c c=1 Nt,c unique topics for all Nc classes of actions. We use the N s t topics shared by all actions for selecting discriminative features. The Ns t shared topics are represented as an Nw×N t dimension matrix Φs. The feature selection can be summarised into two steps: (1) For each feature vk, k = Figure 1: Actions captured in an unconstrained environments, YouTube dataset. From left to right: cycling, diving, soccer juggling, and walking with a dog. Matteo Bregonzio, Shaogang Gong, Tao Xiang 0002 |
BMVC | 3 |
| 2010 | Person Re-Identification by Support Vector RankingabstractSolving the person re-identification problem involves matching observations of individuals across disjoint camera views. The problem becomes particularly hard in a busy public scene as the number of possible matches is very high. This is further compounded by significant appearance changes due to varying lighting conditions, viewing angles and body poses across camera views. To address this problem, existing approaches focus on extracting or learning discriminative features followed by template matching using a distance measure. The novelty of this work is that we reformulate the person reidentification problem as a ranking problem and learn a subspace where the potential true match is given highest ranking rather than any direct distance measure. By doing so, we convert the person re-identification problem from an absolute scoring problem to a relative ranking problem. We further develop an novel Ensemble RankSVM to overcome the scalability limitation problem suffered by existing SVM-based ranking methods. This new model reduces significantly memory usage therefore is much more scalable, whilst maintaining high-level performance. We present extensive experiments to demonstrate the performance gain of the proposed ranking approach over existing template matching and classification models. 1 Bryan James Prosser, Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
BMVC | 3 |
| 2010 | Action categorization by structural probabilistic latent semantic analysis
Jianguo Zhang 0001, Shaogang Gong |
Comput. Vis. Image Underst. | 2 |
| 2010 | Time-Delayed Correlation Analysis for Multi-Camera Activity Understanding
Chen Change Loy, Tao Xiang 0002, Shaogang Gong |
Int. J. Comput. Vis. | 3 |
| 2010 | Action categorization with modified hidden conditional random field
Jianguo Zhang 0001, Shaogang Gong |
Pattern Recognit. | 2 |
| 2010 | Gait recognition without subject cooperation
Khalid Bashir, Tao Xiang 0002, Shaogang Gong |
Pattern Recognit. Lett. | 3 |
| 2009 | Gait Representation Using Flow FieldsabstractGait is characterised by the relative motions between different body parts during walking. However, most recently proposed gait representation approaches such as Gait Energy Image (GEI) and Motion Silhouettes Image (MSI) capture only the motion intensity information whilst ignoring the equally important but less reliable information about the direction of relative motion. They thus essentially sacrifice discriminative power in exchange for robustness. In this paper, we propose a novel gait representation based on optical flow fields computed from normalized and centred person images over a complete gait cycle. In our representation, both the motion intensity and the motion direction information is captured in a set of motion descriptors. To achieve robustness against noise, instead of relying on the exact value of the flow vectors, the flow direction is discretised and a histogram based direction representation is formulated. Compared to the existing model-free gait representations, our representation is not only more discriminative, but also less sensitive to changes in various covariate conditions including clothing, carrying, shoe, and speed. Extensive experiments on both indoor and outdoor public datasets have been carried out to demonstrate that our representation outperforms the state-of-the-art. Khalid Bashir, Tao Xiang 0002, Shaogang Gong |
BMVC | 3 |
| 2009 | Modelling Multi-object Activity by Gaussian ProcessesabstractWe present a new approach for activity modelling and anomaly detection based on non-parametric Gaussian Process (GP) models. Specifically, GP regression models are formulated to learn non-linear relationships between multi-object activity patterns observed from semantically decomposed regions in complex scenes. Predictive distributions are inferred from the regression models to compare with the actual observations for real-time anomaly detection. The use of a flexible, non-parametric model alleviates the difficult problem of selecting appropriate model complexity encountered in parametric models such as Dynamic Bayesian Networks (DBNs). Crucially, our GP models need fewer parameters; they are thus less likely to overfit given sparse data. In addition, our approach is robust to the inevitable noise in activity representation as noise is modelled explicitly in the GP models. Experimental results on a public traffic scene show that our models outperform DBNs in terms of anomaly sensitivity, noise robustness, and flexibility in modelling complex activity. Chen Change Loy, Tao Xiang 0002, Shaogang Gong |
BMVC | 3 |
| 2009 | Head Pose Classification in Crowded ScenesabstractWe propose a novel technique for head pose classification in crowded public space under poor lighting and in low-resolution video images. Unlike previous approaches, we avoid the need for explicit segmentation of skin and hair regions from a head image and implicitly encode spatial information using a grid map for more robustness given lowresolution images. Specifically, a new head pose descriptor is formulated using similarity distance maps by indexing each pixel of a head image to the mean appearance templates of head images at different poses. These distance feature maps are then used to train a multi-class Support Vector Machine for pose classification. Our approach is evaluated against established techniques [3, 13, 14] using the i-LIDS underground scene dataset [9] under challenging lighting and viewing conditions. The results demonstrate that our model gives significant improvement in head pose estimation accuracy, with over 80% pose recognition rate against 32% from the best of existing models. Javier Orozco, Shaogang Gong, Tao Xiang 0002 |
BMVC | 2 |
| 2009 | A Unified Bayesian Framework for Adaptive Visual TrackingabstractTracking is regarded as one of the most fundamental tasks in computer vision. It is used in many computer vision applications in fields such as surveillance, robotic navigation and 3D reconstruction to name but a few. Despite decades of research, the goal of fully automatic tracking of arbitrary types of objects in real world conditions is still an open problem. In this paper, we take a step toward the goal of general real-world tracking, and demonstrate a unified generative model for Bayesian multifeature, adaptive target tracking, or AMFT for short (Adaptive Multiple Feature Tracker). We derive a unified generative model for multi-sensory adaptive tracking which cleanly integrates tracking and the modeling of appearance change across multiple features in the same framework. The unified multi-feature observation model ensures that if one feature is not confident, e.g., color after an object crosses into a region of shadow, it is automatically down-weighted in its contribution to the appearance model update. In this way, without pre-training of specific object models, we achieve an extensible tracker for general object types, robust to real-world problems of clutter, appearance/lighting change and target model drift. The standard modeling assumptions made by a non-adaptive generative model are illustrated by the probabilistic graphical model in Figure 1(a). The unknown target state (e.g., location, size, velocity) xt is assumed to change with time t according to some process parameterized by A. At every time t, we make some noisy observations zt of the target xt (e.g., raw image or color histograms). The target is then tracked online by computing the posterior, p(xt |z1:t) over the true target location recursively. In the case of the Kalman filter (KF), all the distributions involved are Gaussian. In the case of the particle filter (PF), all the distributions involved are represented non-parametrically by a set of samples [1]. The true target model, e.g., the appearance or color histogram to search for, is assumed to be part of the parameters H, i.e., it is known and fixed by an operator or initialized by some external process. In many cases however, the true appearance of the target H may change significantly in time, e.g., the appearance changes when a subject moves between shade and sunlight. This is the case for outdoor surveillance applications and is the motivation for this research. Adaptive trackers [2, 3, 4, 6] have been proposed to update the target appearance online in various heuristic ways. We can formalise this more general modeling assumption generatively, by the generalized dynamic Bayesian network illustrated in Figure 1(b). In contrast to Figure 1(a), the true target model which was previously included in the fixed parameters H, is now included as the the initial condition y0 of a dynamic latent variable yt , formalizing the modeling assumption that the target appearance can change over time. In addition to the target state xt , the target appearance yt will therefore be incrementally and recursively updated as part of the process of inferring the latent variables in this model p(xt ,yt |z1:t). The latent space is of course now greatly expanded, and poses a more challenging inference problem than that of Figure 1(a). In Section 2 of the paper, we detail the specific parametric form of the model and an efficient inference algorithm. We evaluate our method (AMFT) against three contemporary trackers: A standard single feature particle filter (PF), mean-shift (MS) [5] and incremental visual tracking (IVT) [4]. The PF and MS trackers are non-adaptive color-based trackers, while IVT aims for pose and illumination change robustness by performing online adaptation in a subspace appearance model. Note that the AMFT, PF and IVT trackers track object scale, but MS does not. We evaluated these methods on a series of challenging video clips exhibiting a wide variety of data and object types for tracking, including far-field indoor and outdoor pedestrians with and without carried objects, vehicle tracking, and near-field indoor face trackH 1 x2 x3 Emanuel Zelniker, Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
BMVC | 3 |
| 2009 | Associating Groups of PeopleabstractIn a crowded public space, people often walk in groups, either with people they know or strangers. Associating a group of people over space and time can assist understanding individual’s behaviours as it provides vital visual context for matching individuals within the group. Seemingly an ‘easier’ task compared with person matching given more and richer visual content, this problem is in fact very challenging because a group of people can be highly non-rigid with changing relative position of people within the group and severe self-occlusions. In this paper, for the first time, the problem of matching/associating groups of people over large space and time captured in multiple non-overlapping camera views is addressed. Specifically, a novel people group representation and a group matching algorithm are proposed. The former addresses changes in the relative positions of people in a group and the latter deals with variations in illumination and viewpoint across camera views. In addition, we demonstrate a notable enhancement on individual person matching by utilising the group description as visual context. Our methods are validated using the 2008 i-LIDS Multiple-Camera Tracking Scenario (MCTS) dataset on multiple camera views from a busy airport arrival hall. Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
BMVC | 2 |
| 2009 | Recognising action as clouds of space-time interest pointsabstractMuch of recent action recognition research is based on space-time interest points extracted from video using a Bag of Words (BOW) representation. It mainly relies on the discriminative power of individual local space-time descriptors, whilst ignoring potentially valuable information about the global spatio-temporal distribution of interest points. In this paper, we propose a novel action recognition approach which differs significantly from previous interest points based approaches in that only the global spatiotemporal distribution of the interest points are exploited. This is achieved through extracting holistic features from clouds of interest points accumulated over multiple temporal scales followed by automatic feature selection. Our approach avoids the non-trivial problems of selecting the optimal space-time descriptor, clustering algorithm for constructing a codebook, and selecting codebook size faced by previous interest points based methods. Our model is able to capture smooth motions, robust to view changes and occlusions at a low computation cost. Experiments using the KTH and WEIZMANN datasets demonstrate that our approach outperforms most existing methods. Matteo Bregonzio, Shaogang Gong, Tao Xiang 0002 |
CVPR | 2 |
| 2009 | Multi-camera activity correlation analysisabstractWe propose a novel approach for modelling correlations between activities in a busy public space captured by multiple non-overlapping and uncalibrated cameras. In our approach, each camera view is automatically decomposed into semantic regions, across which different spatio-temporal activity patterns are observed. A novel Cross Canonical Correlation Analysis (xCCA) framework is formulated to detect and quantify temporal and causal relationships between regional activities within and across camera views. The approach accomplishes three tasks: (1) estimate the spatial and temporal topology of the camera network; (2) facilitate more robust and accurate person re-identification; (3) perform global activity modelling and video temporal segmentation by linking visual evidence collected across camera views. Our approach differs from the state of the art in that it does not rely on either intra or inter camera tracking. It therefore can be applied to even the most challenging video surveillance settings featured with severe occlusions and extremely low spatial and temporal resolutions. Its effectiveness is demonstrated using 153 hours of videos from 8 cameras installed in a busy underground station. Chen Change Loy, Tao Xiang 0002, Shaogang Gong |
CVPR | 3 |
| 2009 | A Markov Clustering Topic Model for mining behaviour in videoabstractThis paper addresses the problem of fully automated mining of public space video data. A novel Markov Clustering Topic Model (MCTM) is introduced which builds on existing Dynamic Bayesian Network models (e.g. HMMs) and Bayesian topic models (e.g. Latent Dirichlet Allocation), and overcomes their drawbacks on accuracy, robustness and computational efficiency. Specifically, our model profiles complex dynamic scenes by robustly clustering visual events into activities and these activities into global behaviours, and correlates behaviours over time. A collapsed Gibbs sampler is derived for offline learning with unlabeled training data, and significantly, a new approximation to online Bayesian inference is formulated to enable dynamic scene understanding and behaviour mining in new video data online in real-time. The strength of this model is demonstrated by unsupervised learning of dynamic scene models, mining behaviours and detecting salient events in three complex and crowded public scenes. Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
ICCV | 2 |
| 2009 | Modelling activity global temporal dependencies using Time Delayed Probabilistic Graphical ModelabstractWe present a novel approach for detecting global behaviour anomalies in multiple disjoint cameras by learning time delayed dependencies between activities cross camera views. Specifically, we propose to model multi-camera activities using a Time Delayed Probabilistic Graphical Model (TD-PGM) with different nodes representing activities in different semantically decomposed regions from different camera views, and the directed links between nodes encoding causal relationships between the activities. A novel two-stage structure learning algorithm is formulated to learn globally optimised time-delayed dependencies. A new cumulative abnormality score is also introduced to replace the conventional log-likelihood score for gaining significantly more robust and reliable real-time anomaly detection. The effectiveness of the proposed approach is validated using a camera network installed at a busy underground station. Chen Change Loy, Tao Xiang 0002, Shaogang Gong |
ICCV | 3 |
| 2009 | Quantifying contextual information for object detectionabstractContext is critical for minimising ambiguity in object detection. In this work, a novel context modelling framework is proposed without the need of any prior scene segmentation or context annotation. This is achieved by exploring a new polar geometric histogram descriptor for context representation. In order to quantify context, we formulate a new context risk function and a maximum margin context (MMC) model to solve the minimization problem of the risk function. Crucially, the usefulness and goodness of contextual information is evaluated directly and explicitly through a discriminant context inference method and a context confidence function, so that only reliable contextual information that is relevant to object detection is utilised. Experiments on PASCAL VOC2005 and i-LIDS datasets demonstrate that the proposed context modelling approach improves object detection significantly and outperforms a state-of-the-art alternative context model. Wei-Shi Zheng 0001, Shaogang Gong, Tao Xiang 0002 |
ICCV | 2 |
| 2009 | Facial expression recognition based on Local Binary Patterns: A comprehensive study
Caifeng Shan, Shaogang Gong, Peter W. McOwan |
Image Vis. Comput. | 2 |
| 2009 | People detection in low-resolution video with non-stationary background
Jianguo Zhang 0001, Shaogang Gong |
Image Vis. Comput. | 2 |
| 2009 | Combining global, regional and contextual features for automatic image annotation
Tao Mei 0001, Shaogang Gong, Xian-Sheng Hua 0001 |
Pattern Recognit. | 3 |
| 2008 | Feature Selection for Gait Recognition without Subject CooperationabstractThe strength of gait, compared to other biometrics, is that it does not require cooperative subjects. Previoius gait recognition approaches were evaluated using a gallery set consisting of gait sequences of people under similar covariate conditions (i.e. clothing, surface, carrying, and view conditions). This evaluation procedure, however, implies that the gait data are collected in a cooperative manner so that the covariate conditions are known a priori. In this work, the performance of state of the art gait recognition approaches are evaluated without the assumption on cooperative subjects, i.e. the gallery set consists of a mixture of gait sequences under different unknown covariate conditions. The results show that the performance of the existing approaches drop drastically under this more realistic experimental setup. We argue that selecting the most relevant gait features that are invariant to changes in gait covariate conditions is the key to develop a gait recognition system that works without subject cooperation. To that end, we propose a novel gait recognition approach, which performs automatic feature selection on each pair gallery and probe gait sequences, and seamlessly integrates feature selection with an Adaptive Component and Discriminant Analysis (ACDA) for fast recognition. Experiments are carried out to demonstrate that the proposed approach significantly outperforms the existing techniques. 1 Khalid Bashir, Tao Xiang 0002, Shaogang Gong |
BMVC | 3 |
| 2008 | Global Behaviour Inference using Probabilistic Latent Semantic AnalysisabstractWe present a novel framework for inferring global behaviour patterns through modelling behaviour correlations in a wide-area scene and detecting any anomaly in behaviours occurring both locally and globally. Specifically, we propose a semantic scene segmentation model to decompose a wide-area scene into regions where behaviours share similar characteristic and are represented as classes of video events bearing similar features. To model behavioural correlations globally, we investigate both a probabilistic Latent Semantic Analysis (pLSA) model and a two-stage hierarchical pLSA model for global behaviour inference and anomaly detection. The proposed framework is validated by experiments using complex crowded outdoor scenes. 1 Shaogang Gong, Tao Xiang 0002 |
BMVC | 2 |
| 2008 | Video Augmentation for Improving Audio Speech Recognition under NoiseabstractFor the recognition of speech, in particular spoken digits, captured in video with poor sound due to noise, we develop a novel audio-visual fusion technique that performs significantly better than utilising either audio or video signal alone. Specifically, we present an audio-visual intermediate fusion strategy to locate speaker dependant pronounced digits in continuous video recorded with sound. A model template for each digit is represented in a single audio-visual feature space using a set of spatio-temporal visual features at multiple scales together with a set of thirteen Mel Frequency Cepstral Coefficients as audio features. Using a unified structure for both visual and audio feature selection and extraction, we solve the problem of one-to-one correspondence between the audio and visual spaces caused by differences in data sampling rates. To combine the two modalities, we adopt an intermediate fusion strategy by combining the two modalities in a probabilistic sequence matching function, permitting automatic segmentation of a continuous probe video sequence and matching with available model templates. For experiments, the CUAVE [17] database was used to compare our scheme with two alternative methods. The evaluation shows that the proposed approach outperforms the others both in recognition accuracy and robustness in coping with variations in probe sequences. 1 Samuel Pachoud, Shaogang Gong, Andrea Cavallaro |
BMVC | 2 |
| 2008 | Multi-camera Matching using Bi-Directional Cumulative Brightness Transfer FunctionsabstractThe appearance of individuals captured by multiple non-overlapping cameras varies greatly due to pose and illumination changes between camera views. In this paper we address the problem of dealing with illumination changes in order to recover matching of individuals appearing at different camera sites. This task is challenging as accurately mapping colour changes between views requires an exhaustive set of corresponding chromatic brightness values to be collected, which is very difficult in real world scenarios. We propose a Cumulative Brightness Transfer Function (CBTF) for mapping colour between cameras located at different physical sites, which makes better use of the available colour information from a very sparse training set. In addition we develop a bi-directional mapping approach to obtain a more accurate similarity measure between a pair of candidate objects. We evaluate the proposed method using challenging datasets obtained from real world distributed CCTV camera networks. The results demonstrate that our bi-directional CBTF method significantly outperforms existing techniques. 1 Bryan James Prosser, Shaogang Gong, Tao Xiang 0002 |
BMVC | 2 |
| 2008 | Exploiting Periodicity in Recurrent ScenesabstractThere is considerable interest in techniques capable of identifying anomalies and unusual events in busy outdoor scenes, e.g. road junctions. Many approaches achieve this by exploiting deviations in spatial appearance from some expected norm accumulated by a model over time. In this work we show that much can be gained from explicitly modelling temporal aspects of scene activity in detail. We characterize a scene by identifying the fundamental period of change on a spatial block-by-block basis by estimating autocovariance of self-similarity. As our model, we introduce a spatio-temporal grid of histograms built corresponding to some chosen feature. This model is then used to identify objects found in unexpected spatial and temporal locations in subsequent test data. Employing a Phase-Locked Loop technique, we describe a method of ensuring that the spatio-temporal model maintains synchronization with learned scene activity in spite of short-term breakdown in the reliability of acquired data, and long-term change of the mean fundamental period. Results indicate our model to be capable of discrimination between behavioural aspects of cars at a typical road junction sufficiently well to provide useful warnings of adverse activity in real time. 1 David Mark Russell, Shaogang Gong |
BMVC | 2 |
| 2008 | Coherent image annotation by learning semantic distanceabstractConventional approaches to automatic image annotation usually suffer from two problems: (1) They cannot guarantee a good semantic coherence of the annotated words for each image, as they treat each word independently without considering the inherent semantic coherence among the words; (2) They heavily rely on visual similarity for judging semantic similarity. To address the above issues, we propose a novel approach to image annotation which simultaneously learns a semantic distance by capturing the prior annotation knowledge and propagates the annotation of an image as a whole entity. Specifically, a semantic distance function (SDF) is learned for each semantic cluster to measure the semantic similarity based on relative comparison relations of prior annotations. To annotate a new image, the training images in each cluster are ranked according to their SDF values with respect to this image and their corresponding annotations are then propagated to this image as a whole entity to ensure semantic coherence. We evaluate the innovative SDF-based approach on Corel images compared with Support Vector Machine-based approach. The experiments show that SDF-based approach outperforms in terms of semantic coherence, especially when each training image is associated with multiple words. Tao Mei 0001, Xian-Sheng Hua 0001, Shaogang Gong, Shipeng Li 0001 |
CVPR | 4 |
| 2008 | Macro-cuboïd based probabilistic matching for lip-reading digitsabstractIn this paper, we present a spatio-temporal feature representation and a probabilistic matching function to recognise lip movements from pronounced digits. Our model (1) automatically selects spatio-temporal features extracted from 10 digit model templates and (2) matches them with probe video sequences. Spatio-temporal features embed lip movements from pronouncing digits and contain more discriminative information than spatial features alone. A model template for each digit is represented by a set of spatio-temporal features at multiple scales. A probabilistic sequence matching function automatically segments a probe video sequence and matches the most likely sequence of digits recognised in the probe sequence. We demonstrate the proposed approach using the CUAVE database and compare our representational scheme with three alternative methods, based on optical flow, intensity gradient and block matching, respectively. The evaluation shows that the proposed approach outperforms the others in recognition accuracy and is robust in coping with variations in probe sequences. Samuel Pachoud, Shaogang Gong, Andrea Cavallaro |
CVPR | 2 |
| 2008 | Scene Segmentation for Behaviour Correlation
Shaogang Gong, Tao Xiang 0002 |
ECCV (4) | 2 |
| 2008 | Multi-layered Decomposition of Recurrent Scenes
David Mark Russell, Shaogang Gong |
ECCV (3) | 2 |
| 2008 | Feature selection on Gait Energy Image for human identificationabstractIn this paper we address the problem of selecting the most relevant features for human identification by gait. Although gait as a behavioral biometric is concerned with how people walk rather than how people look, most existing gait recognition approaches employ both shape and dynamics information for recognition. This is because shape, as a static appearance feature also contains useful information for identification. However, the inclusion of shape information in the gait features can also introduce variations that will hinder the recognition performance. To address this problem, we develop both supervised and unsupervised feature selection methods to extract the most relevant and informative features from Gait Energy Image (GEI) for human identification. Extensive experiments are carried out which indicate that our feature selection methods significantly improve the performance of gait recognition. Khalid Bashir, Tao Xiang 0002, Shaogang Gong |
ICASSP | 3 |
| 2008 | Incremental and adaptive abnormal behaviour detection
Tao Xiang 0002, Shaogang Gong |
Comput. Vis. Image Underst. | 2 |
| 2008 | Optimising dynamic graphical models for video content analysis
Tao Xiang 0002, Shaogang Gong |
Comput. Vis. Image Underst. | 2 |
| 2008 | Fusing gait and face cues for human gender recognition
Caifeng Shan, Shaogang Gong, Peter W. McOwan |
Neurocomputing | 2 |
| 2008 | Video Behavior Profiling for Anomaly DetectionabstractThis paper aims to address the problem of modelling video behaviour captured in surveillancevideos for the applications of online normal behaviour recognition and anomaly detection. A novelframework is developed for automatic behaviour profiling and online anomaly sampling/detectionwithout any manual labelling of the training dataset. The framework consists of the followingkey components: (1) A compact and effective behaviour representation method is developed basedon discrete scene event detection. The similarity between behaviour patterns are measured basedon modelling each pattern using a Dynamic Bayesian Network (DBN). (2) Natural grouping ofbehaviour patterns is discovered through a novel spectral clustering algorithm with unsupervisedmodel selection and feature selection on the eigenvectors of a normalised affinity matrix. (3) Acomposite generative behaviour model is constructed which is capable of generalising from asmall training set to accommodate variations in unseen normal behaviour patterns. (4) A run-timeaccumulative anomaly measure is introduced to detect abnormal behaviour while normal behaviourpatterns are recognised when sufficient visual evidence has become available based on an onlineLikelihood Ratio Test (LRT) method. This ensures robust and reliable anomaly detection and normalbehaviour recognition at the shortest possible time. The effectiveness and robustness of our approachis demonstrated through experiments using noisy and sparse datasets collected from both indoorand outdoor surveillance scenarios. In particular, it is shown that a behaviour model trained usingan unlabelled dataset is superior to those trained using the same but labelled dataset in detectinganomaly from an unseen video. The experiments also suggest that our online LRT based behaviourrecognition approach is advantageous over the commonly used Maximum Likelihood (ML) methodin differentiating ambiguities among different behaviour classes observed online. Tao Xiang 0002, Shaogang Gong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Spectral clustering with eigenvector selection
Tao Xiang 0002, Shaogang Gong |
Pattern Recognit. | 2 |
| 2008 | Activity based surveillance video content modelling
Tao Xiang 0002, Shaogang Gong |
Pattern Recognit. | 2 |
| 2008 | Generalized Face Super-ResolutionabstractExisting learning-based face super-resolution (hallucination) techniques generate high-resolution images of a single facial modality (i.e., at a fixed expression, pose and illumination) given one or set of low-resolution face images as probe. Here, we present a generalized approach based on a hierarchical tensor (multilinear) space representation for hallucinating high-resolution face images across multiple modalities, achieving generalization to variations in expression and pose. In particular, we formulate a unified tensor which can be reduced to two parts: a global image-based tensor for modeling the mappings among different facial modalities, and a local patch-based multiresolution tensor for incorporating high-resolution image details. For realistic hallucination of unregistered low-resolution faces contained in raw images, we develop an automatic face alignment algorithm capable of pixel-wise alignment by iteratively warping the probing face to its projection in the space of training face images. Our experiments show not only performance superiority over existing benchmark face super-resolution techniques on single modal face hallucination, but also novelty of our approach in coping with multimodal hallucination and its robustness in automatic alignment under practical imaging conditions. Kui Jia, Shaogang Gong |
IEEE Trans. Image Process. | 2 |
| 2007 | Learning gender from human gaits and facesabstractComputer vision based gender classification is an important component in visual surveillance systems. In this paper, we investigate gender classification from human gaits in image sequences, a relatively understudied problem. Moreover, we propose to fuse gait and face for improved gender discrimination. We exploit Canonical Correlation Analysis (CCA), a powerful tool that is well suited for relating two sets of measurements, to fuse the two modalities at the feature level. Experiments demonstrate that our multimodal gender recognition system achieves the superior recognition performance of 97.2% in large datasets. Caifeng Shan, Shaogang Gong, Peter W. McOwan |
AVSS | 2 |
| 2007 | Segmenting Highly Textured Nonstationary BackgroundabstractDetection of unusual objects amongst a highly textured background is a difficult problem, especially when the texture is manifest in the temporal dimension as well. Outdoor scenes involving waving trees or moving water are examples of such a scenario, but are nevertheless frequently encountered in real world vision applications. By defining a simple but rotationally sensitive Local Binary Pattern (LBP) operator and applying it in a probabilistic sense we present a compact but useful feature for tackling moving textures. But as we demonstrate, this alone is not sufficient for good segmentation in difficult circumstances. Cooccurrence of different features in a pixel’s local neighbourhood provides a powerful mechanism for boosting the reliability of the foreground/background decision task. By using the conditional probabilities yielded by pairwise cooccurrence of 4-connected pixels, and casting the problem as one of Combinatorial Optimization, our results show that useful segmentation is possible from challenging dynamic backgrounds. 1 David Mark Russell, Shaogang Gong |
BMVC | 2 |
| 2007 | Beyond Facial Expressions: Learning Human Emotion from Body GesturesabstractVision-based human affect analysis is an interesting and challenging problem, impacting important applications in many areas. In this paper, beyond facial expressions, we investigate affective body gesture analysis in video sequences, a relatively understudied problem. Spatial-temporal features are exploited for modeling of body gestures. Moreover, we present to fuse facial expression and body gesture at the feature level using Canonical Correlation Analysis (CCA). By establishing the relationship between the two modalities, CCA derives a semantic “affect ” space. Experimental results demonstrate the effectiveness of our approaches. 1 Caifeng Shan, Shaogang Gong, Peter W. McOwan |
BMVC | 2 |
| 2007 | Capturing Correlations Among Facial Parts for Facial Expression AnalysisabstractCapturing and analyzing the correlations among facial parts are important for interpreting facial behaviors precisely. In this paper, we exploit Canonical Correlation Analysis (CCA) to model the correlations of facial parts for facial expression analysis. We propose a Matrix-based Canonical Correlation Analysis (MCCA) for better correlation analysis on 2D image or matrix data in general. Extensive experiments have shown that compared to the traditional CCA, MCCA models more accurately correlations among image data with more compact representation using much fewer canonical factors. 1 Caifeng Shan, Shaogang Gong, Peter W. McOwan |
BMVC | 2 |
| 2007 | Conditional Random Field for Natural Scene CategorizationabstractConditional random field (CRF) has been widely used for sequence labeling and segmentation. However, CRF does not offer a straightforward approach to classify whole sequences. On the other hand, hidden conditional random field (HCRF) has been proposed for whole sequences classification by viewing the segment labels as hidden variables. But the objective function of HCRF is non-convex because of its hidden variable structure. In this paper, we propose a classification oriented CRF (COCRF) adapted from HCRF for natural scene categorization by taking an image as an ordered set of local patches. Our approach firstly assigns a topic label to each segment on the training data by the probabilistic latent semantic analysis (PLSA) and train a COCRF model given these topic labels. PLSA provides a higher level of semantic grouping of image patches by considering their co-occurrence relationships while COCRF provides a probabilistic model for the spatial layout structure of image patches. The combination of PLSA and COCRF can not only classify but also interpret scene categories. We tested our approach on two well-known datasets and demonstrated its advantage over existing approaches. 1 Shaogang Gong |
BMVC | 2 |
| 2007 | Translating topics to words for image annotationabstractOne of the classic techniques for image annotation is the language translation model. It views an image as a document, i.e., a set of visual words which are obtained by vector quatitizing the image regions generated by unsupervised image segmentation. Annotating images are achieved by translating visual words to textual words, just like translating a document in English to a document in French. In this paper, we also view an image as a document, but we view the annotation processes as two consecutive processes, i.e., document summarization and translation. In the document summarization process, an image document is firstly summarized into its own visual language, which we called visual topics. The translation process translates these visual topics to textual words. Compared to the original translation model, our visual topics learned by the probabilistic latent semantic analysis (PLSA) approach provide an intermediate abstract level of visual description. We show improved annotation performance on the Corel image dataset. Shaogang Gong |
CIKM | 2 |
| 2007 | Optimizing Distribution-based Matching by Random SubsamplingabstractWe boost the efficiency and robustness of distribution-based matching by random subsampling which results in the minimum number of samples required to achieve a specified probability that a candidate sampling distribution is a good approximation to the model distribution. The improvement is demonstrated with applications to object detection, mean-shift tracking using color distributions and tracking with improved robustness for low-resolution video sequences. The problem of minimizing the number of samples required for robust distribution matching is formulated as a constrained optimization problem with the specified probability as the objective function. We show that surprisingly mean-shift tracking using our method requires very few samples. Our experiments demonstrate that robust tracking can be achieved with even as few as 5 random samples from the distribution of the target candidate. This leads to a considerably reduced computational complexity that is also independent of object size. We show that random subsampling speeds up tracking by two orders of magnitude for typical object sizes. Alex Po Leung, Shaogang Gong |
CVPR | 2 |
| 2007 | Visual inference of human emotion and behaviourabstractWe address the problem of automatic interpretation ofnon-exaggerated human facial and body behaviours captured in video. We illustrate our approach by three examples. (1) We introduce Canonical Correlation Analysis (CCA) and Matrix Canonical Correlation Analysis (MCCA) for capturing and analyzing spatial correlations among non-adjacent facial parts for facial behaviour analysis. (2) We extend Canonical Correlation Analysis to multimodality correlation for bebaviour inference using both facial and body gestures. (3) We model temporal correlation among human movement patterns in a wider space using a mixture of Multi-Observation Hidden Markov Model for human behaviour profiling and behavioural anomaly detection. Shaogang Gong, Caifeng Shan, Tao Xiang 0002 |
ICMI | 1 |
| 2006 | Coupling Face Registration and Super-ResolutionabstractExisting approaches to learning-based face image super-resolution require low-resolution testing inputs manually registered to pre-aligned highresolution training models [9, 12, 13, 5]. This restricts automatic applications to live images and video. In this paper, we propose a multi-resolution patch tensor based model to automatically super-resolve and register low-resolution testing face images. Face candidates are triggered first by a face detector giving the subwindows with their coarse initial positions and scales in a large image frame. This initialises a combined registration and super-resolution process. Rather than manually aligning each coarsely detected face subwindow to some predefined template, based on its position and scale, we scan all the potential face subwindows across different positions and scales, and obtain registration and super-resolution in a simultaneous process. The superresolution result which is optimally correlated to its original low-resolution face subwindow is also guaranteed to be the best super-resolved reconstruction. We verify our approach by experimenting on MIT+CMU face detection dataset, the promising results demonstrate the robustness of our approach on learning-based face super-resolution on real images. 1 Kui Jia, Shaogang Gong, Alex Po Leung |
BMVC | 2 |
| 2006 | Mean-Shift Tracking with Random SamplingabstractIn this work, boosting the efficiency of Mean-Shift Tracking using random sampling is proposed. We obtained the surprising result that mean-shift tracking requires only very few samples. Our experiments demonstrate that robust tracking can be achieved with as few as even 5 random samples from the image of the object. As the computational complexity is considerably reduced and becomes independent of object size, the processor can be used to handle other processing tasks while tracking. It is demonstrated that random sampling significantly reduces the processing time by two orders of magnitude for typical object sizes. Additionally, with random sampling, we propose a new optimal on-line feature selection algorithm for object tracking which maximizes a similarity measure for the weights of the RGB channels. It selects the weights of the RGB channels which discriminate the object and the background the most using Steepest Descent. Moreover, the spatial distribution of pixels representing the object is estimated for spatial weighting. Arbitrary spatial weighting is incorporated into Mean-Shift Tracking to represent objects with arbitrary or changing shapes by picking up non-uniform random samples. Experimental results demonstrate that our tracker with online feature selection and arbitrary spatial weighting outperforms the original mean-shift tracker with improved computational efficiency and tracking accuracy. 1 Alex Po Leung, Shaogang Gong |
BMVC | 2 |
| 2006 | Sparse Multiscale Local Binary PatternsabstractIn a Local Binary Pattern (LBP) representation, circular point features are taken in their entirety as predicates and restricted to uniform patterns with limited scales of small numbers of features in order to avoid large bin complexity. Such a design cannot fully exploit the discriminative capacity of the features available. To address the problem, this paper proposes (1) a pairwise-coupled reformulation of LBP-type classification which involves selecting single-point features for each pair of classes across multiple scales to form compact, contextually-relevant multiscale predicates known as Multiscale Selected Local Binary Features (MSLBF), and (2) a novel binary feature selection procedure, known as Binary Histogram Intersection Minimisation (BHIM) designed to choose features with minimal redundancy. Experiments show the advantages of MSLBF over traditional LBP representation and of BHIM over feature selection schemes such as AdaBoost. 1 Yogesh Raja, Shaogang Gong |
BMVC | 2 |
| 2006 | Minimum Cuts of A Time-Varying BackgroundabstractMotivated by the demand for an effective background model, robust to non-stationary environmental changes in outdoor scenes, we present a technique using Combinatorial Optimization to extract near-optimal background estimates from blocks of temporally localized frames. Using an existing graph cut technique in conjunction with subspace analysis, we demonstrate a novel background model exhibiting results superior to those achievable with the latter technique alone, and especially suitable for background modelling in outdoor situations where variable lighting conditions prevail. 1 David Mark Russell, Shaogang Gong |
BMVC | 2 |
| 2006 | Dynamic Facial Expression Recognition Using A Bayesian Temporal Manifold ModelabstractIn this paper, we propose a novel Bayesian approach to modelling tem-poral transitions of facial expressions represented in a manifold, with the aim of dynamical facial expression recognition in image sequences. A gener-alised expression manifold is derived by embedding image data into a low dimensional subspace using Supervised Locality Preserving Projections. A Bayesian temporal model is formulated to capture the dynamic facial ex-pression transition in the manifold. Our experimental results demonstrate the advantages gained from exploiting explicitly temporal information in ex-pression image sequences resulting in both superior recognition rates and improved robustness against static frame-based recognition methods. 1 Caifeng Shan, Shaogang Gong, Peter W. McOwan |
BMVC | 2 |
| 2006 | Optimal Dynamic Graphs for Video Content AnalysisabstractThis study addresses the problem of learning the optimal structure of a dynamic graphical model for video content analysis given sparse data. We propose a Completed Likelihood AIC (CL-AIC) scoring function that differs from existing ones by optimising explicitly both the explanation and prediction capabilities of a model simultaneously. We demonstrate that CL-AIC is superior to existing scoring functions including BIC, AIC and ICL in building dynamic graph models for video content analysis. 1 Tao Xiang 0002, Shaogang Gong |
BMVC | 2 |
| 2006 | Multi-Resolution Patch Tensor for Facial Expression HallucinationabstractIn this paper, we propose a sequential approach to hallucinate/ synthesize high-resolution images of multiple facial expressions. We propose an idea of multi-resolution tensor for super-resolution, and decompose facial expression images into small local patches. We build a multi-resolution patch tensor across different facial expressions. By unifying the identity parameters and learning the subspace mappings across different resolutions and expressions, we simplify the facial expression hallucination as a problem of parameter recovery in a patch tensor space. We further add a high-frequency component residue using nonparametric patch learning from high-resolution training data. We integrate the sequential statistical modelling into a Bayesian framework, so that given any low-resolution facial image of a single expression, we are able to synthesize multiple facial expression images in high-resolution. We show promising experimental results from both facial expression database and live video sequences. Kui Jia, Shaogang Gong |
CVPR (1) | 2 |
| 2006 | Beyond Tracking: Modelling Activity and Understanding Behaviour
Tao Xiang 0002, Shaogang Gong |
Int. J. Comput. Vis. | 2 |
| 2006 | Model Selection for Unsupervised Learning of Visual Context
Tao Xiang 0002, Shaogang Gong |
Int. J. Comput. Vis. | 2 |
| 2006 | Hallucinating multiple occluded face images of different resolutions
Kui Jia, Shaogang Gong |
Pattern Recognit. Lett. | 2 |
| 2005 | Detecting and quantifying unusual interactions by correlating salient motionabstractA significant problem in scene interpretation is efficient bottom-up extraction and representation of salient features. In this paper, we address the problem of correlating salient motion at a spatio-temporal level and also across spatially separated regions since it is in the interactions that more sophisticated scene interpretation can be found. We show that it is possible to spatio-temporally locate and detect salient motion events and interactions in two contrasting scenarios using the same hierarchical co-occurrence framework. Thus generating a concise description of a dynamic scene from the sequence data alone. Results show it is possible to reduce a highly populated multi-dimensional co-occurrence matrix representing correlations between salient motion regions, to a one dimensional vector with clearly separable unusual activity. The results also show that the method inherently provides a quantifiable measure of the saliency of an interaction through its frequency of occurrence. Hayley Shi-Wen Hung, Shaogang Gong |
AVSS | 2 |
| 2005 | Multi-modal face image super-resolutions in tensor spaceabstractFace images of non-frontal views under poor illumination with low resolution reduce dramatically face recognition accuracy. To overcome these problems, super-resolution techniques can be exploited. In this paper, we present a Bayesian framework to perform multi-modal (such as variations in viewpoint and illumination) face image super-resolutions in tensor space. Given a single modal low-resolution face image, we benefit from the multiple factor interactions of training tensor, and super-resolve its high-resolution reconstructions across different modalities. Instead of performing pixel-domain super-resolutions, we reconstruct the high-resolution face images by computing a maximum likelihood identity parameter vector in high-resolution tensor space. Experiments show promising results of multi-view and multi-illumination face image super-resolutions respectively. Kui Jia, Shaogang Gong |
AVSS | 2 |
| 2005 | Face super-resolution using multiple occluded images of different resolutionsabstractIn this paper, we present a novel learning-based algorithm to super-resolve multiple partially occluded low-resolution face images. By integrating hierarchical patch-wise alignment and inter-frame constraints into a Bayesian framework, we can probabilistically align multiple input images at different resolutions and recursively infer the high-resolution face image. We address the problem of fusing partial imagery information through multiple frames and discuss the new algorithm's effectiveness when encountering occluded low-re solution face images. We show promising results compared to that of existing face hallucination methods. Kui Jia, Shaogang Gong |
AVSS | 2 |
| 2005 | A highly efficient block-based dynamic background modelabstractA block-based dynamic background modelling technique featuring incremental update is presented, with the aim of increasing the sensitivity of background detection by considering the local connectivity of pixels. An accumulative model is maintained on-line by a heuristic algorithm to build up a statistical representation of a dynamic scene from a fixed camera view for foreground-background segmentation. The model is compared with a block-based technique employing cooccurrence [M. Seki et al., June 2003], and shown to be favourable in terms of both performance and conceptual simplicity. Using predominantly simple integer operations, the algorithm is highly amenable to direct implementation in hardware. David Mark Russell, Shaogang Gong |
AVSS | 2 |
| 2005 | Recognizing facial expressions at low resolutionabstractThis paper focuses on recognizing facial expressions at low resolution. We introduce local binary patterns (LBP) as novel low-computation discriminative features for low-resolution facial expression recognition. Compared to Gabor wavelets, LBP features can be derived rapidly in a single scan of raw images, whilst still retaining enough facial information in a compact representation. Support vector machine (SVM) is adopted to classify facial expressions. Extensive experiments on the Cohn-Kanade database demonstrate that the LBP features are effective and efficient for facial expression recognition, and crucially perform robustly and stably over a useful range of low resolutions. Our method yields promising performance when processing compressed low-resolution video sequences from the PETS 2003 dataset. Caifeng Shan, Shaogang Gong, Peter W. McOwan |
AVSS | 2 |
| 2005 | Relevance learning for spectral clustering with applications on image segmentation and video behaviour profilingabstractWe aim to tackle the problem of unsupervised visual learning. A novel relevance learning algorithm is proposed for data clustering using eigenvectors of a data affinity matrix. We show that it is critical to select the relevant eigenvectors for both estimating the optimal number of clusters and performing data clustering especially given noisy and sparse data. The effectiveness of our algorithm is demonstrated on solving two challenging visual data clustering problems: image segmentation and video behaviour profiling. Tao Xiang 0002, Shaogang Gong |
AVSS | 2 |
| 2005 | Multi-Modal Face Image Super-Resolutions in Tensor SpaceabstractFace images of non-frontal views under poor illumination with low resolution reduce dramatically face recognition accuracy. To overcome these problems, super-resolution techniques can be exploited. In this paper, we present a Bayesian framework to perform multi-modal (such as variations in viewpoint and illumination) face image super-resolutions in tensor space. Given a single modal low-resolution face image, we benefit from the multiple factor interactions of training tensor, and super-resolve its high-resolution reconstructions across different modalities. Instead of performing pixel-domain super-resolutions, we reconstruct the high-resolution face images by computing a maximum likelihood identity parameter vector in high-resolution tensor space. Experiments show promising results of multi-view and multiillumination face image super-resolutions respectively. 1 Kui Jia, Shaogang Gong |
BMVC | 2 |
| 2005 | An Optimization Framework for Real-Time Appearance-Based Tracking under Weak PerspectiveabstractIn this work, we present a framework for tracking objects in changing views by finding the subwindow most likely to be the object using Haar-like features selected by AdaBoost as the representation. Probabilistic AdaBoost [14] is used to derive the objective function. In addition, the projective warping of 2D features is used to track 3D objects in non-frontal views in real time. Transformed 2D features can approximate relatively flat object structures such as the two eyes in a face. In this paper, it is shown that, under weak perspective projection, the projective warping of a rectangle feature can be approximated by a similarity transform with an additional free parameter. Since features in non-frontal views are computed on-the-fly by projective transforms under weak perspective projection, our framework requires only frontal-view training samples to track objects in multiple views. 1 Alex Po Leung, Shaogang Gong |
BMVC | 2 |
| 2005 | Conditional Mutual Infomation Based Boosting for Facial Expression RecognitionabstractThis paper proposes a novel approach for facial expression recognition by boosting Local Binary Patterns (LBP) based classifiers. L ow-cost LBP features are introduced to effectively describle local fea tures of face images. A novel learning procedure, Conditional Mutual Infomation based Boosting (CMIB), is proposed. CMIB learns a sequence of weak classifie rs that maximize their mutual information about a candidate class, conditional to the response of any weak classifier already selected; a strong cl assifier is constructed by combining the learned weak classifiers using the Naive-Bayes. Extensive experiments on the Cohn-Kanade database illustrated that LBP features are effective for expression analysis, and CMIB enables much faster training than AdaBoost, and yields a classifier of improved c lassification performance. Caifeng Shan, Shaogang Gong, Peter W. McOwan |
BMVC | 2 |
| 2005 | Online Video Behaviour Abnormality Detection Using Reliability MeasureabstractAn approach is proposed for robust online behaviour recognition and abnormality detection based on discovering natural grouping of bebaviour patterns through unsupervised learning and a time accumulative reliability measure. A novel behaviour learning model and a run-time accumulative reliability measure are introduced to determine both the natural groupings of possible normal behaviour classes without manual labelling and when sufficient visual evidence has become available for differentiating ambiguities among different behaviour classes observed online. This ensures behaviour recognition at the shortest possible time and robust abnormality detection. 1 Tao Xiang 0002, Shaogang Gong |
BMVC | 2 |
| 2005 | 2D Statistical Models of Facial Expressions for Realistic 3D Avatar AnimationabstractWe address the issue of modelling facial expressions for realistic 3D avatar animation. We introduce a hierarchical decomposition of a human face into different components and model them according to their intrinsic functionalities. The parametrisation of the expressions is achieved in a two-level framework. First level accounts for the low level component facial actions and is represented by hierarchical latent variable models. The second level models the final expressions as a combination of subcomponent information extracted from the lower level using combinatorial logic. Finally we produce continuous animation curves that are used to animate 3D avatar in a morph-based fashion. Our approach is entirely based on 2D information extracted from the input source. Lukasz Zalewski, Shaogang Gong |
CVPR (2) | 2 |
| 2005 | Multi-Modal Tensor Face for Simultaneous Super-Resolution and RecognitionabstractFace images of non-frontal views under poor illumination resolution reduce dramatically face recognition accuracy. This is evident most compellingly by the very low recognition rate of all existing face recognition systems when applied to live CCTV camera input. In this paper, we present a Bayesian framework to perform multimodal (such as variations in viewpoint and illumination) face image super-resolution for recognition in tensor space. Given a single modal low-resolution face image, we benefit from the multiple factor interactions of training sensor and super-resolve its high-resolution reconstructions across different modalities for face recognition. Instead of performing pixel-domain super-resolution and recognition independently as two separate sequential processes, we integrate the tasks of super-resolution and recognition by directly computing a maximum likelihood identity parameter vector in high-resolution tensor space for recognition. We show results from multi-modal super-resolution and face recognition experiments across different imaging modalities, using low-resolution images as testing inputs and demonstrate improved recognition rates over standard tensorface and eigenface representations. Kui Jia, Shaogang Gong |
ICCV | 2 |
| 2005 | Visual Learning Given Sparse Data of Unknown ComplexityabstractThis study addresses the problem of unsupervised visual learning. It examines existing popular model order selection criteria before proposes two novel criteria for improving visual learning given sparse data and without any knowledge about model complexity. In particular, a rectified Bayesian information criterion (BICr) and a completed likelihood Akaike's information criterion (CL-AIC) are formulated to estimate the optimal model order (complexity) for learning the dynamic structure of a visual scene. Both criteria are designed to overcome poor model selection by existing popular criteria when the data sample size varies from very small to large. Extensive experiments on learning a dynamic scene structure are carried out to demonstrate the effectiveness of BICr and CL-AIC, compared to that of BIC (Schwarz, 1978), AIC (Akaike, 1973), ICL (Biernacki, 2000) and a MML (Figueiredo and Jain, 2002) based criterion. Tao Xiang 0002, Shaogang Gong |
ICCV | 2 |
| 2005 | Video Behaviour Profiling and Abnormality Detection without Manual LabellingabstractA novel framework is developed for automatic behaviour profiling and abnormality sampling/detection without any manual labelling of the training dataset. Natural grouping of behaviour patterns is discovered through unsupervised model selection and feature selection on the eigenvectors of a normalised affinity matrix. Our experiments demonstrate that a behaviour model trained using an unlabelled dataset is superior to those trained using the same but labelled dataset in detecting abnormality from an unseen video. Tao Xiang 0002, Shaogang Gong |
ICCV | 2 |
| 2005 | Robust facial expression recognition using local binary patternsabstractA novel low-computation discriminative feature space is introduced for facial expression recognition capable of robust performance over a rang of image resolutions. Our approach is based on the simple local binary patterns (LBP) for representing salient micro-patterns of face images. Compared to Gabor wavelets, the LBP features can be extracted faster in a single scan through the raw image and lie in a lower dimensional space, whilst still retaining facial information efficiently. Template matching with weighted Chi square statistic and support vector machine are adopted to classify facial expressions. Extensive experiments on the Cohn-Kanade Database illustrate that the LBP features are effective and efficient for facial expression discrimination. Additionally, experiments on face images with different resolutions show that the LBP features are robust to low-resolution images, which is critical in real-world applications where only low-resolution video input is available. Caifeng Shan, Shaogang Gong, Peter W. McOwan |
ICIP (2) | 2 |
| 2004 | Wavelet-based holistic sequence descriptor for generating video summariesabstractWe propose a video representation, the Video Scene Trajectory, computed from localised temporal-change in order to capture the holistic action content in a sequence. Such a representation is critical for action based genres such as surveillance. We show that an analysis of the trajectory shape is able to produce a temporal segmentation. Furthermore, we build an action based video summary using the Top Discriminative Active Pixels in the segments. Experiments are presented on real-world outdoor surveillance scenes. Andrew Graves, Shaogang Gong |
BMVC | 2 |
| 2004 | Quantifying Temporal SaliencyabstractA significant problem in automatic scene interpretation is the ability to per-form contextually meaningful segmentation of both static and moving images using a bottom-up approach. We examine and propose an extension to Kadir and Brady’s Scale Saliency Algorithm for quantifying temporal saliency and performing automatic spatial and temporal scale selection. 1 Hayley Shi-Wen Hung, Shaogang Gong |
BMVC | 2 |
| 2004 | Activity Based Video Content Trajectory Representation and SegmentationabstractA novel approach is developed to segment continuous CCTV recordings according to the activities captured in the scene. This approach differs from previous approaches which are mostly based on shot change detection and shot grouping. Video content is represented by constructing a cumulative multi-event histogram over time. An on-line segmentation algorithm is then proposed to detect breakpoints in the video content, which is more robust to noise and computationally much more efficient compared to existing on-line segmentation algorithms. Its performance is compared to that of two off-line algorithms on surveillance videos monitoring an aircraft ramp area. Our experiments demonstrate that this activity based video content representation is superior to a colour histogram based representation for surveillance video segmentation. 1 Tao Xiang 0002, Shaogang Gong |
BMVC | 2 |
| 2004 | Support vector machine based multi-view face detection and recognition
Yongmin Li 0001, Shaogang Gong, Jamie Sherrah, Heather M. Liddell |
Image Vis. Comput. | 2 |
| 2003 | Discovering Bayesian Causality among Visual Events in a Complex Outdoor SceneabstractModelling events is one of the key problems in dynamic scene understanding when salient and autonomous visual changes occurring in a scene need to be characterised as a set of different object temporal events. We propose an approach to understand complex outdoor scenarios which is based on modelling temporally correlated events using dynamic Bayesian networks (DBNs). A partially coupled hidden Markov model (PCHMM) is exploited whose topology is determined automatically using the Bayesian information criterion (BIC). Causality discovery and events modelling are also tackled using a multi-observation hidden Markov model (MOHMM). Tao Xiang 0002, Shaogang Gong |
AVSS | 2 |
| 2003 | Spotting Scene Change for Indexing Surveillance VideoabstractWe address the issue of forming a pre-attentive mechanism that can be used to analyse surveillance sequences. We address the problem of spotting scene change by performing temporal segmentation on long video sequences with little colour information and observed content. This is typical in surveillance sequences. Our approach: (1) employs sustained temporal change computed for local neighbourhoods in the image frames; (2) defines a frame activity similarity metric that accounts for local spatial and temporal displacement of change; and (3) monitors the similarity over a wide period to detect changes in emphasis that are then identified as scene breaks. 1 Andrew Graves, Shaogang Gong |
BMVC | 2 |
| 2003 | Outdoor Activity Recognition using Multi-Linked Temporal ProcessesabstractWe develop Dynamically Multi-Linked Hidden Markov Models (DML-HMMs) for interpreting group activities involving multiple objects captured in an outdoor scene. The models are based on the discovery of salient dynamic interlinks among multiple different object events. A layered hierarchical DML-HMM is built using Schwarz’s Bayesian Information Criterion (BIC) based factorisation resulting in its topology being intrinsically determined by the underlying causality and temporal order among different object events. Our experiments demonstrate that the performance of a DML-HMM on modelling group activities in a noisy outdoor scene is superior compared to that of a Coupled Hidden Markov Model (CHMM). Tao Xiang 0002, Shaogang Gong, Dennis Parkinson |
BMVC | 2 |
| 2003 | Recognition of Group Activities using Dynamic Probabilistic NetworksabstractDynamic Probabilistic Networks (DPNs) are exploited for modeling the temporal relationships among a set of different object temporal events in the scene for a coherent and robust scene-level behaviour interpretation. In particular, we develop a Dynamically Multi-Linked Hidden Markov Model (DML-HMM) to interpret group activities involving multiple objects captured in an outdoor scene. The model is based on the discovery of salient dynamic interlinks among multiple temporal events using DPNs. Object temporal events are detected and labeled using Gaussian Mixture Models with automatic model order selection. A DML-HMM is built using Schwarz's Bayesian Information Criterion based factorisation resulting in its topology being intrinsically determined by the underlying causality and temporal order among different object events. Our experiments demonstrate that its performance on modelling group activities in a noisy outdoor scene is superior compared to that of a Multi-Observation Hidden Markov Model (MOHMM), a Parallel Hidden Markov Model (PaHMM) and a Coupled Hidden Markov Model (CHMM). Shaogang Gong, Tao Xiang 0002 |
ICCV | 1 |
| 2003 | Constructing Facial Identity Surfaces for Recognition
Yongmin Li 0001, Shaogang Gong, Heather M. Liddell |
Int. J. Comput. Vis. | 2 |
| 2003 | Learning pixel-wise signal energy for understanding semantics
Jeffrey Ng Sing Kwong, Shaogang Gong |
Image Vis. Comput. | 2 |
| 2003 | Recognising trajectories of facial identities using kernel discriminant analysis
Yongmin Li 0001, Shaogang Gong, Heather M. Liddell |
Image Vis. Comput. | 2 |
| 2002 | Autonomous Visual Events Detection and Classification without Explicit Object-Centred Segmentation and TrackingabstractModelling events is one of the key problems in dynamic scene analysis when salient and autonomous visual changes occuring in a scene need to be characterised effectively as meaningful events. We propose a new approach for modelling such temporal events based on the local intensity temporal history of pixels. The method provides a computationally very effective temporal measure for detecting autonomous events. Events are represented and detected first at the pixel level and then at a blob level (grouped pixels) autonomously. The Expectation-Maximisation (EM) algorithm is employed to cluster events with automatic model order selection using modified Minimum Description Length (MDL). Experiments are presented to demonstrate that meaningful clusters of blob-level events can be formed without object segmentation and tracking. 1 Tao Xiang 0002, Shaogang Gong, Dennis Parkinson |
BMVC | 2 |
| 2002 | Learning Intrinsic Video Content Using Levenshtein Distance in Graph Partitioning
Jeffrey Ng, Shaogang Gong |
ECCV (4) | 2 |
| 2002 | Understanding visual behaviour
Shaogang Gong, Hilary Buxton |
Image Vis. Comput. | 1 |
| 2002 | On the semantics of visual behaviour, structured events and trajectories of human action
Shaogang Gong, Jeffrey Ng Sing Kwong, Jamie Sherrah |
Image Vis. Comput. | 1 |
| 2002 | Corresponding dynamic appearances
Shaogang Gong, Alexandra Psarrou, Sami Romdhani |
Image Vis. Comput. | 1 |
| 2002 | Composite support vector machines for detection of faces across views and pose estimation
Jeffrey Ng Sing Kwong, Shaogang Gong |
Image Vis. Comput. | 2 |
| 2002 | The dynamics of linear combinations: tracking 3D skeletons of human subjects
Eng-Jon Ong, Shaogang Gong |
Image Vis. Comput. | 2 |
| 2002 | Recognition of human gestures and behaviour based on motion trajectories
Alexandra Psarrou, Shaogang Gong, Michael Walter 0004 |
Image Vis. Comput. | 2 |
| 2001 | Recognising Trajectories of Facial Identities Using Kernel Discriminant AnalysisabstractWe present a comprehensive approach to address three challenging problems in face recognition: modelling faces across multi-views, extracting the nonlinear discriminating features, and recognising moving faces dynamically in image sequences. A multi-view dynamic face model is designed to extract the shape-and-pose-free facial texture patterns. Kernel discriminant analysis, which employs the kernel technique to perform linear discriminant analysis in a high-dimensional feature space, is developed to extract the significant nonlinear features which maximise the between-class variance and minimise the within-class variance. Finally, an identity surface based face recognition is performed dynamically from video input by matching object and model trajectories. q 2003 Elsevier B.V. All rights reserved. Yongmin Li 0001, Shaogang Gong, Heather M. Liddell |
BMVC | 2 |
| 2001 | Learning Pixel-Wise Signal Energy for Understanding SemanticsabstractVisual interpretation of events requires both an appropriate representation of change occurring in the scene and the application of semantics for differentiating between different types of change. Conventional approaches for tracking objects and modelling object dynamics make use of either temporal region-correlation or pre-learnt shape or appearance models. We propose a new pixel-level approach for learning the temporal characteristics of change at individual pixels. Gaussian mixture models are used to model slow long-term changes in pixel distributions while pixel energy histories are used to extract fast-change signatures from short-term events and modelled by CONDENSATION matching. Jeffrey Ng, Shaogang Gong |
BMVC | 2 |
| 2001 | Data Driven Model Acquisition using Minimum Description Length
Michael Walter 0004, Alexandra Psarrou, Shaogang Gong |
BMVC | 3 |
| 2001 | Constructing Facial Identity Surfaces in a Nonlinear Discriminating SpaceabstractRecognising face with large pose variation is more challenging than that in a fixed view, e.g. frontal-view, due to the severe non-linearity caused by rotation in depth, self-shading and self-occlusion. To address this problem, a multi-view dynamic face model is designed to extract the shape-and-pose-free facial texture patterns from multi-view face images. Kernel Discriminant Analysis is developed to extract the significant non-linear discriminating features which maximise the between-class variance and minimise the within-class variance. By using the kernel technique, this process is equivalent to a Linear Discriminant Analysis in a high-dimensional feature space which can be solved conveniently. The identity surfaces are then constructed from these non-linear discriminating features. Face recognition can be performed dynamically from an image sequence by matching an object trajectory and model trajectories on the identity surfaces. Yongmin Li 0001, Shaogang Gong, Heather M. Liddell |
CVPR (2) | 2 |
| 2001 | Modelling Faces Dynamically across Views and Over Time
Yongmin Li 0001, Shaogang Gong, Heather M. Liddell |
ICCV | 2 |
| 2001 | Continuous Global Evidence-Based Bayesian Modality Fusion for Simultaneous Tracking of Multiple ObjectsabstractRobust, real-time tracking of objects from visual data requires probabilistic fusion of multiple visual cues. Previous approaches have either been ad hoc or relied on a Bayesian network with discrete spatial variables which suffers from discretisation and computational complexity problems. We present a new Bayesian modality fusion network that uses continuous domain variables. The network architecture distinguishes between cues that are necessary or unnecessary for the object's presence. Computationally expensive and inexpensive modalities are also handled differently to minimise cost. The method provides a formal, tractable and robust probabilistic method for simultaneously tracking multiple objects. While instantaneous inference is exact, approximation is required for propagation over time. Jamie Sherrah, Shaogang Gong |
ICCV | 2 |
| 2001 | Face distributions in similarity space under varying head pose
Jamie Sherrah, Shaogang Gong, Eng-Jon Ong |
Image Vis. Comput. | 2 |
| 2001 | Fusion of perceptual cues for robust tracking of head pose and position
Jamie Sherrah, Shaogang Gong |
Pattern Recognit. | 2 |
| 2000 | Tracking Multiple People Under Occlusion Using Multiple CamerasabstractWe describe a system for tracking multiple people with multiple cameras based on fusion of multiple cues. Face trackers are used to self-calibrate our system. Epipolar geometry and landmarks are employed to disambiguate the tracking problem. The correlation of visual information between different cameras is learnt using Support Vector Regression and Hierarchical Principal Component Analysis to estimate the subject appearance across cameras. The joint features of subjects extracted from multiple cameras are tracked and used as a model to re-track people once the subjects are lost tracking in the system. Results demonstrate that our system can deal with the occlusion. 1 Ting-Hsun Chang, Shaogang Gong, Eng-Jon Ong |
BMVC | 2 |
| 2000 | Recognising the Dynamics of Faces across Multiple ViewsabstractWe present an integrated framework for dynamic face detection and recognition, where head pose is estimated using Support Vector Regression, face detection is performed by Support Vector Classification, and recognition is carried out in a feature space constructed by Linear Discriminant Analysis. Unlike most traditional approaches to matching the patterns from static face images, we model the dynamics of human faces from video sequences in a consistent spatio-temporal context, i.e. recognition is accomplished by matching an object trajectory to a set of identity model trajectories in feature space. The model trajectories are synthesized from only a few views which sparsely cover the view sphere. Compared with the static face matching techniques, this approach is more robust and accurate under a coarse correspondence of face images, and has potential to visual interaction and advanced human behaviour recognition in real-world scenarios. 1 Introduction The issue of face rec... Yongmin Li 0001, Shaogang Gong, Heather M. Liddell |
BMVC | 2 |
| 2000 | Quantifying Ambiguities in Inferring Vector-Based 3D ModelsabstractThis paper presents a framework for directly addressing issues arising from self-occlusions and ambiguities due to the lack of depth information in vector-based representations. Visual data directly observed from an image are used to indirectly recover the parameters of an underlying dynamic model of an articulated object. The proposed framework allows us to learn the ambiguities of a representation from training examples. The resulting model is then used to measure the ambiguities of each estimated underlying model parameter given the available visual information. This provides an indication of how much we can “trust ” the visual data for estimating certain parts of the model. We then provide a working example of multi-view data fusion for tracking 3D skeletons of articulated objects in a multi-camera environment. 1 Eng-Jon Ong, Shaogang Gong |
BMVC | 2 |
| 2000 | Resolving Visual Uncertainty and Occlusion through Probabilistic ReasoningabstractTracking interacting human body parts from a single two-dimensional view is difficult due to occlusion, ambiguity and spatio-temporal discontinuities. We present a Bayesian network method for this task. The method is not reliant upon spatio-temporal continuity, but exploits it when present. Our inference-based tracking model is compared with a CONDENSATION model aug-mented with a probabilistic exclusion mechanism. We show that the Bayesian network has the advantages of fully modelling the state space, explicitly rep-resenting domain knowledge, and handling complex interactions between variables in a globally consistent and computationally effective manner. 1 Jamie Sherrah, Shaogang Gong |
BMVC | 2 |
| 2000 | On Utilising Template and Feature-Based Correspondence in Multi-view Appearance Models
Sami Romdhani, Alexandra Psarrou, Shaogang Gong |
ECCV (1) | 3 |
| 2000 | Tracking Discontinuous Motion Using Bayesian Inference
Jamie Sherrah, Shaogang Gong |
ECCV (2) | 2 |
| 2000 | Support Vector Regression and Classification Based Multi-View Face Detection and RecognitionabstractA support vector machine-based multi-view face detection and recognition framework is described. Face detection is carried out by constructing several detectors, each of them in charge of one specific view. The symmetrical property of face images is employed to simplify the complexity of the modelling. The estimation of head pose, which is achieved by using the support vector regression technique, provides crucial information for choosing the appropriate face detector. This helps to improve the accuracy and reduce the computation in multi-view face detection compared to other methods. For video sequences, further computational reduction can be achieved by using a pose change smoothing strategy. When face detectors find a face in frontal view, a support vector machine-based multi-class classifier is activated for face recognition. All the above issues are integrated under a support vector machine framework. Test results on four video sequences are presented, among them the detection rate is above 95%, recognition accuracy is above 90%, average pose estimation error is around 10/spl deg/, and the full detection and recognition speed is up to 4 frames/second on a Pentium II 300 PC. Yongmin Li 0001, Shaogang Gong, Heather M. Liddell |
FG | 2 |
| 2000 | A Generic Face Appearance Model of Shape and Texture under Very Large Pose Variations from Profile to Profile ViewsabstractModelling the appearance of 3D objects under very large pose variations relies on recovering correspondence between local features and the texture variations across views. However, changes in object 3D pose introduce self-occlusions and cause nonlinear variations in both the shape and the texture of object appearance. In this paper we introduce a pose-invariant appearance model that utilises a generic-view shape template for alignment, Kernel PCA for modelling shape and texture nonlinearities across views of large pose variations, and a neural network for model fitting to new images. Sami Romdhani, Alexandra Psarrou, Shaogang Gong |
ICPR | 3 |
| 2000 | VIGOUR: A System for Tracking and Recognition of Multiple People and their ActivitiesabstractTracking multiple people and their behaviours is a fundamental task for visually mediated interaction. We present VIGOUR, a platform for simultaneously tracking of multiple people and recognition of their behaviours for high-level interpretation. Through perceptual integration, different types of visual information are fused to combine the benefits of different techniques. Robust low-level visual cues such as skin colour and motion are used to focus attention and facilitate real-time tracking. VIGOUR tracks behaviours using gestures and head pose to produce a high-level behaviour representation for subsequent interpretation. The system is able to track three people and recognise their gestures simultaneously in real time. Jamie Sherrah, Shaogang Gong |
ICPR | 2 |
| 2000 | Interpretation of Group Behavior in Visually Mediated InteractionabstractWhile full computer understanding of dynamic visual scenes containing several people may be currently unattainable, we propose a computationally efficient approach to determine areas of interest in such scenes. We present methods for modelling and interpretation of multi-person human behaviour in real time to control video cameras for visually mediated interaction. Jamie Sherrah, Shaogang Gong, A. Jonathan Howell, Hilary Buxton |
ICPR | 2 |
| 2000 | Multi-view face detection using support vector machines and eigenspace modellingabstractAn approach to multi-view face detection based on head pose estimation is presented in this paper. Support vector regression is employed to solve the problem of pose estimation. Three methods, the eigenface method the support vector machine (SVM) based method, and a combination of the two methods, are investigated. The eigenface method, which seeks to estimate the overall probability distribution of patterns to be recognised, is fast but less accurate because of the overlap of confidence distributions between face and non-face classes. On the other hand, the SVM method, which tries to model the boundary of two classes to be classified is more accurate but slower as the number of support vectors is normally large. The combined method can achieve an improved performance by speeding up the computation and keeping the accuracy to a preset level. It can be used to automatically detect and track faces in face verification and identification systems. Yongmin Li 0001, Shaogang Gong, Jamie Sherrah, Heather M. Liddell |
KES | 2 |
| 1999 | Learning Support Vector Machines for A Multi-View Face ModelabstractSupport Vector Machines have shown great potential for learning classification functions that can be applied to object recognition. In this work, we extend SVMs to model the appearance of human faces which undergo nonlinear change across multiple views. The approach uses inherent factors in the nature of the input images and the SVM classification algorithm to perform both multi-view face detection and pose estimation. Jeffrey Ng Sing Kwong, Shaogang Gong |
BMVC | 2 |
| 1999 | A Dynamic 3D Human Model using Hybrid 2D-3D Representations in Hierarchical PCA SpaceabstractWe propose a novel framework for a hybrid 2D-3D dynamic human model with which robust matching and tracking of a 3D skeleton model of a human body among multiple views can be performed. We describe a method that measures the image ambiguity at each view. The 3D skeleton model and the correspondence between the model and its 2D images are learnt using hierarchical principal component analysis. Tracking in individual views is performed based on CONDENSATION. Eng-Jon Ong, Shaogang Gong |
BMVC | 2 |
| 1999 | A Multi-View Nonlinear Active Shape Model Using Kernel PCAabstractRecovering the shape of any 3D object using multiple 2D views requires establishing correspondence between feature points at different views. How-ever changes in viewpoint introduce self-occlusions, resulting nonlinear vari-ations in the shape and inconsistent 2D features between views. Here we introduce a multi-view nonlinear shape model utilising 2D view-dependent constraint without explicit reference to 3D structures. For nonlinear model transformation, we adopt Kernel PCA based on Support Vector Machines. 1 Sami Romdhani, Shaogang Gong, Alexandra Psarrou |
BMVC | 2 |
| 1999 | Fusion of Perceptual Cues using Covariance EstimationabstractThe paradigm of perceptual integration provides robust solutions to computer vision problems. By combining the outputs of multiple vision modules, the assumptions and constraints of each module are factored out to result in a more robust system overall. The integration of dierent modules can be regarded as a form of data fusion. To this end, we propose a framework for fusing dierent information sources through estimation of covariance from observations. The framework is demonstrated in a pose tracking method that fuses similarity-to-prototypes measures and skin colour to track head pose and face position. The use of data fusion through covariance introduces constraints that allow the tracker to robustly estimate head pose and track face position. Contact Author Jamie R. Sherrah Contact Address Department of Computer Science Queen Mary and Westeld College Mile End, E1 4NS London UK Contact email [email protected] Contact Tel. (+44) (0)171 975 5230 Contact Fax (+44) (0)181 980 Bri... Jamie Sherrah, Shaogang Gong |
BMVC | 2 |
| 1999 | Understanding Pose Discrimination in Similarity SpaceabstractIdentity-independent estimation of head pose from prototype im-ages is a perplexing task, requiring pose-invariant face detection. The problem is exacerbated by changes in illumination, identity and facial position. Facial images must be transformed in such a way as to em-phasise dierences in pose, while suppressing dierences in identity. We investigate appropriate transformations for use with a similarity-to-prototypes philosophy. The results show that orientation-selective Gabor lters enhance dierences in pose, and that dierent lter ori-entations are optimal at dierent poses. In contrast, PCA was found to provide an identity-invariant representation in which similarities can be calculated more robustly. We also investigate the angular resolution at which pose changes can be resolved using our methods. An angular resolution of 10 was found to be suciently discriminable at some poses but not at others, while 20 is quite acceptable at most poses. 1 Jamie Sherrah, Shaogang Gong, Eng-Jon Ong |
BMVC | 2 |
| 1999 | Learning Prior and Observation Augmented Density Models for Behaviour RecognitionabstractRecognition of human behaviours requires modeling the underlying spatial and temporal structures of their motion patterns. Such structures are intrinsically probabilistic and therefore should be modelled as stochastic processes. In this paper we introduce a framework to recognise behaviours based on both learning prior and continuous propagation of density models of behaviour patterns. Prior is learned from training sequences using hidden Markov models and density models are augmented by current visual observation. Michael Walter 0004, Alexandra Psarrou, Shaogang Gong |
BMVC | 3 |
| 1999 | Recognition of Temporal Structures: Learning Prior and Propagating Observation Augmented Densities via Hidden Markov StatesabstractAn algorithm is described for modelling and recognising temporal structures of visual activities. The method is based on (1) learning prior probabilistic knowledge using hidden Markov models, (2) automatic temporal clustering of hidden Markov states based on expectation maximisation and (3) using observation augmented conditional density distributions to reduce the number of samples required for propagation and therefore improve recognition speed and robustness. Shaogang Gong, Michael Walter 0004, Alexandra Psarrou |
ICCV | 1 |
| 1999 | Tracking colour objects using adaptive mixture models
Stephen J. McKenna, Yogesh Raja, Shaogang Gong |
Image Vis. Comput. | 3 |
| 1998 | Appearance-Based Face Recognition under Large Head Rotations in Depth
Shaogang Gong, Eng-Jon Ong, Peter J. Loft |
ACCV (2) | 1 |
| 1998 | Face Recognition from Sequences Using Models of Identity
Stephen J. McKenna, Shaogang Gong |
ACCV (1) | 2 |
| 1998 | Object Tracking Using Adaptive Color Mixture Models
Stephen J. McKenna, Yogesh Raja, Shaogang Gong |
ACCV (1) | 3 |
| 1998 | Segmentation and Tracking Using Color Mixture Models
Yogesh Raja, Stephen J. McKenna, Shaogang Gong |
ACCV (1) | 3 |
| 1998 | Learning to Associate Faces across Views in Vector Space of Similarities to PrototypesabstractWe present a method for learning appearance models that can be used to recognise and track both 3D head pose and identities of novel subjects with continuous head movement across the view-sphere. We describe an automatic face data acquisition system based on a magnetic sensor and a calibrated camera. The system enabled us to obtain systematically a database of face images with labelled 3D poses across a view-sphere of \\Sigma90 ffi yaw and \\Sigma30 ffi tilt at intervals of 10 ffi . The database was used to learn appearance models of unseen faces based on similarity measures to prototype faces. The method is computationally efficient and enables real-time performance with ease. 1 Introduction To be able to recognise faces of moving people not only requires the ability to label novel face images with known identities, but also needs detecting and tracking of faces over time [1]. We refer to this as the task of associating faces. We adopt the view such a task can be better achieved... Shaogang Gong, Eng-Jon Ong, Stephen J. McKenna |
BMVC | 1 |
| 1998 | Gesture Recognition for Visually Mediated Interaction using Probabilistic Event TrajectoriesabstractAn approach to gesture recognition is presented in which gestures are modelled probabilistically as sequences of visual events. These events are matched to visual input using probabilistic models estimated from motion feature trajectories. The features used are motion image moments. The method was applied to a set of gestures defined within the context of an application in visually mediated interaction in which they would be used to control an active teleconferencing camera. The approach is computationally efficient, allowing real-time performance to be obtained. 1 Stephen J. McKenna, Shaogang Gong |
BMVC | 2 |
| 1998 | Colour Model Selection and Adaption in Dynamic Scenes
Yogesh Raja, Stephen J. McKenna, Shaogang Gong |
ECCV (1) | 3 |
| 1998 | View-Based Adaptive Affine Tracking
Fernando De la Torre, Shaogang Gong, Stephen J. McKenna |
ECCV (1) | 2 |
| 1998 | Tracking and Segmenting People in Varying Lighting Conditions Using Colour
Yogesh Raja, Stephen J. McKenna, Shaogang Gong |
FG | 3 |
| 1998 | View Alignment with Dynamically Updated Affine Tracking
Fernando De la Torre, Shaogang Gong, Stephen J. McKenna |
FG | 2 |
| 1998 | Modelling facial colour and identity with Gaussian mixtures
Stephen J. McKenna, Shaogang Gong, Yogesh Raja |
Pattern Recognit. | 2 |
| 1997 | Face Recognition in Dynamic Scenes
Stephen J. McKenna, Shaogang Gong, Yogesh Raja |
BMVC | 2 |
| 1996 | Face Tracking and Pose RepresentationabstractWe describe a dynamic face tracking system based on an integrated motion-based object tracking and model-based face detection framework. The motion-based tracker focuses attention for the face detector whilst the latter aids the tracking process. The system produces segmented face sequences from complex scenes with poor viewing conditions in surveillance applications. We also investigate a Gabor wavelet transform as a representation scheme for capturing head rotations in depth. Principal components analysis was used to visualise the manifolds described by pose changes. Qualitative results are given. 1 Introduction In order to analyse and recognise peoples' faces in realistically unconstrained environments, robust tracking and segmentation are needed to provide sequences of normalised face images. Although such a normalisation process is often treated as a separate preprocessing step, it is an inherent part of face recognition. The ability of a system to produce normalised face sequenc... Stephen J. McKenna, Shaogang Gong, J. J. Collins |
BMVC | 2 |
| 1996 | An Investigation into Face Pose DistributionsabstractVisual perception of faces is invariant under many transformations, perhaps the most problematic of which is pose change (face rotating in depth). We use a variation of Gabor wavelet transform (GWT) as a representation framework for investigating face pose measurement. Dimensionality reduction using principal components analysis (PCA) enables pose changes to be visualised as manifolds in low-dimensional subspaces and provides a useful mechanism for investigating these changes. The effectiveness of measuring face pose with GWT representations was examined using PCA. We discuss our experimental results and draw a few preliminary conclusions. Shaogang Gong, Stephen J. McKenna, J. J. Collins |
FG | 1 |
| 1996 | Tracking FacesabstractRobust tracking and segmentation of faces is a prerequisite for face analysis and recognition. In this paper we describe an approach to this problem which is well suited to surveillance applications with poorly constrained viewing conditions. It integrates motion-based tracking with model based face detection to produce segmented face sequences from complex scenes containing several people. The motion of moving image contours was estimated using temporal convolution and a temporally consistent list of moving objects was maintained. Objects were tracked using Kalman filters. Faces were detected using a neural network. The essence of the system is that the motion tracker is able to focus attention for a face detection network whilst the latter is used to aid the tracking process. Stephen J. McKenna, Shaogang Gong |
FG | 2 |
| 1995 | Visual Surveillance in a Dynamic and Uncertain World
Hilary Buxton, Shaogang Gong |
Artif. Intell. | 2 |
| 1993 | From Contextual Knowledge to Computational ConstraintsabstractIn this work we address the issue of focused computation in computer vision for effectiveness and efficiency- In particular, we propose a scheme that links scene-oriented contextual knowledge with the computational constraints required in visual motion segmentation and tracking. The approach uses Ra.yesia.ii belief revision techniques to map explicit scene knowledge onto implicit causal dependent constraints in controlling computational parameters. We discuss our experimental results from applying this method in improving existing techniques in traffic surveillance applications. 1 Shaogang Gong, Hilary Buxton |
BMVC | 1 |
| 1992 | On the Visual Expectations of Moving Objects
Shaogang Gong, Hilary Buxton |
ECAI | 1 |
| 1990 | Parallel Computation Of Optic Flow
Shaogang Gong, J. Michael Brady |
ECCV | 1 |