VLDB 2026 Research / reviewers in the wild / expert
Weiming Hu 0004
dblp:41/6824-4
· DBLP profile ↗
336ranked-venue papers
32as first author
113since 2021 · last 2026
0000-0001-9237-8825ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 205 · 8 first-author · 67 since 2021Artificial intelligence and machine learning · 204 · 17 first-author · 75 since 2021Databases, data management, data science and information retrieval · 32 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 11 · 5 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HDGS: Hierarchical Dynamic Gaussian Splatting for Urban Driving ScenesabstractThis paper tackles the challenging task of achieving storage-efficient yet high-fidelity motion representation in large-scale dynamic 3D Gaussian Splatting. Our motivation stems from the truth that existing urban-scale methods, which rely on massive and unstructured individual Gaussians for scene modeling, face a critical scalability bottleneck. Inspired by recent advances in the 3DGS-based compression beyond autonomous driving, we address this challenge by leveraging the compression capability of anchor-driven methods. However, this is non-trivial as our exploratory experiments reveal that the direct application of this paradigm to dynamic, large-scale urban scenes results in performance degradation. We attribute this phenomenon to the hierarchical anchor design that severely loses dynamic information. To this end, we propose Hierarchical Dynamic Gaussian Splatting (HDGS), a novel framework designed to adapt the anchor-based Gaussian paradigm to 4D urban environments. We first establish a local support network to reinforce inter-anchor consistency, mitigating geometric and appearance fractures caused by supervision attenuation in deep hierarchies. Then, we handle heterogeneous object motion via coarse-to-fine decomposition, where high-level anchors model coarse dynamics and low-level anchors refine them with residual deformations. Third, we introduce a hybrid supervision scheme that fuses global geometric constraints and local pixel-level cues to alleviate geometrically inconsistent reconstruction under sparse LiDAR. Extensive experiments show that HDGS reduces storage by 69.0% while maintaining or even improving rendering fidelity compared to state-of-the-art methods. Fudong Ge, Hanshi Wang, Weiming Hu 0004 |
AAAI | 6 |
| 2026 | Integrating Diverse Assignment Strategies into DETRsabstractLabel assignment is a critical component in object detectors, particularly within DETR-style frameworks where the one-to-one matching strategy, despite its end-to-end elegance, suffers from slow convergence due to sparse supervision. While recent works have explored one-to-many assignments to enrich supervisory signals, they often introduce complex, architecture-specific modifications and typically focus on a single auxiliary strategy, lacking a unified and scalable design. In this paper, we first systematically investigate the effects of ``one-to-many'' supervision and reveal a surprising insight that performance gains are driven not by the sheer quantity of supervision, but by the diversity of the assignment strategies employed. This finding suggests that a more elegant, parameter-efficient approach is attainable. Building on this insight, we propose LoRA-DETR, a flexible and lightweight framework that seamlessly integrates diverse assignment strategies into any DETR-style detector. Our method augments the primary network with multiple Low-Rank Adaptation (LoRA) branches during training, each instantiating a different one-to-many assignment rule. These branches act as auxiliary modules that inject rich, varied supervisory gradients into the main model and are discarded during inference, thus incurring no additional computational cost. This design promotes robust joint optimization while maintaining the architectural simplicity of the original detector. Extensive experiments on different baselines validate the effectiveness of our approach. Our work presents a new paradigm for enhancing detectors, demonstrating that diverse ``one-to-many'' supervision can be integrated to achieve state-of-the-art results without compromising model elegance. Hanshi Wang, Fudong Ge, Guan Luo, Weiming Hu 0004 |
AAAI | 6 |
| 2026 | MMhops-R1: Multimodal Multi-hop ReasoningabstractThe ability to perform multi-modal multi-hop reasoning by iteratively integrating information across various modalities and external knowledge is critical for addressing complex real-world challenges. However, existing Multi-modal Large Language Models (MLLMs) are predominantly limited to single-step reasoning, as existing benchmarks lack the complexity needed to evaluate and drive multi-hop abilities. To bridge this gap, we introduce MMhops, a novel, large-scale benchmark designed to systematically evaluate and foster multi-modal multi-hop reasoning. MMhops dataset comprises two challenging task formats, Bridging and Comparison, which necessitate that models dynamically construct complex reasoning chains by integrating external knowledge. To tackle the challenges posed by MMhops, we propose MMhops-R1, a novel multi-modal Retrieval-Augmented Generation (mRAG) framework for dynamic reasoning. Our framework utilizes reinforcement learning to optimize the model for autonomously planning reasoning paths, formulating targeted queries, and synthesizing multi-level information. Comprehensive experiments demonstrate that MMhops-R1 significantly outperforms strong baselines on MMhops, highlighting that dynamic planning and multi-modal knowledge integration are crucial for complex reasoning. Moreover, MMhops-R1 demonstrates strong generalization to tasks requiring fixed-hop reasoning, underscoring the robustness of our dynamic planning approach. Ziqi Zhang 0010, Zongyang Ma, Bing Li 0001, Chunfeng Yuan, Guangting Wang, Fengyun Rao, Ying Shan, Weiming Hu 0004 |
AAAI | 10 |
| 2026 | SCG-SSC: Semantic Scene Completion via Self-and-Cross Gated Fusion of Depth Maps and Semantic Priors
Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge |
ICPR (9) | 4 |
| 2026 | Open-Tag: A Generative Framework for Open-World Multimodal Tagging
Ziqi Zhang 0010, Zongyang Ma, Peijin Wang, Bing Li 0001, Chunfeng Yuan, Weiming Hu 0004 |
Int. J. Comput. Vis. | 6 |
| 2026 | MSFI: Multi-timescale spatio-temporal features integration in spiking neural networks
Dengfeng Xue, Chunfeng Yuan, Man Yao, Wei Liu 0153, Li Yang 0014, Bing Li 0001, Weiming Hu 0004, Haoliang Sun, Zhetao Li |
Neural Networks | 9 |
| 2026 | Reinforcement Learning-Based Sequential Parameter Tuning for Image Signal ProcessingabstractHardware image signal processing (ISP) transforms RAW inputs into high-quality RGB images through a series of processing modules, each with numerous tunable parameters. Traditionally, these parameters are manually tuned by imaging experts, a time-consuming and subjective process. Recent deep learning approaches predict ISP parameters, but often treat the process as a black box and overlook the intrinsic relationships among ISP modules. To address these fundamental issues, we introduce a novel ISP parameter optimization model based on single-agent reinforcement learning (RL) (i.e., SARL-ISP), formulating the hardware ISP parameter tuning as a sequential optimization problem. During the optimization process, the agent updates ISP parameter tuning strategies for different tasks through interaction with the environment. In order to explore the influence of the sequential structure of hardware ISP modules and the coupling relationships among ISP parameters on the tuning process, we further propose a sequential ISP framework based on collaborative multi-agent RL (i.e., MARL-ISP). Specifically, the serialized parameter tuning module (SPTM) realistically simulates the process of manual prediction and module pipeline. Additionally, the feature selection module (FSM) facilitates the transmission and fusion of agent features, thereby selecting more appropriate feature inputs for downstream tasks. Extensive experiments across various tasks (e.g., object detection, instance segmentation) validate the effectiveness and efficiency of our models. Even with minimal training data, our models also outperform current state-of-the-art methods in both quantitative metrics and qualitative evaluations. Bing Li 0001, Congyan Lang, Zhikun Zhao, Juan Wang 0012, Weihua Xiong, Weiming Hu 0004, Long Cheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Multi-modal face anti-spoofing via self-supervised learning
Yufan Liu 0001, Lai Jiang 0004, Shengxi Li, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jinlong Lin |
Pattern Recognit. Lett. | 7 |
| 2026 | Iter3DDet: Depth-Guided Iterative Fusion and Refinement for Monocular 3D Object DetectionabstractMonocular 3D object detection offers significant potential for autonomous systems due to its inherent cost-effectiveness and scalability. While DETR-based architectures excel in 2D vision tasks, critical limitations persist in extending them effectively to monocular 3D detection, as evidenced in existing frameworks like MonoDETR and MonoDGP. These methods typically suffer from inefficient serial fusion of multimodal features and lack iterative refinement mechanisms, limiting their performance, especially for mid-to-long range targets. To overcome these shortcomings, we propose Iter3DDet, a novel depth-guided iterative refinement framework that integrates fine-grained feature fusion to significantly enhance detection performance. The core novelty of our approach lies in two key innovations: (1) A hybrid feature encoder combining MonoDGP’s region segmentation head with MonoDETR’s visual backbone, augmented by a multi-scale context attention module that dynamically aggregates structural and semantic cues across pyramid levels, eliminating heuristic fusion rules; (2) A depth-guided adaptive cross-modal decoder that iteratively fuses depth and context features through prioritized attention mechanisms, coupled with a novel iterative refinement training strategy that progressively refines 3D detection hypotheses, substantially improving accuracy across targets of varying difficulty levels. Extensive experiments on the KITTI, nuScenes, and Waymo benchmarks demonstrate Iter3DDet’s state-of-the-art performance, validating the effectiveness of our iterative refinement paradigm. The code will be open-sourced at https://github.com/PCwenyue. Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Motion-Guided Disentanglement for Point Cloud Masked AutoencodersabstractMasked autoencoders have been extended beyond images, but random masking often fails to capture non-uniform motion regions in inherently disordered and irregular data like point cloud videos. In this paper, we propose a motion-guided disentanglement method (MGD) to improve masked autoencoders for point cloud video representation learning. Specifically, we begin by estimating motion intensity using an optimal transport approach, which guides the separate masking of dynamic and static regions. This motion-guided masking ensures balanced coverage, addressing the limitations of random masking in capturing non-uniformly distributed motion regions. Furthermore, we disentangle the prediction tasks into motion prediction for high-motion point tubes and appearance reconstruction for low-motion ones. This disentanglement enables the model to more effectively capture both motion and appearance in point cloud videos. We conducted experiments on four widely used point cloud video datasets—NTU RGB+D, MSR-Action3D, NvGesture, and SHREC’17—which demonstrate that our approach consistently improves masked autoencoders for point cloud video representation learning, achieving new state-of-the-art results. Code will be publicly available on GitHub. Haoran Wang 0001, Shaqing Song, Baosheng Yu, Tong Jia 0001, Dongyue Chen 0001, Chunfeng Yuan, Weiming Hu 0004, Haibin Ling |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Deepfake Detection via Exploring Degradation InconsistencyabstractThe detection of face forgery has become increasingly vital due to the severe security concerns posed by face manipulation techniques. While recent studies on forgery detection have demonstrated promising results when the training and testing samples come from the same domains, the problem remains challenging when attempting to extend the detector to unseen methods. In this work, we propose an innovative approach to enhance the generalization capability of forgery detection methods by exploring degradation inconsistency clues interspersed between the background and the manipulated face regions. Our motivation stems from the observation that digital photos undergo different degradation during acquisition and transmission, resulting in backgrounds and faces from different sources containing distinct degradation patterns in the forged faces. The proposed framework, termed the Degradation Consistency Learning Framework, integrates two core components: a data generation network that modulates degradation transformations to obtain tampered facial images, and a detection network that mines degradation inconsistency clues from both spatial and frequency domains. These two components are tightly coupled through adversarial training, forming a dynamic architecture akin to a Generative Adversarial Network (GAN). Experimental results on different benchmark and evaluation protocols (i.e., indataset and cross-dataset) have demonstrated the effectiveness of our method. Weiming Bai, Yufan Liu 0001, Aixi Zhang, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | DRDFNet: A Degradation-Aware Restoration and Detail-Preserving Fusion Network for Infrared and Visible ImageabstractMulti-source image fusion combines infrared and visible information to improve scene perception in applications such as drone reconnaissance and autonomous driving. However, most existing infrared-visible image fusion methods are developed under ideal imaging assumptions. In adverse environments, visible images often lose structural and textural details, whereas infrared images are affected by noise, stripe artifacts, and low contrast, leading to degraded fusion quality and weakened downstream perception performance. To address these limitations, we propose a unified Degradation-aware Restoration and Detail-preserving Fusion Network (DRDFNet), which consists of a Degradation-Aware Restoration Transformer and a Detail-Preserving Fusion Mamba. The restoration branch uses a Compound Degradation Restoration Module (CDRM) to remove complex degradations, while the fusion branch employs a Dynamic Feature Fusion Module (DFFM) to integrate local complementary cues and global correlations across modalities. A two-stage training strategy is further introduced to reduce the optimization conflict between restoration and fusion. In addition, we construct DIVIF, a large-scale degraded IVIF benchmark generated by a physics-based imaging simulator. Experiments on the DIVIF and AWMM-100k benchmarks demonstrate that DRDFNet achieves robust and competitive performance compared with SOTA methods. Both the dataset and source code will be made publicly available at https://github.com/Liupeng97/DRDFNet. Peng Liu 0024, An Wei, Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002 |
IEEE Trans. Image Process. | 7 |
| 2025 | Towards More Discriminative Feature Learning in SNNs with Temporal-Self-Erasing SupervisionabstractSpiking Neural Networks (SNNs) are biologically inspired models that process visual inputs over multiple time steps. However, they often struggle with limited feature discrimination along the temporal dimension due to inherent spatiotemporal invariance. This limitation arises from the redundant activation of certain regions and shared supervision for multiple time steps, constraining the network’s ability to adapt and learn diverse features. To address this challenge, we propose a novel Temporal-Self-Erasing (TSE) supervision method that dynamically adapts the learning regions of interest for different time steps. The TSE method operates by identifying highly activated regions from predictions across multiple time steps and adaptively suppressing them during model training, thereby encouraging the network to focus on less activated yet potentially informative regions. This approach not only enhances the feature discrimination capability of SNNs but also facilitates more effective multi-time-step inference by exploiting more semantic information. Experimental results on benchmark datasets demonstrate that our TSE method significantly improves the classification accuracy and robustness of SNNs. Wei Liu 0153, Li Yang 0014, Mingxuan Zhao, Dengfeng Xue, Shuxun Wang, Boyu Cai, Bing Li 0001, Weiming Hu 0004 |
AAAI | 10 |
| 2025 | Visual-Instructed Degradation Diffusion for All-in-One Image RestorationabstractImage restoration tasks like deblurring, denoising, and dehazing usually need distinct models for each degradation type, restricting their generalization in real-world scenarios with mixed or unknown degradations. In this work, we propose Defusion, a novel all-in-one image restoration framework that utilizes visual instruction-guided degradation diffusion. Unlike existing methods that rely on task-specific models or ambiguous text-based priors, Defusion constructs explicit visual instructions that align with the visual degradation patterns. These instructions are grounded by applying degradations to standardized visual elements, capturing intrinsic degradation features while agnostic to image semantics. Defusion then uses these visual instructions to guide a diffusion-based model that operates directly in the degradation space, where it reconstructs high-quality images by denoising the degradation effects with enhanced stability and generalizability. Comprehensive experiments demonstrate that Defusion outperforms state-of-the-art methods across diverse image restoration tasks, including complex and real-world degradations. Wenyang Luo, Haina Qin, Zewen Chen, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
CVPR | 9 |
| 2025 | Reversing Flow for Image RestorationabstractImage restoration aims to recover high-quality (HQ) images from degraded low-quality (LQ) ones by reversing the effects of degradation. Existing generative models for image restoration, including diffusion and score-based models, often treat the degradation process as a stochastic transformation, which introduces inefficiency and complexity. In this work, we propose ResFlow, a novel image restoration framework that models the degradation process as a deterministic path using continuous normalizing flows. ResFlow augments the degradation process with an auxiliary process that disambiguates the uncertainty in HQ prediction to enable reversible modeling of the degradation process. ResFlow adopts entropy-preserving flow paths and learns the augmented degradation flow by matching the velocity field. ResFlow significantly improves the performance and speed of image restoration, completing the task in fewer than four sampling steps. Extensive experiments demonstrate that ResFlow achieves state-of-the-art results across various image restoration benchmarks, offering a practical and efficient solution for real-world applications. Haina Qin, Wenyang Luo, Jingdong Chen, Ming Yang 0007, Bing Li 0001, Weiming Hu 0004 |
CVPR | 8 |
| 2025 | Dual-Head Feature Enhancement for Graph-Based Cross-View Multi-object Tracking
Wen-Juan Li, Weiming Hu 0004 |
ICANN (2) | 4 |
| 2025 | VisionMath: Vision-Form Mathematical Problem-Solving
Zongyang Ma, Ziqi Zhang 0010, Zhongang Oi, Chunfeng Yuan, Shaojie Zhu, Chengxiang Zhuo, Bing Li 0001, Ye Liu 0002, Zang Li, Ying Shan, Weiming Hu 0004 |
ICCV | 12 |
| 2025 | Height-Fidelity Dense Global Fusion for Multi-Modal 3D Object Detection
Hanshi Wang, Weiming Hu 0004 |
ICCV | 3 |
| 2025 | DeepTAGE: Deep Temporal-Aligned Gradient Enhancement for Optimizing Spiking Neural NetworksabstractSpiking Neural Networks (SNNs), with their biologically inspired spatio-temporal dynamics and spike-driven processing, are emerging as a promising low-power alternative to traditional Artificial Neural Networks (ANNs). However, the complex neuronal dynamics and non-differentiable spike communication mechanisms in SNNs present substantial challenges for efficient training. By analyzing the membrane potentials in spiking neurons, we found that their distributions can increasingly deviate from the firing threshold as time progresses, which tends to cause diminished backpropagation gradients and unbalanced optimization. To address these challenges, we propose Deep Temporal-Aligned Gradient Enhancement (DeepTAGE), a novel approach that improves optimization gradients in SNNs from both internal surrogate gradient functions and external supervision methods. Our DeepTAGE dynamically adjusts surrogate gradients in accordance with the membrane potential distribution across different time steps, enhancing their respective gradients in a temporal-aligned manner that promotes balanced training. Moreover, to mitigate issues of gradient vanishing or deviating during backpropagation, DeepTAGE incorporates deep supervision at both spatial (network stages) and temporal (time steps) levels to ensure more effective and robust network optimization. Importantly, our method can be seamlessly integrated into existing SNN architectures without imposing additional inference costs or requiring extra control modules. We validate the efficacy of DeepTAGE through extensive experiments on static benchmarks (CIFAR10, CIFAR100, and ImageNet-1k) and a neuromorphic dataset (DVS-CIFAR10), demonstrating significant performance improvements. Wei Liu 0153, Li Yang 0014, Mingxuan Zhao, Shuxun Wang, Bing Li 0001, Weiming Hu 0004 |
ICLR | 8 |
| 2025 | SSTrack: Sample-interval Scheduling for Lightweight Visual Object TrackingabstractIn recent years, CPU real-time object tracking has gained significant attention due to its broad applications such as UAV-tracking. To maintain computational efficiency, most existing CPU real-time object trackers rely on lightweight backbones and employ a single initial template image without intermediate online templates. Although the appearance variance between the template and the search is larger under this single template setting, the representation ability of lightweight backbones is weaker which poses a challenge when training lightweight object trackers. To address this issue, we propose SSTrack, a new easier-to-harder training schedule for the lightweight object tracker. From the data perspective, our method designed a success-aware sample scheduler that gradually increases difficult training samples with longer template-search time intervals and reduces the amount of the easier samples so the training cost remains unchanged. From the optimization perspective, we utilized a gradient scaling strategy that retains the original training objective of easier samples despite the reduction in their quantities. With the collective effort from both perspectives, our method achieves State-of-the-Art CPU-real-time accuracy on 5 UAV-tracking benchmarks and 5 general object tracking benchmarks. Codes and models will be available at https://github.com/Kou-99/SSTrack. Yutong Kou, Shubo Lin, Weiming Hu 0004 |
IJCAI | 5 |
| 2025 | Noise-Optimized Distribution Distillation for Dataset Condensation
Tongfei Liu, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Chenguang Ma |
ACM Multimedia | 4 |
| 2025 | SynCL: A Synergistic Training Strategy with Instance-Aware Contrastive Learning for End-to-End Multi-Camera 3D TrackingabstractWhile existing query-based 3D end-to-end visual trackers integrate detection and tracking via the *tracking-by-attention* paradigm, these two chicken-and-egg tasks encounter optimization difficulties when sharing the same parameters. Our findings reveal that these difficulties arise due to two inherent constraints on the self-attention mechanism, i.e., over-deduplication for object queries and self-centric attention for track queries. In contrast, removing self-attention mechanism not only minimally impacts regression predictions of the tracker, but also tends to generate more latent candidate boxes. Based on these analyses, we present SynCL, a novel plug-and-play synergistic training strategy designed to co-facilitate multi-task learning for detection and tracking. Specifically, we propose a Task-specific Hybrid Matching module for a weight-shared cross-attention-based decoder that matches the targets of track queries with multiple object queries to exploit promising candidates overlooked by the self-attention mechanism and the bipartite matching. To flexibly select optimal candidates for the one-to-many matching, we also design a Dynamic Query Filtering module controlled by model training status. Moreover, we introduce Instance-aware Contrastive Learning to break through the barrier of self-centric attention for track queries, effectively bridging the gap between detection and tracking. Without additional inference costs, SynCL consistently delivers improvements in various benchmarks and achieves state-of-the-art performance with $58.9\%$ AMOTA on the nuScenes dataset. Code and raw results are available at <https://github.com/shubolin028/SynCL>. Shubo Lin, Yutong Kou, Zirui Wu, Shaoru Wang, Bing Li 0001, Weiming Hu 0004 |
NeurIPS | 6 |
| 2025 | Online Segment Any 3D Thing as Instance TrackingabstractOnline, real-time, and fine-grained 3D segmentation constitutes a fundamental capability for embodied intelligent agents to perceive and comprehend their operational environments. Recent advancements employ predefined object queries to aggregate semantic information from Vision Foundation Models (VFMs) outputs that are lifted into 3D point clouds, facilitating spatial information propagation through inter-query interactions.
Nevertheless, perception, whether human or robotic, is an inherently dynamic process, rendering temporal understanding a critical yet overlooked dimension within these prevailing query-based pipelines. This deficiency in temporal reasoning can exacerbate issues such as the over-segmentation commonly produced by VFMs, necessitating more handcrafted post-processing. Therefore, to further unlock the temporal environmental perception capabilities of embodied agents, our work reconceptualizes online 3D segmentation as an instance tracking problem (AutoSeg3D). Our core strategy involves utilizing object queries for temporal information propagation, where long-term instance association promotes the coherence of features and object identities, while short-term instance update enriches
instant observations. Given that viewpoint variations in embodied robotics often lead to partial object visibility across frames, this mechanism aids the model in developing a holistic object understanding beyond incomplete instantaneous views. Furthermore, we introduce spatial consistency learning to mitigate the fragmentation problem inherent in VFMs, yielding more comprehensive instance information for enhancing the efficacy of both long-term and short-term temporal learning. The temporal information exchange and consistency learning facilitated by these sparse object queries not only enhance spatial comprehension but also circumvent the computational burden associated with dense temporal point cloud interactions. Our method establishes a new state-of-the-art, surpassing ESAM by 2.8 AP on ScanNet200 and delivering consistent gains on ScanNet, SceneNN, and 3RScan datasets, corroborating that identity-aware temporal reasoning is a crucial, previously underemphasized component for robust 3D segmentation in real-time embodied intelligence. Code is at https://github.com/AutoLab-SAI-SJTU/AutoSeg3D. Hanshi Wang, Zijian Cai, Weiming Hu 0004 |
NeurIPS | 5 |
| 2025 | Each Complexity Deserves a Pruning PolicyabstractThe established redundancy in visual tokens within large vision–language models (LVLMs) allows for pruning to effectively reduce their substantial computational demands. Empirical evidence from previous works indicates that visual tokens in later decoder stages receive less attention than shallow layers. Then, previous methods typically employ heuristics layer-specific pruning strategies where, although the number of tokens removed may differ across decoder layers, the overall pruning schedule is fixed and applied uniformly to all input samples and tasks, failing to align token elimination with the model’s holistic reasoning trajectory. Cognitive science indicates that human visual processing often begins with broad exploration to accumulate evidence before narrowing focus as the target becomes distinct. Our experiments reveal an analogous pattern in LVLMs. This observation strongly suggests that neither a fixed pruning schedule nor a heuristics layer-wise strategy can optimally accommodate the diverse complexities inherent in different inputs. To overcome this limitation, we introduce Complexity-Adaptive Pruning (AutoPrune), which is a training-free, plug-and-play framework that tailors pruning policies to varying sample and task complexities. Specifically, AutoPrune quantifies the mutual information between visual and textual tokens, and then projects this signal to a budget-constrained logistic retention curve. Each such logistic curve, defined by its unique shape, is shown to effectively correspond with the specific complexity of different tasks, and can easily guarantee adherence to a pre-defined computational constraints. We evaluate AutoPrune not only on standard vision-language tasks but also on Vision-Language-Action (VLA) models for autonomous driving. Notably, when applied to LLaVA-1.5-7B, our method prunes 89% of visual tokens and reduces inference FLOPs by 76.8%, but still retaining 96.7% of the original accuracy averaged over all tasks. This corresponds to a 9.1% improvement over the recent work PDrop (CVPR'2025), demonstrating the effectivenes. Code is available at https://github.com/AutoLab-SAI-SJTU/AutoPrune. Hanshi Wang, Yufan Liu 0001, Weiming Hu 0004 |
NeurIPS | 6 |
| 2025 | MI-TRQR: Mutual Information-Based Temporal Redundancy Quantification and Reduction for Energy-Efficient Spiking Neural NetworksabstractBrain-inspired spiking neural networks (SNNs) provide energy-efficient computation through event-driven processing. However, the shared weights across multiple timesteps lead to serious temporal feature redundancy, limiting both efficiency and performance. This issue is further aggravated when processing static images due to the duplicated input. To mitigate this problem, we propose a parameter-free and plug-and-play module named Mutual Information-based Temporal Redundancy Quantification and Reduction (MI-TRQR), constructing energy-efficient SNNs. Specifically, Mutual Information (MI) is properly introduced to quantify redundancy between discrete spike features at different timesteps on two spatial scales: pixel (local) and the entire spatial features (global). Based on the multi-scale redundancy quantification, we apply a probabilistic masking strategy to remove redundant spikes. The final representation is subsequently recalibrated to account for the spike removal. Extensive experimental results demonstrate that our MI-TRQR achieves sparser spiking firing, higher energy efficiency, and better performance concurrently with different SNN architectures in tasks of neuromorphic data classification, static data classification, and time-series forecasting. Notably, MI-TRQR increases accuracy by \textbf{1.7\%} on CIFAR10-DVS with 4 timesteps while reducing energy cost by \textbf{37.5\%}. Our codes are available at https://github.com/dfxue/MI-TRQR. Dengfeng Xue, Yifan Lu 0001, Chunfeng Yuan, Yufan Liu 0001, Wei Liu 0153, Man Yao, Li Yang 0014, Bing Li 0001, Stephen J. Maybank, Weiming Hu 0004, Zhetao Li |
NeurIPS | 12 |
| 2025 | An Experimental Study on Exploring Strong Lightweight Vision Transformers via Masked Image Modeling Pre-training
Shubo Lin, Shaoru Wang, Yutong Kou, Congxuan Zhang, Xiaoqin Zhang 0002, Yizheng Wang, Weiming Hu 0004 |
Int. J. Comput. Vis. | 10 |
| 2025 | Two-stream transformer tracking with messengers
Miaobo Qiu, Wenyang Luo, Tongfei Liu, Yanqin Jiang, Jiaming Yan, Weiming Hu 0004, Stephen J. Maybank |
Image Vis. Comput. | 8 |
| 2025 | FiGVCL: Fine-Grained Benchmark and Method for Video Copy LocalizationabstractContent-based video copy localization (VCL) aims to detect and locate copied segments in pairs of videos. VCL requires fine-grained video analysis to robustly identify copied segments that have been edited. Despite recent progress, the prohibitive cost of annotating copied segments and the lack of a fine-grained benchmark hinder the development of effective VCL systems. In this work, we annotate a new real-world dataset, FiGVCL, with challenging scenarios designed to evaluate VCL methods. FiGVCL is carefully annotated to preserve the temporal correspondences observed in copied segments. Moreover, we propose a novel fine-grained VCL benchmark metric based on temporal correspondences to improve discriminability. Finally, we design a simple but effective baseline model that uses fine-grained local embeddings for accurate copied segment localization. We also present an unsupervised training strategy that outperforms previous supervised VCL methods. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | iESTA: Instance-Enhanced Spatial-Temporal Alignment for Video Copy LocalizationabstractVideo copy Segment Localization (VSL) requires the identification of the temporal segments within a pair of videos that contain copied content. Current methods primarily focus on global temporal modeling, overlooking the complementarity of global semantic and local fine-grained features, which limits their effectiveness. Some related methods attempt to incorporate local spatial information but often disrupt spatial semantic structures, resulting in less accurate matching. To address these issues, we propose the Instance-Enhanced Spatial-Temporal Alignment Framework (iESTA), based on a proper representation granularity that integrates instance-level local features and semantic global features. Specifically, the Instance-relation Graph (IRG) is constructed to capture instance-level features and fine-grained interactions, preserving local information integrity and better representing the video feature space in a proper granularity. An instance-GNN structure is designed to refine these graph representations. For global features, we enhance the representation of semantic information, capturing temporal relationships within videos using a Transformer framework. Additionally, we design a Complementarity-perception Alignment Module (CAM) to effectively process and integrate complementary spatial-temporal information, producing accurate frame-to-frame alignment maps. Our approach also incorporates a differentiable Dynamic Time Warping (DTW) method to utilize latent temporal alignments as weak supervisory signals, improving the accuracy of the matching process. Experimental results indicate that our proposed iESTA outperforms state-of-the-art methods on both the small-scale dataset VCDB and the large-scale dataset VCSL. Xinmiao Ding, Jinming Lou, Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Task-Aware Attentional Dynamic Alignment for Few-Shot Compressed Video ClassificationabstractWe present a novel Task-aware Attentional Dynamic Alignment (TADA) framework for visual-based few-shot video classification (FSVC) that addresses two key challenges in this field: efficiency and nuanced spatio-temporal reasoning. Existing methods are often hindered by computationally expensive video decoding processes and neglect the temporal order of videos. In contrast, our method harnesses compressed domain data to extract rich spatio-temporal cues at a fraction of the cost of traditional video processing methods. Specifically, we propose an embedding module to extract informative features from compressed domain data while minimizing computational overheads. Furthermore, to exploit the temporal order of frames, we develop a prototypical ADA module to align and classify videos with an explicit temporal order constraint. Our framework also incorporates a contextual mixer to enrich video embeddings with task-specific context. Extensive experiments on multiple datasets demonstrate that TADA achieves state-of-the-art performance and outperforms existing methods in accuracy and efficiency. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | DABF-Net: A Dual-Branch Attention-Guided and Bi-Directional Feature Enhancement Network for Infrared Small-Target Detection With Air-to-Ground BenchmarkabstractInfrared small-target detection (IRSTD) is a critical, yet challenging task with significant applications in both military and civilian domains. Despite advancements in existing methods, two major limitations remain: the difficulty of achieving an optimal balance between detection probability and false alarm rate, and the lack of specialized datasets for air-to-ground scenarios. To address these limitations, this article presents a dual-pronged solution. At the algorithmic level, we propose a novel dual-branch attention-guided and bi-directional feature enhancement network (DABF-Net). First, we design a dual-branch high-low frequency attention (DHLA), which enhances the discriminability between the target and the background by preserving high-frequency edge features and modeling low-frequency contextual information in a complementary manner. Subsequently, we construct a bi-directional fusion module (BFM) to optimize multiscale feature compatibility while suppressing redundant information propagation. Furthermore, we introduce a small-target feature enhancement branch (STEB) employing space-to-depth (SPD) convolution and a feature integration module (FIM) to amplify latent target signatures through exponentially expanding receptive fields. At the data level, we contribute the NCHU-A2G-SIRST benchmark, the first comprehensive dataset specifically designed for the air-to-ground IRSTD task. The dataset contains four different scenes with two types of annotations, enabling evaluation and training of detection models in real-world conditions. Extensive experiments on several challenging datasets, including the self-built NCHU-A2G-SIRST dataset and three public datasets (NCHU-SIRST, NUAA-SIRST, and IRSTD-1K), demonstrate that the DABF-Net can outperform many state-of-the-art competing methods. The code and dataset are publicly available athttps://github.com/PCwenyue/DABF-Net Fagan Wang, Congxuan Zhang, Peng Liu 0024, Zhen Chen 0004, Weiming Hu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Self-Supervised Monocular Depth Estimation With Dual-Path Encoders and Offset Field InterpolationabstractAlthough self-supervised learning approaches have demonstrated tremendous potential in multi-frame depth estimation scenarios, existing methods struggle to perform well in cases involving dynamic targets and static ego-camera conditions. To address this issue, we propose a self-supervised monocular depth estimation method featuring dual-path encoders and learnable offset interpolation (LOI). First, we construct a dual-path encoding scheme that utilizes residual and transformer blocks to extract both single- and multi-frame features from the input frames. We design a contrastive learning strategy to effectively decouple single- and multi-frame features, enabling weighted fusion guided by a confidence map. Next, we explore two distinct decoding heads for simultaneously generating low-resolution predictions and offset fields. We then design an LOI module to directly upsample a low-resolution depth map to a full-resolution map. This one-step decoding framework enables accurate and efficient depth prediction. Finally, we evaluate our proposed method on the KITTI and Cityscapes benchmarks, conducting a comprehensive comparison with state-of-the-art approaches. The experimental results demonstrate that our DualDepth method achieves competitive performance in terms of both estimation accuracy and efficiency. Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge |
IEEE Trans. Image Process. | 4 |
| 2025 | Content-Decoupled Contrastive Learning-Based Implicit Degradation Modeling for Blind Image Super-ResolutionabstractImplicit degradation modeling-based blind super-resolution (SR) has attracted more increasing attention in the community due to its excellent generalization to complex degradation scenarios and wide application range. How to extract more discriminative degradation representations and fully adapt them to specific image features is the key to this task. In this paper, we propose a new Content-decoupled Contrastive Learning-based blind image super-resolution (CdCL) framework following the typical blind SR pipeline. This framework introduces negative-free contrastive learning technique for the first time to model the implicit degradation representation, in which a new cyclic shift sampling strategy is designed to ensure decoupling between content features and degradation features from the data perspective, thereby improving the purity and discriminability of the learned implicit degradation space. In addition, we propose a detail-aware implicit degradation adapting module that can better adapt degradation representations to specific LR features by enhancing the basic adaptation unit's perception of image details, significantly reducing the overall SR model complexity. Extensive experiments on synthetic and real data show that our method achieves highly competitive quantitative and qualitative results in various degradation settings while obviously reducing parameters and computational costs, validating the feasibility of designing practical and lightweight blind SR tools. Codes and models will be available at https://github.com/Fieldhunter/CdCL. Jiang Yuan, Bo Wang 0147, Weiming Hu 0004 |
IEEE Trans. Image Process. | 4 |
| 2024 | Set Prediction Guided by Semantic Concepts for Diverse Video CaptioningabstractDiverse video captioning aims to generate a set of sentences to describe the given video in various aspects. Mainstream methods are trained with independent pairs of a video and a caption from its ground-truth set without exploiting the intra-set relationship, resulting in low diversity of generated captions. Different from them, we formulate diverse captioning into a semantic-concept-guided set prediction (SCG-SP) problem by fitting the predicted caption set to the ground-truth set, where the set-level relationship is fully captured. Specifically, our set prediction consists of two synergistic tasks, i.e., caption generation and an auxiliary task of concept combination prediction providing extra semantic supervision. Each caption in the set is attached to a concept combination indicating the primary semantic content of the caption and facilitating element alignment in set prediction. Furthermore, we apply a diversity regularization term on concepts to encourage the model to generate semantically diverse captions with various concept combinations. These two tasks share multiple semantics-specific encodings as input, which are obtained by iterative interaction between visual features and conceptual queries. The correspondence between the generated captions and specific concept combinations further guarantees the interpretability of our model. Extensive experiments on benchmark datasets show that the proposed SCG-SP achieves state-of-the-art (SOTA) performance under both relevance and diversity metrics. Yifan Lu 0001, Ziqi Zhang 0010, Chunfeng Yuan, Yan Wang 0153, Bing Li 0001, Weiming Hu 0004 |
AAAI | 7 |
| 2024 | Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-trainingabstractIn vision-language pre-training (VLP), masked image modeling (MIM) has recently been introduced for fine-grained cross-modal alignment. However, in most existing methods, the reconstruction targets for MIM lack high-level semantics, and text is not sufficiently involved in masked modeling. These two drawbacks limit the effect of MIM in facilitating cross-modal semantic alignment. In this work, we propose a semantics-enhanced cross-modal MIM framework (SemMIM) for vision-language representation learning. Specifically, to provide more semantically meaningful supervision for MIM, we propose a local semantics enhancing approach, which harvest high-level semantics from global image features via self-supervised agreement learning and transfer them to local patch encodings by sharing the encoding space. Moreover, to achieve deep involvement of text during the entire MIM process, we propose a text-guided masking strategy and devise an efficient way of injecting textual information in both masked modeling and reconstruction target acquisition. Experimental results validate that our method improves the effectiveness of the MIM task in facilitating cross-modal semantic alignment. Compared to previous VLP models with similar model size and data scale, our SemMIM model achieves state-of-the-art or competitive performance on multiple downstream vision-language tasks. Yaya Shi, Haiyang Xu 0001, Chunfeng Yuan, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Bing Li 0001, Weiming Hu 0004 |
LREC/COLING | 11 |
| 2024 | Unifying Latent and Lexicon Representations for Effective Video-Text RetrievalabstractIn video-text retrieval, most existing methods adopt the dual-encoder architecture for fast retrieval, which employs two individual encoders to extract global latent representations for videos and texts. However, they face challenges in capturing fine-grained semantic concepts. In this work, we propose the UNIFY framework, which learns lexicon representations to capture fine-grained semantics and combines the strengths of latent and lexicon representations for video-text retrieval. Specifically, we map videos and texts into a pre-defined lexicon space, where each dimension corresponds to a semantic concept. A two-stage semantics grounding approach is proposed to activate semantically relevant dimensions and suppress irrelevant dimensions. The learned lexicon representations can thus reflect fine-grained semantics of videos and texts. Furthermore, to leverage the complementarity between latent and lexicon representations, we propose a unified learning scheme to facilitate mutual learning via structure sharing and self-distillation. Experimental results show our UNIFY framework largely outperforms previous video-text retrieval methods, with 4.8% and 8.2% Recall@1 improvement on MSR-VTT and DiDeMo respectively. Yaya Shi, Haiyang Xu 0001, Chunfeng Yuan, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Bing Li 0001, Weiming Hu 0004 |
LREC/COLING | 11 |
| 2024 | How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?abstractDominant dual-encoder models enable efficient image-text retrieval but suffer from limited accuracy, while the cross-encoder models offer higher accuracy at the expense of efficiency. Distilling cross-modality matching knowledge from cross-encoder to dual-encoder provides a natural approach to harness their strengths. Thus, we investigate the following valuable question: how to make cross-encoder a good teacher for dual-encoder? Our findings are threefold: (1) Cross-modal similarity score distribution of cross-encoder is more concentrated, while the result of dual-encoder is nearly normal, making vanilla logit distillation less effective. However, ranking distillation remains practical, as it is not affected by the score distribution. (2) Only the relative order between hard negatives conveys valid knowledge, while the order information between easy negatives has little significance. (3) Maintaining the coordination between distillation loss and dual-encoder training loss is beneficial for knowledge transfer. Based on these findings, we propose a novel Contrastive Partial Ranking Distillation (CPRD) method, which implements the objective of mimicking relative order between hard negative samples with contrastive learning. This approach coordinates with the training of the dual-encoder, effectively transferring valid knowledge from the cross-encoder to the dual-encoder. Extensive experiments on image-text retrieval and ranking tasks show that our method surpasses other distillation methods and significantly improves the accuracy of dual-encoder. Zongyang Ma, Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Bing Li 0001, Junfu Pu, Ying Shan, Xiaojuan Qi 0001, Weiming Hu 0004 |
CVPR | 10 |
| 2024 | A-Teacher: Asymmetric Network for 3D Semi-Supervised Object DetectionabstractThis work proposes the first online asymmetric semi-supervised framework, namely A-Teacher, for LiDAR -based 3D object detection. Our motivation stems from the observation that 1) existing symmetric teacher-student methods for semi-supervised 3D object detection have characterized simplicity, but impede the distillation performance between teacher and student because of the demand for an identical model structure and input data format. 2) The offline asymmetric methods with a complex teacher model, constructed differently, can generate more precise pseudo labels, but is challenging to jointly optimize the teacher and student model. Consequently, in this paper, we devise a different path from the conventional paradigm, which can harness the capacity of a strong teacher while preserving the advantages of jointly updating the whole framework. The essence is the proposed attention-based refinement model that can be seamlessly integrated into a vanilla teacher. The refinement model works in the divide-and-conquer manner that respectively handles three challenging scenarios including 1) objects detected in the current timestamp but with sub-optimal box quality, 2) objects are missed in the current timestamp but are detected in supporting frames, 3) objects are neglected in all frames. It is worth noting that even while tackling these complex cases, our model retains the efficiency of the online semi-supervised framework. Experimental results on Waymo [38] show that our method out-performs previous state-of-the-art HSSDA [17] for 4.7 on mAP (L1) while consuming fewer training resources. Hanshi Wang, Weiming Hu 0004 |
CVPR | 4 |
| 2024 | PromptIQA: Boosting the Performance and Generalization for No-Reference Image Quality Assessment via Prompts
Zewen Chen, Haina Qin, Juan Wang 0012, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Liang Wang 0001 |
ECCV (1) | 6 |
| 2024 | EA-VTR: Event-Aware Video-Text Retrieval
Zongyang Ma, Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Bing Li 0001, Yingmin Luo, Xu Li 0015, Xiaojuan Qi 0001, Ying Shan, Weiming Hu 0004 |
ECCV (52) | 11 |
| 2024 | STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians
Yifei Zeng, Yanqin Jiang, Siyu Zhu 0001, Yuanxun Lu, Youtian Lin, Hao Zhu 0004, Weiming Hu 0004, Xun Cao, Yao Yao 0008 |
ECCV (36) | 7 |
| 2024 | MIBench: Evaluating Multimodal Large Language Models over Multiple ImagesabstractHaowei Liu, Xi Zhang, Haiyang Xu, Yaya Shi, Chaoya Jiang, Ming Yan, Ji Zhang, Fei Huang, Chunfeng Yuan, Bing Li, Weiming Hu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Haiyang Xu 0001, Yaya Shi, Chaoya Jiang, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004 |
EMNLP | 11 |
| 2024 | Real-Time Monocular Depth Estimation on Embedded SystemsabstractDepth sensing is of paramount importance for unmanned aerial and autonomous vehicles. Nonetheless, contemporary monocular depth estimation methods employing complex deep neural networks within Convolutional Neural Networks are inadequately expedient for real-time inference on embedded platforms. This paper endeavors to surmount this challenge by proposing two efficient and lightweight architectures, RT-MonoDepth and RT-MonoDepth-S, thereby mitigating computational complexity and latency. Our methodologies not only attain accuracy comparable to prior depth estimation methods but also yield faster inference speeds. Specifically, RT-MonoDepth and RT-MonoDepth-S achieve frame rates of 18.4&30.5 FPS on NVIDIA Jetson Nano and 253.0&364.1 FPS on Jetson AGX Orin, utilizing a single RGB image of resolution $640 \times 192$. The experimental results underscore the superior accuracy and faster inference speed of our methods in comparison to existing fast monocular depth estimation methodologies on the KITTI dataset. Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Liyue Ge |
ICIP | 4 |
| 2024 | Consistent4D: Consistent 360° Dynamic Object Generation from Monocular VideoabstractIn this paper, we present Consistent4D, a novel approach for generating 4D dynamic objects from uncalibrated monocular videos. Uniquely, we cast the 360-degree dynamic object reconstruction as a 4D generation problem, eliminating the need for tedious multi-view data collection and camera calibration. This is achieved by leveraging the object-level 3D-aware image diffusion model as the primary supervision signal for training dynamic Neural Radiance Fields (DyNeRF). Specifically, we propose a cascade DyNeRF to facilitate stable convergence and temporal continuity under the time-discrete supervision signal. To achieve spatial and temporal consistency of the 4D generation, an interpolation-driven consistency loss is further introduced, which aligns the rendered frames with the interpolated frames from a pre-trained video interpolation model. Extensive experiments show that the proposed Consistent4D significantly outperforms previous 4D reconstruction approaches as well as per-frame 3D generation approaches, opening up new possibilities for 4D dynamic object generation from a single-view uncalibrated video. Project page: https://consistent4d.github.io Yanqin Jiang, Li Zhang 0040, Weiming Hu 0004, Yao Yao 0008 |
ICLR | 4 |
| 2024 | Learn from Noise: Detecting Deepfakes via Regional Noise ConsistencyabstractFace forgery detection becomes increasingly crucial due to the serious security issues caused by face manipulation techniques. Various methods primarily concentrate on the features specific to certain generation techniques, potentially resulting in overfitting to the distinctive fingerprint characteristics of those manipulation techniques, thus undermining their generalizability. In contrast, our investigation reveals a prevalent phenomenon wherein regional noise consistency is disrupted during the integration of synthesized faces into source images, regardless of specific manipulation techniques. Motivated by this observation, we introduce the Regional Noise Consistency Learning Framework (RNCL), a novel approach designed to discern manipulated faces. Central to RNCL are two pivotal modules: Noise Consistency Enhancement (NCE) and Pyramidal Noise Consistency Learning (PNCL). The NCE module facilitates channel-wise and spatial-wise feature enhancement by exploiting noise inconsistencies between facial and non-facial regions. Complementarily, the PNCL module constructs a noise consistency pyramid to analyze enhanced features across multiple scales, enabling adaptive multi-scale feature integration. Leveraging the NCE and PNCL modules, our framework effectively transforms noise information into useful forgery cues, significantly enhancing forgery detection performance. Experimental results demonstrate that our method achieves state-of-the-art performance on standard benchmarks. The code will be publicly available. Weiming Bai, Yufan Liu 0001, Bo Wang 0147, Chengwei Peng, Weiming Hu 0004, Bing Li 0001 |
IJCNN | 7 |
| 2024 | BEV2PR: BEV-Enhanced Visual Place Recognition with Structural CuesabstractIn this paper, we propose a new image-based visual place recognition (VPR) framework by exploiting the structural cues in bird’s-eye view (BEV) from a single monocular camera. The motivation arises from two key observations about place recognition methods based on both appearance and structure: 1) For the methods relying on LiDAR sensors, the integration of LiDAR in robotic systems has led to increased expenses, while the alignment of data between different sensors is also a major challenge. 2) Other image-/camera-based methods, involving integrating RGB images and their derived variants (e.g., pseudo depth images, pseudo 3D point clouds), exhibit several limitations, such as the failure to effectively exploit the explicit spatial relationships between different objects. To tackle the above issues, we design a new BEV-enhanced VPR framework, namely BEV2PR, generating a composite descriptor with both visual cues and spatial awareness based on a single camera. The key points lie in: 1) We use BEV features as an explicit source of structural knowledge in constructing global features. 2) The lower layers of the pretrained backbone from BEV generation are shared for visual and structural streams in VPR, facilitating the learning of fine-grained local features in the visual stream. 3) The complementary visual and structural features can jointly enhance VPR performance. Our BEV2PR framework enables consistent performance improvements over several popular aggregation modules for RGB global features. The experiments on our collected VPR-NuScenes dataset demonstrate an absolute gain of 2.47% on Recall@1 for the strong Conv-AP baseline to achieve the best performance in our setting, and notably, a 18.06% gain on the hard set. The code and dataset will be available at https://github.com/FudongGe/BEV2PR. Fudong Ge, Shuhan Shen, Weiming Hu 0004 |
IROS | 4 |
| 2024 | NFT1000: A Cross-Modal Dataset For Non-Fungible Token RetrievalabstractWith the rise of "Metaverse" and "Web 3.0", Non-Fungible Token (NFT) has emerged as a kind of pivotal digital asset, garnering significant attention. By the end of March 2024, more than 1.7 billion NFTs have been minted across various blockchain platforms. To effectively locate a desired NFT, conducting searches within a vast array of NFTs is essential. The challenge in NFT retrieval is heightened due to the high degree of similarity among different NFTs, regarding regional and semantic aspects. In this paper, we will introduce a benchmark dataset named "NFT Top1000 Visual-Text Dataset" (NFT1000), containing 7.56 million image-text pairs, and being collected from 1000 most famous PFP1 NFT collections2 by sales volume on the Ethereum blockchain. Based on this dataset and leveraging the CLIP series of pre-trained models as our foundation, we propose the dynamic masking fine-tuning scheme. This innovative approach results in a 7.4\% improvement in the top1 accuracy rate, while utilizing merely 13\% of the total training data (0.79 million vs. 6.1 million). We also propose a robust metric Comprehensive Variance Index (CVI) to assess the similarity and retrieval difficulty of visual-text pairs data. The dataset will be released as an open-source resource. For more details, please refer to: https://github.com/ShuxunoO/NFT-Net.git. Shuxun Wang, Yunfei Lei, Ziqi Zhang 0010, Wei Liu 0153, Li Yang 0014, Bing Li 0001, Weiming Hu 0004 |
ACM Multimedia | 10 |
| 2024 | Animate3D: Animating Any 3D Model with Multi-view Video DiffusionabstractRecent advances in 4D generation mainly focus on generating 4D content by distilling pre-trained text or single-view image conditioned models. It is inconvenient for them to take advantage of various off-the-shelf 3D assets with multi-view attributes, and their results suffer from spatiotemporal inconsistency owing to the inherent ambiguity in the supervision signals. In this work, we present Animate3D, a novel framework for animating any static 3D model. The core idea is two-fold: 1) We propose a novel multi-view video diffusion model (MV-VDM) conditioned on multi-view renderings of the static 3D object, which is trained on our presented large-scale multi-view video dataset (MV-Video). 2) Based on MV-VDM, we introduce a framework combining reconstruction and 4D Score Distillation Sampling (4D-SDS) to leverage the multi-view video diffusion priors for animating 3D objects. Specifically, for MV-VDM, we design a new spatiotemporal attention module to enhance spatial and temporal consistency by integrating 3D and video diffusion models. Additionally, we leverage the static 3D model’s multi-view renderings as conditions to preserve its identity. For animating 3D models, an effective two-stage pipeline is proposed: we first reconstruct coarse motions directly from generated multi-view videos, followed by the introduced 4D-SDS to model fine-level motions. Benefiting from accurate motion learning, we could achieve straightforward mesh animation. Qualitative and quantitative experiments demonstrate that Animate3D significantly outperforms previous approaches. Data, code, and models are open-released. Yanqin Jiang, Chaohui Yu, Chenjie Cao, Fan Wang 0019, Weiming Hu 0004 |
NeurIPS | 5 |
| 2024 | VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector QuantizationabstractBird's-eye-view (BEV) map layout estimation requires an accurate and full understanding of the semantics for the environmental elements around the ego car to make the results coherent and realistic. Due to the challenges posed by occlusion, unfavourable imaging conditions and low resolution, \emph{generating} the BEV semantic maps corresponding to corrupted or invalid areas in the perspective view (PV) is appealing very recently. \emph{The question is how to align the PV features with the generative models to facilitate the map estimation}. In this paper, we propose to utilize a generative model similar to the Vector Quantized-Variational AutoEncoder (VQ-VAE) to acquire prior knowledge for the high-level BEV semantics in the tokenized discrete space. Thanks to the obtained BEV tokens accompanied with a codebook embedding encapsulating the semantics for different BEV elements in the groundtruth maps, we are able to directly align the sparse backbone image features with the obtained BEV tokens from the discrete representation learning based on a specialized token decoder module, and finally generate high-quality BEV maps with the BEV codebook embedding serving as a bridge between PV and BEV. We evaluate the BEV map layout estimation performance of our model, termed VQ-Map, on both the nuScenes and Argoverse benchmarks, achieving 62.2/47.6 mean IoU for surround-view/monocular evaluation on nuScenes, as well as 73.4 IoU for monocular evaluation on Argoverse, which all set a new record for this map layout estimation task. The code and models are available on \url{https://github.com/Z1zyw/VQ-Map}. Fudong Ge, Guan Luo, Bing Li 0001, Zhaoxiang Zhang 0001, Haibin Ling, Weiming Hu 0004 |
NeurIPS | 8 |
| 2024 | Cross-Architecture Knowledge Distillation
Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jingting Ding, Liang Li 0006, Stephen J. Maybank |
Int. J. Comput. Vis. | 4 |
| 2024 | Joint Learning of Audio-Visual Saliency Prediction and Sound Source Localization on Multi-face Videos
Minglang Qiao, Yufan Liu 0001, Mai Xu, Xin Deng 0002, Bing Li 0001, Weiming Hu 0004, Ali Borji |
Int. J. Comput. Vis. | 6 |
| 2024 | DCFNet: Discriminant Correlation Filters Network for Visual Tracking
Weiming Hu 0004, Qiang Wang 0051, Bing Li 0001, Stephen J. Maybank |
J. Comput. Sci. Technol. | 1 |
| 2024 | Recursive Least-Squares Estimator-Aided Online Learning for Visual TrackingabstractTracking visual objects from a single initial exemplar in the testing phase has been broadly cast as a one-/few-shot problem, i.e., one-shot learning for initial adaptation and few-shot learning for online adaptation. The recent few-shot online adaptation methods incorporate the prior knowledge from large amounts of annotated training data via complex meta-learning optimization in the offline phase. This helps the online deep trackers to achieve fast adaptation and reduce overfitting risk in tracking. In this paper, we propose a simple yet effective recursive least-squares estimator-aided online learning approach for few-shot online adaptation without requiring offline training. It allows an in-built memory retention mechanism for the model to remember the knowledge about the object seen before, and thus the seen data can be safely removed from training. This also bears certain similarities to the emerging continual learning field in preventing catastrophic forgetting. This mechanism enables us to unveil the power of modern online deep trackers without incurring too much extra computational cost. We evaluate our approach based on two networks in the online learning families for tracking, i.e., multi-layer perceptrons in RT-MDNet and convolutional neural networks in DiMP. The consistent improvements on several challenging tracking benchmarks demonstrate its effectiveness and efficiency. Yan Lu 0001, Xiaojuan Qi 0001, Yutong Kou, Bing Li 0001, Liang Li 0006, Weiming Hu 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | One-Stage Anchor-Free Online Multiple Target Tracking With Deformable Local Attention and Task-Aware PredictionabstractThe tracking-by-detection paradigm currently dominates multiple target tracking algorithms. It usually includes three tasks: target detection, appearance feature embedding, and data association. Carrying out these three tasks successively usually leads to lower tracking efficiency. In this paper, we propose a one-stage anchor-free multiple task learning framework which carries out target detection and appearance feature embedding in parallel to substantially increase the tracking speed. This framework simultaneously predicts a target detection and produces a feature embedding for each location, by sharing a pyramid of feature maps. We propose a deformable local attention module which utilizes the correlations between features at different locations within a target to obtain more discriminative features. We further propose a task-aware prediction module which utilizes deformable convolutions to select the most suitable locations for the different tasks. At the selected locations, classification of samples into foreground or background, appearance feature embedding, and target box regression are carried out. Two effective training strategies, regression range overlapping and sample reweighting, are proposed to reduce missed detections in dense scenes. Ambiguous samples whose identities are difficult to determine are effectively dealt with to obtain more accurate feature embedding of target appearance. An appearance-enhanced non-maximum suppression is proposed to reduce over-suppression of true targets in crowded scenes. Based on the one-stage anchor-free network with the deformable local attention module and the task-aware prediction module, we implement a new online multiple target tracker. Experimental results show that our tracker achieves a very fast speed while maintaining a high tracking accuracy. Weiming Hu 0004, Shaoru Wang, Zongwei Zhou, Yangxi Li, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Chinese Title Generation for Short Videos: Dataset, Metric and AlgorithmabstractPrevious work for video captioning aims to objectively describe the video content but the captions lack human interest and attractiveness, limiting its practical application scenarios. The intention of video title generation (video titling) is to produce attractive titles, but there is a lack of benchmarks. This work offers CREATE, the first large-scale Chinese shoRt vidEo retrievAl and Title gEneration dataset, to assist research and applications in video titling, video captioning, and video retrieval in Chinese. CREATE comprises a high-quality labeled 210 K dataset and two web-scale 3 M and 10 M pre-training datasets, covering 51 categories, 50K+ tags, 537K+ manually annotated titles and captions, and 10M+ short videos with original video information. This work presents ACTEr, a unique Attractiveness-Consensus-based Title Evaluation, to objectively evaluate the quality of video title generation. This metric measures the semantic correlation between the candidate (model-generated title) and references (manual-labeled titles) and introduces attractive consensus weights to assess the attractiveness and relevance of the video title. Accordingly, this work proposes a novel multi-modal ALignment WIth Generation model, ALWIG, as one strong baseline to aid future model development. With the help of a tag-driven video-text alignment module and a GPT-based generation module, this model achieves video titling, captioning, and retrieval simultaneously. We believe that the release of the CREATE dataset, ACTEr metric, and ALWIG model will encourage in-depth research on the analysis and creation of Chinese short videos. Ziqi Zhang 0010, Zongyang Ma, Chunfeng Yuan, Peijin Wang, Zhongang Qi, Chenglei Hao, Bing Li 0001, Ying Shan, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2024 | DARTScore: DuAl-Reconstruction Transformer for Video Captioning EvaluationabstractVideo captioning evaluation aims at assessing the semantic consistency between video and candidate text, which should include measurement from two aspects: faithfulness (whether the information conveyed by candidate is correct w.r.t. video) and comprehensiveness (whether the main video content is covered by candidate). However, previous approaches have difficulty in evaluating faithfulness and comprehensiveness due to heavy reliance on references or heterogeneous of visual and textual data. In this paper, we propose a vision-involved evaluation metric based on a novel DuAl-Reconstruction Transformer, named DARTScore. DARTScore formulates the caption evaluation task as a dual-reconstruction problem to evaluate both faithfulness and comprehensiveness explicitly. Since the word in a candidate is usually related to several frames, DARTScore adaptively collects relevant frames to reconstruct the word and computes the reconstruction accuracy as faithfulness to inherently reflect whether the word information is contained in the video. In the inversive way, DARTScore reconstructs each frame with relevant words to evaluate comprehensiveness. By integrating fine-grained bidirectional reconstruction accuracies, DARTScore drills into each word in candidate and each frame in video to fully evaluate the semantic consistency. Furthermore, we collect and annotate two Chinese datasets with a large domain gap, named CRAETE-EVAL and VATEX-ZH-EVAL, to systematically evaluate existing metrics and fill the blank of Chinese video captioning evaluation. Experimental results show that DARTScore achieves higher correlation with human judgments, has lower reference reliance, and generalizes well to data from different domains. Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Ying Shan, Bing Li 0001, Weiming Hu 0004, Xiaohu Qie |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | IterDepth: Iterative Residual Refinement for Outdoor Self-Supervised Multi-Frame Monocular Depth EstimationabstractSelf-supervised monocular depth estimation has been a challenging task in computer vision for a long time, and it relies on only monocular or stereo video for its supervision. To address the challenge, we propose a novel multi-frame monocular depth estimation method called IterDepth, which is based on an iterative residual refinement network. IterDepth extracts depth features from consecutive frames and computes a 3D cost volume measuring the difference between current and previous features transformed by PoseCNN (pose estimation convolutional neural network). We reformulate depth prediction as a residual learning problem, revamping the dominating depth regression to enable high-accuracy multi-frame monocular depth estimation. Specifically, we design a gated recurrent depth fusion unit that seamlessly blends depth features from the cost volume, image features, and the depth prediction. The unit updates the hidden states and refines the depth map through iterative refinement, achieving more accurate predictions than existing methods. Our experiments on the KITTI dataset demonstrate that IterDepth is$7\times $faster in terms of FPS (frames per second) than the recent state-of-the-art DepthFormer model with competitive performance. We also test IterDepth on the Cityscapes dataset to showcase its generalization capability in other real-world environments. Moreover, IterDepth can balance accuracy and computational efficiency by adjusting the number of refinement iterations and performs competitively with other CNN-based monocular depth estimation approaches. Source code is available athttps://github.com/PCwenyue/IterDepth-TCSVT. Zhen Chen 0004, Congxuan Zhang, Weiming Hu 0004, Bing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | ACR-Net: Learning High-Accuracy Optical Flow via Adaptive-Aware Correlation Recurrent NetworkabstractAlthough recurrent network-based optical flow estimation methods have shown great success in recent years, most of these methods have difficulty handling large displacements and occlusions because the existing recurrent networks are usually restricted to coarse-resolution single-scale models while ignoring the multiscale features brought by hierarchical concepts in previous coarse-to-fine approaches. In this paper, we propose an adaptive-aware correlation recurrent network for optical flow estimation, named ACR-Net, which preserves fine motion features with a single-scale resolution recurrent framework and adaptively incorporates multiscale features at different stages to achieve high-accuracy optical flow estimation. First, our proposed self-adaptation scale-aware correlation module can incorporate the adaptive correlation of multiscale inter- and intra-motion features, which makes the features more discriminative for capturing long-range dependencies between pixels. Second, our presented adaptive-aware motion module can effectively extract the required features of different kinds of motion from multilevel correspondence. Third, our introduced cross-guide motion and fusion modules can accurately guide the propagation of reliable pixels towards unreliable pixels and dynamically determine the most suitable expression to address the occlusion challenges. Comprehensive experiments demonstrate that ACR-Net outperforms existing two-view models, striking a good balance between speed and accuracy and achieving the best performance on the MPI-Sintel final pass and KITTI-2015 test datasets. The code will be made publicly available. Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ke Lu 0002, Liyue Ge, Zige Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | PolarFormer: Multi-Camera 3D Object Detection with Polar Transformerabstract3D object detection in autonomous driving aims to reason “what” and “where” the objects of interest present in a 3D world. Following the conventional wisdom of previous 2D object detection, existing methods often adopt the canonical Cartesian coordinate system with perpendicular axis. However, we conjugate that this does not fit the nature of the ego car’s perspective, as each onboard camera perceives the world in shape of wedge intrinsic to the imaging geometry with radical (non perpendicular) axis. Hence, in this paper we advocate the exploitation of the Polar coordinate system and propose a new Polar Transformer (PolarFormer) for more accurate 3D object detection in the bird’s-eye-view (BEV) taking as input only multi-camera 2D images. Specifically, we design a cross-attention based Polar detection head without restriction to the shape of input structure to deal with irregular Polar grids. For tackling the unconstrained object scale variations along Polar’s distance dimension, we further introduce a multi-scale Polar representation learning strategy. As a result, our model can make best use of the Polar representation rasterized via attending to the corresponding image observation in a sequence-to-sequence fashion subject to the geometric constraints. Thorough experiments on the nuScenes dataset demonstrate that our PolarFormer outperforms significantly state-of-the-art 3D object detection alternatives. Yanqin Jiang, Li Zhang 0040, Zhenwei Miao, Xiatian Zhu, Weiming Hu 0004, Yu-Gang Jiang 0001 |
AAAI | 6 |
| 2023 | AUNet: Learning Relations Between Action Units for Face Forgery DetectionabstractFace forgery detection becomes increasingly crucial due to the serious security issues caused by face manipulation techniques. Recent studies in deepfake detection have yielded promising results when the training and testing face forgeries are from the same domain. However, the problem remains challenging when one tries to generalize the detector to forgeries created by unseen methods during training. Observing that face manipulation may alter the relation between different facial action units (AU), we propose the Action-Units Relation Learning framework to improve the generality of forgery detection. In specific, it consists of the Action Units Relation Transformer (ART) and the Tampered AU Prediction (TAP). The ART constructs the relation between different AUs with AU-agnostic Branch and AU-specific Branch, which complement each other and work together to exploit forgery clues. In the Tampered AU Prediction, we tamper AU-related regions at the image level and develop challenging pseudo samples at the feature level. The model is then trained to predict the tampered AU regions with the generated location-specific supervision. Experimental results demonstrate that our method can achieve state-of-the-art performance in both the in-dataset and cross-dataset evaluations. Weiming Bai, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
CVPR | 5 |
| 2023 | ViLEM: Visual-Language Error Modeling for Image-Text RetrievalabstractDominant pre-training works for image-text retrieval adopt “dual-encoder” architecture to enable high efficiency, where two encoders are used to extract image and text representations and contrastive learning is employed for global alignment. However, coarse-grained global alignment ignores detailed semantic associations between image and text. In this work, we propose a novel proxy task, named Visual-Language Error Modeling (ViLEM), to inject detailed image-text association into “dual-encoder” model by “proofreading” each word in the text against the corresponding image. Specifically, we first edit the image-paired text to automatically generate diverse plausible negative texts with pre-trained language models. ViLEM then enforces the model to discriminate the correctness of each word in the plausible negative texts and further correct the wrong words via resorting to image information. Further-more, we propose a multi-granularity interaction framework to perform ViLEM via interacting text features with both global and local image features, which associates local text semantics with both high-level visual context and multi-level local visual information. Our method surpasses state-of-the-art “dual-encoder” methods by a large margin on the image-text retrieval task and significantly improves discriminativeness to local textual semantics. Our model can also generalize well to video-text retrieval. Zongyang Ma, Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Ying Shan, Bing Li 0001, Weiming Hu 0004, Xiaohu Qie |
CVPR | 8 |
| 2023 | Learning to Exploit the Sequence-Specific Prior Knowledge for Image Processing Pipelines OptimizationabstractThe hardware image signal processing (ISP) pipeline is the intermediate layer between the imaging sensor and the downstream application, processing the sensor signal into an RGB image. The ISP is less programmable and consists of a series of processing modules. Each processing module handles a subtask and contains a set of tunable hyperparameters. A large number of hyperparameters form a complex mapping with the ISP output. The industry typically relies on manual and time-consuming hyperparameter tuning by image experts, biased towards human perception. Recently, several automatic ISP hyperparameter optimization methods using downstream evaluation metrics come into sight. However, existing methods for ISP tuning treat the high-dimensional parameter space as a global space for optimization and prediction all at once without inducing the structure knowledge of ISP. To this end, we propose a sequential ISP hyperparameter prediction framework that utilizes the sequential relationship within ISP modules and the similarity among parameters to guide the model sequence process. We validate the proposed method on object detection, image segmentation, and image quality tasks. Haina Qin, Longfei Han, Weihua Xiong, Juan Wang 0012, Bing Li 0001, Weiming Hu 0004 |
CVPR | 7 |
| 2023 | Learning from the Raw Domain: Cross Modality Distillation for Compressed Video Action RecognitionabstractVideo action recognition is faced with the challenges of both huge computation burden and performance requirements. Using compressed domain data, which saves much decoding computation, is a possible solution. Unfortunately, existing compressed-domain-based (CD) methods fail to obtain high performance, compared with state-of-the-art (SOTA) raw-domain-based (RD) methods. In order to solve the problem, we propose a cross-modality knowledge distillation method to force the CD model to learn the knowledge from the RD model. In particular, spatial knowledge and temporal knowledge are first constructed to align feature space between the raw domain and the compressed domain. Then, an adaptively multi-path knowledge learning scheme is presented to help the CD model learn in a more efficient way. Experiments verify the effectiveness of the proposed method in large-scale and small-scale datasets. Yufan Liu 0001, Jiajiong Cao, Weiming Bai, Bing Li 0001, Weiming Hu 0004 |
ICASSP | 5 |
| 2023 | Order-Prompted Tag Sequence Generation for Video TaggingabstractVideo Tagging intends to infer multiple tags spanning relevant content for a given video. Typically, video tags are freely defined and uploaded by a variety of users, so they have two characteristics: abundant in quantity and disordered intra-video. It is difficult for the existing multilabel classification and generation methods to adapt directly to this task. This paper proposes a novel generative model, Order-Prompted Tag Sequence Generation (OP-TSG), according to the above characteristics. It regards video tagging as a tag sequence generation problem guided by sample-dependent order prompts. These prompts are semantically aligned with tags and enable to decouple tag generation order, making the model focus on modeling the tag dependencies. Moreover, the word-based generation strategy enables the model to generate novel tags. To verify the effectiveness and generalization of the proposed method, a Chinese video tagging benchmark CREATE-tagging, and an English image tagging benchmark Pexel-tagging are established. Extensive results show that OP-TSG is significantly superior to other methods, especially the results on rare tags improve by 3.3% and 3% over SOTA methods on CREATE-tagging and Pexel-tagging, and novel tags generated on CREATE-tagging exhibit a tag gain of 7.04%. Zongyang Ma, Ziqi Zhang 0010, Zhongang Qi, Yingmin Luo, Zekun Li 0006, Chunfeng Yuan, Bing Li 0001, Xiaohu Qie, Ying Shan, Weiming Hu 0004 |
ICCV | 11 |
| 2023 | A Closer Look at Self-Supervised Lightweight Vision TransformersabstractSelf-supervised learning on large-scale Vision Transformers (ViTs) as pre-training methods has achieved promising downstream performance. Yet, how much these pre-training paradigms promote lightweight ViTs' performance is considerably less studied. In this work, we develop and benchmark several self-supervised pre-training methods on image classification tasks and some downstream dense prediction tasks. We surprisingly find that if proper pre-training is adopted, even vanilla lightweight ViTs show comparable performance to previous SOTA networks with delicate architecture design. It breaks the recently popular conception that vanilla ViTs are not suitable for vision tasks in lightweight regimes. We also point out some defects of such pre-training, e.g., failing to benefit from large-scale pre-training data and showing inferior performance on data-insufficient downstream tasks. Furthermore, we analyze and clearly show the effect of such pre-training by analyzing the properties of the layer representation and attention maps for related models. Finally, based on the above analyses, a distillation strategy during pre-training is developed, which leads to further downstream performance improvement for MAE-based pre-training. Code is available at https://github.com/wangsr126/mae-lite. Shaoru Wang, Xiaoqin Zhang 0002, Weiming Hu 0004 |
ICML | 5 |
| 2023 | Learning Semantics-Grounded Vocabulary Representation for Video-Text RetrievalabstractPrevious dual-encoder pre-training methods for video-text retrieval employ contrastive learning for cross-modal alignment in a latent space. However, such learned latent spaces often result in modality gap problem [26]. In this paper, we introduce a novel SemVTR framework designed to learn semantics-grounded video-text representations in a vocabulary space, in which each dimension corresponds to a semantic concept represented by a word. The representation is obtained by grounding video and text into semantically-related dimensions with high activation values. As video-text pairs share grounded dimensions, their vocabulary representations are expected to cluster together and thus alleviate modality gap problem. So, the crux of our method lies in grounding video and text into vocabulary space. Specifically, we propose a Multi-Granularity Video Semantics Grounding approach and a Textual Semantics Preserving training strategy. The visualization illustrates that SemVTR obtains semantics-gronded vocabulary representation and also alleviates the modality gap problem. SemVTR significantly outperforms existing methods on four video-text retrieval benchmarks. Yaya Shi, Haiyang Xu 0001, Zongyang Ma, Qinghao Ye, Anwen Hu, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Zhengjun Zha |
ACM Multimedia | 12 |
| 2023 | ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual TrackingabstractRecently, the transformer has enabled the speed-oriented trackers to approach state-of-the-art (SOTA) performance with high-speed thanks to the smaller input size or the lighter feature extraction backbone, though they still substantially lag behind their corresponding performance-oriented versions. In this paper, we demonstrate that it is possible to narrow or even close this gap while achieving high tracking speed based on the smaller input size. To this end, we non-uniformly resize the cropped image to have a smaller input size while the resolution of the area where the target is more likely to appear is higher and vice versa. This enables us to solve the dilemma of attending to a larger visual field while retaining more raw information for the target despite a smaller input size. Our formulation for the non-uniform resizing can be efficiently solved through quadratic programming (QP) and naturally integrated into most of the crop-based local trackers. Comprehensive experiments on five challenging datasets based on two kinds of transformer trackers, \ie, OSTrack and TransT, demonstrate consistent improvements over them. In particular, applying our method to the speed-oriented version of OSTrack even outperforms its performance-oriented counterpart by 0.6\% AUC on TNL2K, while running 50\% faster and saving over 55\% MACs. Codes and models are available at https://github.com/Kou-99/ZoomTrack. Yutong Kou, Bing Li 0001, Gang Wang 0031, Weiming Hu 0004, Yizheng Wang, Liang Li 0006 |
NeurIPS | 5 |
| 2023 | Exploiting Contextual Objects and Relations for 3D Visual Groundingabstract3D visual grounding, the task of identifying visual objects in 3D scenes based on natural language inputs, plays a critical role in enabling machines to understand and engage with the real-world environment. However, this task is challenging due to the necessity to capture 3D contextual information to distinguish target objects from complex 3D scenes. The absence of annotations for contextual objects and relations further exacerbates the difficulties. In this paper, we propose a novel model, CORE-3DVG, to address these challenges by explicitly learning about contextual objects and relations. Our method accomplishes 3D visual grounding via three sequential modular networks, including a text-guided object detection network, a relation matching network, and a target identification network. During training, we introduce a pseudo-label self-generation strategy and a weakly-supervised method to facilitate the learning of contextual objects and relations, respectively. The proposed techniques allow the networks to focus more effectively on referred objects within 3D scenes by understanding their context better. We validate our model on the challenging Nr3D, Sr3D, and ScanRefer datasets and demonstrate state-of-the-art performance. Our code will be public at https://github.com/yangli18/CORE-3DVG. Li Yang 0014, Chunfeng Yuan, Ziqi Zhang 0010, Zhongang Qi, Wei Liu 0153, Ying Shan, Bing Li 0001, Weiping Yang, Yan Wang 0153, Weiming Hu 0004 |
NeurIPS | 12 |
| 2023 | Hierarchical Curriculum Learning for No-Reference Image Quality Assessment
Juan Wang 0012, Zewen Chen, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004 |
Int. J. Comput. Vis. | 6 |
| 2023 | SiamMask: A Framework for Fast Online Object Tracking and SegmentationabstractIn this article, we introduce SiamMask, a framework to perform both visual object tracking and video object segmentation, in real-time, with the same simple method. We improve the offline training procedure of popular fully-convolutional Siamese approaches by augmenting their losses with a binary segmentation task. Once the offline training is completed, SiamMask only requires a single bounding box for initialization and can simultaneously carry out visual object tracking and segmentation at high frame-rates. Moreover, we show that it is possible to extend the framework to handle multiple object tracking and segmentation by simply re-using the multi-task model in a cascaded fashion. Experimental results show that our approach has high processing efficiency, at around 55 frames per second. It yields real-time state-of-the art results on visual-object tracking benchmarks, while at the same time demonstrating competitive performance at a high speed for video object segmentation benchmarks. Weiming Hu 0004, Qiang Wang 0051, Li Zhang 0040, Luca Bertinetto, Philip Torr 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Ranking-Based Color Constancy With Limited Training SamplesabstractComputational color constancy is an important component of Image Signal Processors (ISP) for white balancing in many imaging devices. Recently, deep convolutional neural networks (CNN) have been introduced for color constancy. They achieve prominent performance improvements comparing with those statistics or shallow learning-based methods. However, the need for a large number of training samples, a high computational cost and a huge model size make CNN-based methods unsuitable for deployment on low-resource ISPs for real-time applications. In order to overcome these limitations and to achieve comparable performance to CNN-based methods, an efficient method is defined for selecting the optimal simple statistics-based method (SM) for each image. To this end, we propose a novel ranking-based color constancy method (RCC) that formulates the selection of the optimal SM method as a label ranking problem. RCC designs a specific ranking loss function, and uses a low rank constraint to control the model complexity and a grouped sparse constraint for feature selection. Finally, we apply the RCC model to predict the order of the candidate SM methods for a test image, and then estimate its illumination using the predicted optimal SM method (or fusing the results estimated by the top k SM methods). Comprehensive experiment results show that the proposed RCC outperforms nearly all the shallow learning-based methods and achieves comparable performance to (sometimes even better performance than) deep CNN-based methods with only 1/2000 of the model size and training time. RCC also shows good robustness to limited training samples and good generalization crossing cameras. Furthermore, to remove the dependence on the ground truth illumination, we extend RCC to obtain a novel ranking-based method without ground truth illumination (RCC_NO) that learns the ranking model using simple partial binary preference annotations provided by untrained annotators rather than experts. RCC_NO also achieves better performance than the SM methods and most shallow learning-based methods with low costs of sample collection and illumination measurement. Bing Li 0001, Haina Qin, Weihua Xiong, Yangxi Li, Songhe Feng, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Learning to Explore Distillability and Sparsability: A Joint Framework for Model CompressionabstractDeep learning shows excellent performance usually at the expense of heavy computation. Recently, model compression has become a popular way of reducing the computation. Compression can be achieved using knowledge distillation or filter pruning. Knowledge distillation improves the accuracy of a lightweight network, while filter pruning removes redundant architecture in a cumbersome network. They are two different ways of achieving model compression, but few methods simultaneously consider both of them. In this paper, we revisit model compression and define two attributes of a model: distillability and sparsability, which reflect how much useful knowledge can be distilled and how many pruned ratios can be obtained, respectively. Guided by our observations and considering both accuracy and model size, a dynamically distillability-and-sparsability learning framework (DDSL) is introduced for model compression. DDSL consists of teacher, student and dean. Knowledge is distilled from the teacher to guide the student. The dean controls the training process by dynamically adjusting the distillation supervision and the sparsity supervision in a meta-learning framework. An alternating direction method of multiplier (ADMM)-based knowledge distillation-with-pruning (KDP) joint optimization algorithm is proposed to train the model. Extensive experimental results show that DDSL outperforms 24 state-of-the-art methods, including both knowledge distillation and filter pruning methods. Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Self-Prior Guided Pixel Adversarial Networks for Blind Image InpaintingabstractBlind image inpainting involves two critical aspects, i.e., "where to inpaint" and "how to inpaint". Knowing "where to inpaint" can eliminate the interference arising from corrupted pixel values; a good "how to inpaint" strategy yields high-quality inpainted results robust to various corruptions. In existing methods, these two aspects usually lack explicit and separate consideration. This paper fully explores these two aspects and proposes a self-prior guided inpainting network (SIN). The self-priors are obtained by detecting semantic-discontinuous regions and by predicting global semantic structures of the input image. On the one hand, the self-priors are incorporated into the SIN, which enables the SIN to perceive valid context information from uncorrupted regions and to synthesize semantic-aware textures for corrupted regions. On the other hand, the self-priors are reformulated to provide a pixel-wise adversarial feedback and a high-level semantic structure feedback, which can promote the semantic continuity of inpainted images. Experimental results demonstrate that our method achieves state-of-the-art performance in metric scores and in visual quality. It has an advantage over many existing methods that assume "where to inpaint" is known in advance. Extensive experiments on a series of related image restoration tasks validate the effectiveness of our method in obtaining high-quality inpainting. Juan Wang 0012, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Multi-scale self-attention-based feature enhancement for detection of targets with small image sizes
Xingliang Hu, Bing Li 0001, Congxuan Zhang, Weiming Hu 0004 |
Pattern Recognit. Lett. | 5 |
| 2023 | Dynamic adjustment of hyperparameters for anchor-based detection of objects with large image size differences
Xinliang Hu, Da Teng, Bing Li 0001, Congxuan Zhang, Weiming Hu 0004 |
Pattern Recognit. Lett. | 6 |
| 2023 | Jointing Recurrent Across-Channel and Spatial Attention for Multi-Object Tracking With Block-Erasing Data AugmentationabstractAlthough deep-learning-based multi-object tracking (MOT) approaches have achieved remarkable performances in terms of accuracy and efficiency, the issue of object occlusions remains an open challenge for most one-shot MOT methods. To address the problem of object occlusions, in this paper we present a recurrent across-channel and spatial attention-based one-shot multi-object tracking method with block-erasing data augmentation. First, we construct a multiattention feature learning module, named RASFL, that combines recurrent across -channel attention with spatial attention. The RASFL extracts both the correlations of the feature channels and the differences of the spatial locations to improve the accuracy of the re-identification (Re-ID) task. Second, we adopt a block-erasing data augmentation strategy to handle object occlusions by using random pixel blocks to simulate occlusion cases during the network training process. This block-erasing data augmentation assists the network to be more robust under object occlusions. By integrating the proposed RASFL module and the block-erasing data augmentation strategy into a one-shot online MOT system, we build an accurate and robust MOT model called DcMOT. Finally, we run our method on the MOT16, MOT17 and MOT20 datasets to conduct a comprehensive comparison with some of the state-of-the-art MOT methods. The experimental results demonstrate that the proposed DcMOT model achieves a competitive performance in terms of both accuracy and efficiency; with especially good performances in the occlusion cases. Keyu Deng, Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Bing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | TranSkeleton: Hierarchical Spatial-Temporal Transformer for Skeleton-Based Action RecognitionabstractIn skeleton-based action recognition, it has been a dominant paradigm to extract motion features with temporal convolution and model spatial correlations with graph convolution. However, it’s difficult for temporal convolution to capture long-range dependencies effectively. Meanwhile, commonly used multi-branch graph convolution leads to high complexity. In this paper, we propose TranSkeleton, a powerful Transformer framework which neatly unifies the spatial and temporal modeling of skeleton sequences. For temporal modeling, we propose a novel partition-aggregation temporal Transformer. It works with hierarchical temporal partition and aggregation, and can capture both long-range dependencies and subtle temporal structures effectively. A difference-aware aggregation approach is designed to reduce information loss during temporal aggregation. For spatial modeling, we propose a topology-aware spatial Transformer which utilizes the prior information of human body topology to facilitate spatial correlation modeling. Extensive experiments on two challenging benchmark datasets demonstrate that TranSkeleton notably outperforms the state of the arts. Yongcheng Liu, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | A Robust Infrared Small Target Detection Method Jointing Multiple Information and Noise Prediction: Algorithm and BenchmarkabstractInfrared small target detection plays an important role in many military and civilian applications. Despite the great advances made by infrared small target detection studies in recent years, most of the existing methods have difficulty in balancing detection probabilities and false alarms. Moreover, there are only a few public datasets for infrared small targets, which limits the development of infrared small target detection research. To address the abovementioned issues, in this paper, we propose a robust infrared small target detection method that joins multiple pieces of information and noise predictions, named MINP-Net. Specifically, we first design a gradient and contextual information extraction module to extract multiscale features from an input infrared image. Second, we construct a noise prediction network to model the background noise. Third, we plan a regional positioning branch to provide a coarse target location to decrease the false alarm ratio. In addition, we build a new infrared small target detection benchmark to advance the research in this field, named the NCHU-Seg dataset. To the best of our knowledge, the NCHU-Seg dataset is the largest real-world scene dataset for evaluating infrared small target segmentation methods. For a comprehensive evaluation, we compare our method with some of the state-of-the-art methods on both the well-known NUAA-SIRST dataset and our NCHU-Seg dataset. The experimental results demonstrate that the proposed MINP-Net method performs better in terms of detection effectiveness and segmentation accuracy and effectively balances the detection probabilities and false alarms with complex backgrounds. (The code and dataset are available at https://github.com/PCwenyue.). Siqiang Meng, Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Learning Video-Text Aligned Representations for Video CaptioningabstractVideo captioning requires that the model has the abilities of video understanding, video-text alignment, and text generation. Due to the semantic gap between vision and language, conducting video-text alignment is a crucial step to reduce the semantic gap, which maps the representations from the visual to the language domain. However, the existing methods often overlook this step, so the decoder has to directly take the visual representations as input, which increases the decoder’s workload and limits its ability to generate semantically correct captions. In this paper, we propose a video-text alignment module with a retrieval unit and an alignment unit to learn video-text aligned representations for video captioning. Specifically, we firstly propose a retrieval unit to retrieve sentences as additional input which is used as the semantic anchor between visual scene and language description. Then, we employ an alignment unit with the input of the video and retrieved sentences to conduct the video-text alignment. The representations of two modal inputs are aligned in a shared semantic space. The obtained video-text aligned representations are used to generate semantically correct captions. Moreover, retrieved sentences provide rich semantic concepts which are helpful for generating distinctive captions. Experiments on two public benchmarks, i.e., VATEX and MSR-VTT, demonstrate that our method outperforms state-of-the-art performances by a large margin. The qualitative analysis shows that our method generates correct and distinctive captions. Yaya Shi, Haiyang Xu 0001, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Zhengjun Zha |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | One More Check: Making "Fake Background" Be Tracked AgainabstractThe one-shot multi-object tracking, which integrates object detection and ID embedding extraction into a unified network, has achieved groundbreaking results in recent years. However, current one-shot trackers solely rely on single-frame detections to predict candidate bounding boxes, which may be unreliable when facing disastrous visual degradation, e.g., motion blur, occlusions. Once a target bounding box is mistakenly classified as background by the detector, the temporal consistency of its corresponding tracklet will be no longer maintained. In this paper, we set out to restore the bounding boxes misclassified as ``fake background'' by proposing a re-check network. The re-check network innovatively expands the role of ID embedding from data association to motion forecasting by effectively propagating previous tracklets to the current frame with a small overhead. Note that the propagation results are yielded by an independent and efficient embedding search, preventing the model from over-relying on detection results. Eventually, it helps to reload the ``fake background'' and repair the broken tracklets. Building on a strong baseline CSTrack, we construct a new one-shot tracker and achieve favorable gains by 70.7 ➡ 76.4, 70.6 ➡ 76.3 MOTA on MOT16 and MOT17, respectively. It also reaches a new state-of-the-art MOTA and IDF1 performance. Code is released at https://github.com/JudasDie/SOTS. Bing Li 0001, Weiming Hu 0004 |
AAAI | 5 |
| 2022 | Teacher-Guided Learning for Blind Image Quality Assessment
Zewen Chen, Juan Wang 0012, Bing Li 0001, Chunfeng Yuan, Weihua Xiong, Weiming Hu 0004 |
ACCV (3) | 7 |
| 2022 | Cross-Architecture Knowledge Distillation
Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jingting Ding, Liang Li 0006 |
ACCV (5) | 4 |
| 2022 | Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge DistillationabstractOpen- vocabulary object detection aims to detect novel object categories beyond the training set. The advanced open- vocabulary two-stage detectors employ instance-level visual-to- visual knowledge distillation to align the visual space of the detector with the semantic space of the Pre-trained Visual-Language Model (PVLM). However, in the more efficient one-stage detector, the absence of class-agnostic object proposals hinders the knowledge distil-lation on unseen objects, leading to severe performance degradation. In this paper, we propose a hierarchical visual-language knowledge distillation method, i.e., Hi-erKD, for open-vocabulary one-stage detection. Specifi-cally, a global-level knowledge distillation is explored to transfer the knowledge of unseen categories from the PVLM to the detector. Moreover, we combine the proposed global-level knowledge distillation and the common instance-level knowledge distillation to learn the knowledge of seen and unseen categories simultaneously. Extensive experiments on MS-COCO show that our method significantly surpasses the previous best one-stage detector with 11.9% and 6.7% AP50 gains under the zero-shot detection and generalized zero-shot detection settings, and reduces the AP50performance gap from 14% to 7.3% compared to the best two-stage detector. Code will be released at this url11https://qithub.com/menqqiDyanqqe/HierKD. Zongyang Ma, Guan Luo, Liang Li 0006, Shaoru Wang, Congxuan Zhang, Weiming Hu 0004 |
CVPR | 8 |
| 2022 | EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding MatchingabstractCurrent metrics for video captioning are mostly based on the text-level comparison between reference and candidate captions. However, they have some insuperable drawbacks, e.g., they cannot handle videos without references, and they may result in biased evaluation due to the one-to-many nature of video-to-text and the neglect of visual relevance. From the human evaluator's viewpoint, a high-quality caption should be consistent with the provided video, but not necessarily be similar to the reference in literal or semantics. Inspired by human evaluation, we propose EMScore (Embedding Matching-based score), a novel reference-free metric for video captioning, which directly measures similarity between video and candidate captions. Benefiting from the recent development of large-scale pre-training models, we exploit a well pre-trained vision-language model to extract visual and linguistic embeddings for computing EMScore. Specifically, EMScore combines matching scores of both coarse-grained (video and caption) and fine-grained (frames and words) levels, which takes the overall understanding and detailed characteristics of the video into account. Furthermore, considering the potential information gain, EMScore can be flexibly extended to the conditions where human-labeled references are available. Last but not least, we collect VATEX-EVAL and ActivityNet-FOIl datasets to systematically evaluate the existing metrics. VATEX-EVAL experiments demonstrate that EMScore has higher human correlation and lower reference dependency. ActivityNet-FOIL experiment verifies that EMScore can effectively identify “hallucinating” captions. Code and datasets are available at https://github.com/shiyaya/emscore. Yaya Shi, Xu Yang 0001, Haiyang Xu 0001, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Zhengjun Zha |
CVPR | 6 |
| 2022 | Improving Visual Grounding with Visual-Linguistic Verification and Iterative ReasoningabstractVisual grounding is a task to locate the target indicated by a natural language expression. Existing methods extend the generic object detection framework to this problem. They base the visual grounding on the features from pre-generated proposals or anchors, and fuse these features with the text embeddings to locate the target mentioned by the text. However, modeling the visual features from these predefined locations may fail to fully exploit the visual context and attribute information in the text query, which limits their performance. In this paper, we propose a transformer-based framework for accurate visual grounding by establishing text-conditioned discriminative features and performing multi-stage cross-modal reasoning. Specifically, we develop a visual-linguistic verification module to focus the visual features on regions relevant to the textual descriptions while suppressing the unrelated areas. A language-guided feature encoder is also devised to aggregate the visual contexts of the target object to improve the object's distinctiveness. To retrieve the target from the encoded visual features, we further propose a multi-stage cross-modal decoder to iteratively speculate on the correlations between the image and text for accurate target localization. Extensive experiments on five widely used datasets validate the efficacy of our proposed components and demonstrate state-of-the-art performance. Li Yang 0014, Chunfeng Yuan, Wei Liu 0153, Bing Li 0001, Weiming Hu 0004 |
CVPR | 6 |
| 2022 | Attention-Aware Learning for Hyperparameter Prediction in Image Processing Pipelines
Haina Qin, Longfei Han, Juan Wang 0012, Congxuan Zhang, Bing Li 0001, Weiming Hu 0004 |
ECCV (19) | 7 |
| 2022 | Learnable Pixel Clustering Via Structure and Semantic Dual Constraints for Unsupervised Image SegmentationabstractUnsupervised image segmentation is a challenge task, since a high-quality segmented image should perceive not only local object structures but also certain semantics without any annotations. In this paper, we propose a novel encoder-decoder pixel clustering framework with dual constraints to incorporate local structure and global semantic information for guiding pixel feature learning in a self-supervised manner. On one hand, a Local Structure Constraint (LStC) is constructed based on fine-grained superpixels, which improves the boundary perception of pixel features by keeping intra-superpixel feature consistency and largening inter-superpixel feature distance. On the other hand, a new Global Semantic Constraint (GSeC) is proposed via adapting the mutual information maximization technique to the single-image setting, and it strengthens the global semantic perception of pixel features and thus improves the segmenting integrity of objects. Finally, based on the learned pixel features, a smoothing component is employed to achieve semantically meaningful pixel clustering. The experimental evaluation on BSDS500 and PASCAL Context datasets show the superiority of our method on region and boundary qualities. Bo Wang 0147, Shiang Wang, Chunfeng Yuan, Zhonghai Wu, Bing Li 0001, Weiming Hu 0004, Jeffrey Xiong |
ICIP | 6 |
| 2022 | Inter-Intra Cross-Modality Self-Supervised Video Representation Learning by Contrastive ClusteringabstractThis paper introduces an online self-supervised method that leverages inter- and intra-level variance for video representation learning. Most existing methods tend to focus on instance-level or inter-variance encoding but ignore the intra-variance existing in clips. The key observation to solving this problem is the underlying correlation between visual and audio, in which the distribution of flow patterns in feature space is diverse, but expresses complementary similar semantics. And in the semantic feature space, the horizontal dimension of the feature matrix could be regarded as cluster labels. These cluster labels should be consistent for different modalities of the same video clip. Based on this idea, we propose an end-to-end inter-intra cross-modality contrastive clustering scheme to simultaneously optimize the inter- and intra-level contrastive loss. Experiments show that our proposed approach is able to considerably outperform previous methods for self-supervised learning on HMDB51 and UCF101 when applied to video retrieval and action recognition tasks. Jiutong Wei, Guan Luo, Bing Li 0001, Weiming Hu 0004 |
ICPR | 4 |
| 2022 | Learning Target-aware Representation for Visual Tracking via Informative InteractionsabstractWe introduce a novel backbone architecture to improve target-perception ability of feature representation for tracking. Having observed de facto frameworks perform feature matching simply using the backbone outputs for target localization, there is no direct feedback from the matching module to the backbone network, especially the shallow layers. Concretely, only the matching module can directly access the target information, while the representation learning of candidate frame is blind to the reference target. Therefore, the accumulated target-irrelevant interference in shallow stages may degrade the feature quality of deeper layers. In this paper, we approach the problem by conducting multiple branch-wise interactions inside the Siamese-like backbone networks (InBN). The core of InBN is a general interaction modeler (GIM) that injects the target information to different stages of the backbone network, leading to better target-perception of candidate feature representation with negligible computation cost. The proposed GIM module and InBN mechanism are general and applicable to different backbone types including CNN and Transformer for improvements, as evidenced on multiple benchmarks. In particular, the CNN version improves the baseline with 3.2/6.9 absolute gains of SUC on LaSOT/TNL2K. The Transformer version obtains SUC of 65.7/52.0 on LaSOT/TNL2K, which are on par with recent SOTAs. Mingzhe Guo, Heng Fan 0001, Liping Jing, Yilin Lyu, Bing Li 0001, Weiming Hu 0004 |
IJCAI | 7 |
| 2022 | Long-Short Term Cross-Transformer in Compressed Domain for Few-Shot Video ClassificationabstractCompared with image few-shot learning, most of the existing few-shot video classification methods perform worse on feature matching, because they fail to sufficiently exploit the temporal information and relation. Specifically, frames are usually evenly sampled, which may miss important frames. On the other hand, the heuristic model simply encodes the equally treated frames in sequence, which results in the lack of both long-term and short-term temporal modeling and interaction. To alleviate these limitations, we take advantage of the compressed domain knowledge and propose a long-short term Cross-Transformer (LSTC) for few-shot video classification. For short terms, the motion vector (MV) contains temporal cues and reflects the importance of each frame. For long terms, a video can be natively divided into a sequence of GOPs (Group Of Picture). Using this compressed domain knowledge helps to obtain a more accurate spatial-temporal feature space. Consequently, we design the long-short term selection module, short-term module, and long-term module to comprise the LSTC. Long-short term selection is performed to select informative compressed domain data. Long/short-term modules are utilized to sufficiently exploit the temporal information so that the query and support can be well-matched by cross-attention. Experimental results show the superiority of our method on various datasets. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Yanan Miao, Yangxi Li |
IJCAI | 4 |
| 2022 | Interaction-Aware Spatio-Temporal Pyramid Attention Networks for Action ClassificationabstractFor CNN-based visual action recognition, the accuracy may be increased if local key action regions are focused on. The task of self-attention is to focus on key features and ignore irrelevant information. So, self-attention is useful for action recognition. However, current self-attention methods usually ignore correlations among local feature vectors at spatial positions in CNN feature maps. In this paper, we propose an effective interaction-aware self-attention model which can extract information about the interactions between feature vectors to learn attention maps. Since the different layers in a network capture feature maps at different scales, we introduce a spatial pyramid with the feature maps at different layers for attention modeling. The multi-scale information is utilized to obtain more accurate attention scores. These attention scores are used to weight the local feature vectors of the feature maps and then calculate attentional feature maps. Since the number of feature maps input to the spatial pyramid attention layer is unrestricted, we easily extend this attention layer to a spatio-temporal version. Our model can be embedded in any general CNN to form a video-level end-to-end attention network for action recognition. Several methods are investigated to combine the RGB and flow streams to obtain accurate predictions of human actions. Experimental results show that our method achieves state-of-the-art results on the datasets UCF101, HMDB51, Kinetics-400, and untrimmed Charades. Weiming Hu 0004, Chunfeng Yuan, Bing Li 0001, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Parallel multiscale context-based edge-preserving optical flow estimation with occlusion detection
Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Ming Li 0056 |
Signal Process. Image Commun. | 4 |
| 2022 | SDTP: Semantic-Aware Decoupled Transformer Pyramid for Dense Image PredictionabstractAlthough transformer has achieved great progress on computer vision tasks, the scale variation in dense image prediction is still the key challenge. Few effective multi-scale techniques are applied in transformer and there are two main limitations in the current methods. On the one hand, self-attention module in vanilla transformer fails to sufficiently exploit the diversity of semantic information because of its rigid mechanism. On the other hand, it is difficult to build attention and interaction among different levels due to the heavy computational burden. To alleviate this problem, we first revisit multi-scale problem in dense prediction, verifying the significance of diverse semantic representation and multi-scale interaction, and exploring the adaptation of transformer to pyramidal structure. Inspired by these findings, we propose a novel Semantic-aware Decoupled Transformer Pyramid (SDTP) for dense image prediction, consisting of Intra-level Semantic Promotion (ISP), Cross-level Decoupled Interaction (CDI) and Attention Refinement Function (ARF). ISP explores the semantic diversity in different receptive space through more flexible self-attention strategy. CDI builds the global attention and interaction among different levels in decoupled space which also solves the problem of heavy computation. Besides, ARF is further added to refine the attention in transformer. Experimental results demonstrate the validity and generality of the proposed method, which outperforms the state-of-the-art by a significant margin in dense image prediction tasks. Furthermore, the proposed components are all plug-and-play, which can be embedded in other methods. Zekun Li 0006, Yufan Liu 0001, Bing Li 0001, Bailan Feng, Kebin Wu, Chengwei Peng, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2022 | A Simple and Strong Baseline for Universal Targeted Attacks on Siamese Visual TrackingabstractSiamese trackers are shown to be vulnerable to adversarial attacks recently. However, the existing attack methods craft the perturbations for each video independently, which comes at a non-negligible computational cost. In this paper, we show the existence of universal perturbations that can enable the targeted attack, e.g., forcing a tracker to follow the ground-truth trajectory with specified offsets, to be video-agnostic and free from inference in a network. Specifically, we attack a tracker by adding a universal translucent perturbation to the template image and adding afake target, i.e., a small universal adversarial patch, into the search images adhering to the predefined trajectory, so that the tracker outputs the location and size of thefake targetinstead of the real target. Our approach allows perturbing a novel video to come at no additional cost except the mere addition operations – and not require gradient optimization or network inference. Experimental results on several datasets demonstrate that our approach can effectively fool the Siamese trackers in a targeted attack manner. We show that the proposed perturbations are not only universal across videos, but also generalize well across different trackers. Such perturbations are therefore doubly universal, both with respect to the data and the network architectures. Our code is available athttps://github.com/lizhenbang56/SiamAttack. Zhenbang Li, Yaya Shi, Shaoru Wang, Bing Li 0001, Pengpeng Liang, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2022 | MRDDANet: A Multiscale Residual Dense Dual Attention Network for SAR Image DenoisingabstractSynthetic aperture radar (SAR), due to its inherent characteristics, will produce speckle noise, which results in the deterioration of image quality, so the removal of speckle in SAR image is very important for the subsequent high-level image processing. In order to balance the relationship between denoising and texture preservation, we propose a multiscale residual dense dual attention network (MRDDANet) for SAR image denoising. This algorithm can effectively suppress the speckle while fully retaining the texture details of the image. In MRDDANet, shallow features are extracted from the noisy images by multiscale modules with different kernel sizes, and then, the extracted shallow features are mapped to the residual dense dual-attention network to obtain the deep features of SAR image. Finally, the final denoising image is generated through global residual learning. MRDDANet has advantages of both multiscale blocks and residual dense dual attention networks. The dense connection can fully extract features in the image, and the dual-channel attention enables MRDDANet to pay more attention to noise information, which is beneficial to remove noise and keep the details of the original image at the same time. Compared with state-of-the-art algorithms, the results of the experiment indicate that our method not only improves various objective indicators but also shows great advantages in visual effects. Shuaiqi Liu 0001, Luyao Zhang 0004, Bing Li 0001, Weiming Hu 0004, Yudong Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | SSAU-Net: A Spectral-Spatial Attention-Based U-Net for Hyperspectral Image FusionabstractCompared with traditional remoting image, there is a large amount of spectral information in the hyperspectral image (HSI), which makes HSI better reflect the actual condition of surface features. However, due to the limitations of imaging conditions, HSI tends to have a lower spatial resolution. In order to overcome this issue, we propose a spectral-spatial attention-based U-Net named SSAU-Net for HSI and multispectral image (MSI) fusion. The SSAU-Net constructs a spectral-spatial attention module by a coordinate-attention (CA) module and an efficient pyramid split attention (ESPA) module, which can enhance the image’s spectral information and spatial information. Meanwhile, the proposed network fully extracts the shallow and deep features of the images, and finally generates high-resolution (HR) hyperspectral images. Compared with state-of-the-art HSI-MSI fusion methods, the experimental results verify that the proposed method has a better subjective and objective fusion effect. Shuaiqi Liu 0001, Shichong Zhang, Bing Li 0001, Weiming Hu 0004, Yudong Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Rethinking the Competition Between Detection and ReID in Multiobject TrackingabstractDue to balanced accuracy and speed, one-shot models which jointly learn detection and identification embeddings, have drawn great attention in multi-object tracking (MOT). However, the inherent differences and relations between detection and re-identification (ReID) are unconsciously overlooked because of treating them as two isolated tasks in the one-shot tracking paradigm. This leads to inferior performance compared with existing two-stage methods. In this paper, we first dissect the reasoning process for these two tasks, which reveals that the competition between them inevitably would destroy task-dependent representations learning. To tackle this problem, we propose a novel reciprocal network (REN) with a self-relation and cross-relation design so that to impel each branch to better learn task-dependent representations. The proposed model aims to alleviate the deleterious tasks competition, meanwhile improve the cooperation between detection and ReID. Furthermore, we introduce a scale-aware attention network (SAAN) that prevents semantic level misalignment to improve the association capability of ID embeddings. By integrating the two delicately designed networks into a one-shot online MOT system, we construct a strong MOT tracker, namely CSTrack. Our tracker achieves the state-of-the-art performance on MOT16, MOT17 and MOT20 datasets, without other bells and whistles. Moreover, CSTrack is efficient and runs at 16.4 FPS on a single modern GPU, and its lightweight version even runs at 34.6 FPS. The complete code has been released at https://github.com/JudasDie/SOTS. Bing Li 0001, Shuyuan Zhu, Weiming Hu 0004 |
IEEE Trans. Image Process. | 6 |
| 2022 | Narrowing the Gap: Improved Detector Training With Noisy Location AnnotationsabstractDeep learning methods require massive of annotated data for optimizing parameters. For example, datasets attached with accurate bounding box annotations are essential for modern object detection tasks. However, labeling with such pixel-wise accuracy is laborious and time-consuming, and elaborate labeling procedures are indispensable for reducing man-made noise, involving annotation review and acceptance testing. In this paper, we focus on the impact of noisy location annotations on the performance of object detection approaches and aim to, on the user side, reduce the adverse effect of the noise. First, noticeable performance degradation is experimentally observed for both one-stage and two-stage detectors when noise is introduced to the bounding box annotations. For instance, our synthesized noise results in performance decrease from 38.9% AP to 33.6% AP for FCOS detector on COCO test split, and 37.8%AP to 33.7%AP for Faster R-CNN. Second, a self-correction technique based on a Bayesian filter for prediction ensemble is proposed to better exploit the noisy location annotations following a Teacher-Student learning paradigm. Experiments for both synthesized and real-world scenarios consistently demonstrate the effectiveness of our approach, e.g., our method increases the degraded performance of the FCOS detector from 33.6% AP to 35.6% AP on COCO. Shaoru Wang, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Image Process. | 4 |
| 2022 | PDNet: Toward Better One-Stage Object Detection With Prediction DecouplingabstractRecent one-stage object detectors follow a per-pixel prediction approach that predicts both the object category scores and boundary positions from every single grid location. However, the most suitable positions for inferring different targets, i.e., the object category and boundaries, are generally different. Predicting all these targets from the same grid location thus may lead to sub-optimal results. In this paper, we analyze the suitable inference positions for object category and boundaries, and propose a prediction-target-decoupled detector named PDNet to establish a more flexible detection paradigm. Our PDNet with the prediction decoupling mechanism encodes different targets separately in different locations. A learnable prediction collection module is devised with two sets of dynamic points, i.e., dynamic boundary points and semantic points, to collect and aggregate the predictions from the favorable regions for localization and classification. We adopt a two-step strategy to learn these dynamic point positions, where the prior positions are estimated for different targets first, and the network further predicts residual offsets to the positions with better perceptions of the object properties. Extensive experiments on the MS COCO benchmark demonstrate the effectiveness and efficiency of our method. With a single ResNeXt-64x4d-101-DCN as the backbone, our detector achieves 50.1 AP with single-scale testing, which outperforms the state-of-the-art methods by an appreciable margin under the same experimental settings. Moreover, our detector is highly efficient as a one-stage framework. Our code will be public. Li Yang 0014, Shaoru Wang, Chunfeng Yuan, Ziqi Zhang 0010, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Image Process. | 7 |
| 2022 | 3DCANN: A Spatio-Temporal Convolution Attention Neural Network for EEG Emotion RecognitionabstractSince electroencephalogram (EEG) signals can truly reflect human emotional state, emotion recognition based on EEG has turned into a critical branch in the field of artificial intelligence. Aiming at the disparity of EEG signals in various emotional states, we propose a new deep learning model named three-dimension convolution attention neural network (3DCANN) for EEG emotion recognition in this paper. The 3DCANN model is composed of spatio-temporal feature extraction module and EEG channel attention weight learning module, which can extract the dynamic relation well among multi-channel EEG signals and the internal spatial relation of multi-channel EEG signals during continuous period time. In this model, the spatio-temporal features are fused with the weights of dual attention learning, and the fused features are input into the softmax classifier for emotion classification. In addition, we utilize SJTU Emotion EEG Dataset (SEED) to appraise the feasibility and effectiveness of the proposed algorithm. Finally, experimental results display that the 3DCANN method has superior performance over the state-of-the-art models in EEG emotion recognition. Shuaiqi Liu 0001, Xu Wang 0029, Bing Li 0001, Weiming Hu 0004, Yudong Zhang 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Self-Attention-Based Multiscale Feature Learning Optical Flow With Occlusion Feature Map PredictionabstractEven though optical flow approaches based on convolutional neural networks have achieved remarkable performance with respect to both accuracy and efficiency, large displacements and motion occlusions remain challenges for most existing learning-based models. To address the abovementioned issues, we propose in this paper a self-attention-based multiscale feature learning optical flow computation method with occlusion feature map prediction. First, we exploit a self-attention mechanism-based multiscale feature learning module to compensate for large displacement optical flows, and the presented module is able to capture long-range dependencies from the input frames. Second, we design a simple but effective self-learning module to acquire an occlusion feature map, in which the predicted occlusion map is utilized to correct the optical flow estimation in occluded areas. Third, we explore a hybrid loss function that integrates the photometric and smoothness losses into the classical endpoint error (EPE)-based loss to ensure the accuracy and robustness of the presented network. Finally, we compare the proposed method with some state-of-the-art approaches using the MPI-Sintel and KITTI test databases. The experimental results demonstrate that the proposed method achieved competitive performance with respect to both accuracy and robustness, and it produced the better results compared to other methods under large displacements and motion occlusions. Congxuan Zhang, Zhongkai Zhou, Zhen Chen 0004, Weiming Hu 0004, Ming Li 0056, Shaofeng Jiang |
IEEE Trans. Multim. | 4 |
| 2021 | DPFPS: Dynamic and Progressive Filter Pruning for Compressing Convolutional Neural Networks from ScratchabstractFilter pruning is a commonly used method for compressing Convolutional Neural Networks (ConvNets), due to its friendly hardware supporting and flexibility. However, existing methods mostly need a cumbersome procedure, which brings many extra hyper-parameters and training epochs. This is because only using sparsity and pruning stages cannot obtain a satisfying performance. Besides, many works do not consider the difference of pruning ratio across different layers. To overcome these limitations, we propose a novel dynamic and progressive filter pruning (DPFPS) scheme that directly learns a structured sparsity network from Scratch. In particular, DPFPS imposes a new structured sparsity-inducing regularization specifically upon the expected pruning parameters in a dynamic sparsity manner. The dynamic sparsity scheme determines sparsity allocation ratios of different layers and a Taylor series based channel sensitivity criteria is presented to identify the expected pruning parameters. Moreover, we increase the structured sparsity-inducing penalty in a progressive manner. This helps the model to be sparse gradually instead of forcing the model to be sparse at the beginning. Our method solves the pruning ratio based optimization problem by an iterative soft-thresholding algorithm (ISTA) with dynamic sparsity. At the end of the training, we only need to remove the redundant parameters without other stages, such as fine-tuning. Extensive experimental results show that the proposed method is competitive with 11 state-of-the-art methods on both small-scale and large-scale datasets (i.e., CIFAR and ImageNet). Specifically, on ImageNet, we achieve a 44.97% pruning ratio of FLOPs by compressing ResNet-101, even with an increase of 0.12% Top-5 accuracy. Our pruned models and codes are released at https://github.com/taoxvzi/DPFPS. Xiaofeng Ruan, Yufan Liu 0001, Bing Li 0001, Chunfeng Yuan, Weiming Hu 0004 |
AAAI | 5 |
| 2021 | Open-Book Video Captioning With Retrieve-Copy-Generate NetworkabstractIn this paper, we convert traditional video captioning task into a new paradigm, i.e., Open-book Video Captioning, which generates natural language under the prompts of video-content-relevant sentences, not limited to the video itself. To address the open-book video captioning problem, we propose a novel Retrieve-Copy-Generate network, where a pluggable video-to-text retriever is constructed to retrieve sentences as hints from the training corpus effectively, and a copy-mechanism generator is introduced to extract expressions from multi-retrieved sentences dynamically. The two modules can be trained end-to-end or separately, which is flexible and extensible. Our framework co-ordinates the conventional retrieval-based methods with orthodox encoder-decoder methods, which can not only draw on the diverse expressions in the retrieved sentences but also generate natural and accurate content of the video. Extensive experiments on several benchmark datasets show that our proposed approach surpasses the state-of-the-art performance, indicating the effectiveness and promising of the proposed paradigm in the task of video captioning. Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Ying Shan, Bing Li 0001, Weiming Hu 0004 |
CVPR | 7 |
| 2021 | Practical Face Swapping Detection Based on Identity Spatial ConstraintsabstractThe generalization of face swapping detectors against unseen face manipulation methods is important to practical applications. Most existing methods based on convolutional neural networks (CNN) simply map the facial images to real/fake binary labels and achieve high performance on the known forgeries, but they almost fail to detect new manipulation methods. In order to improve the generalization of face swapping detection, this work concentrates on a practical scenario to protect specific persons by proposing a novel face swapping detector requiring a reference image. To this end, we design a new detection framework based on identity spatial constraints (DISC), which consists of a backbone network and an identity semantic encoder (ISE). When inspecting an image of a particular person, the ISE utilizes a real facial image of that person as the reference to constrain the backbone to focus on the identity-related facial areas, so as to exploit the intrinsic discriminative clues to the forgery in the query image. Cross-dataset evaluations on five large-scale face forgery datasets show that DISC significantly improves the performance against unseen manipulation methods and is robust against the distortions. Compared to the existing detection methods, the AUC scores achieve 10%~40% performance improvements. Bo Wang 0147, Bing Li 0001, Weiming Hu 0004 |
IJCB | 4 |
| 2021 | Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionabstractGraph convolutional networks (GCNs) have been widely used and achieved remarkable results in skeleton-based action recognition. In GCNs, graph topology dominates feature aggregation and therefore is the key to extracting representative features. In this work, we propose a novel Channel-wise Topology Refinement Graph Convolution (CTR-GC) to dynamically learn different topologies and effectively aggregate joint features in different channels for skeleton-based action recognition. The proposed CTR-GC models channel-wise topologies through learning a shared topology as a generic prior for all channels and refining it with channel-specific correlations for each channel. Our refinement method introduces few extra parameters and significantly reduces the difficulty of modeling channel-wise topologies. Furthermore, via reformulating graph convolutions into a unified form, we find that CTR-GC relaxes strict constraints of graph convolutions, leading to stronger representation capability. Combining CTR-GC with temporal modeling modules, we develop a powerful graph convolutional network named CTR-GCN which notably outperforms state-of-the-art methods on the NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets.1 Ziqi Zhang 0010, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004 |
ICCV | 6 |
| 2021 | Learn to Match: Automatic Matching Network Design for Visual TrackingabstractSiamese tracking has achieved groundbreaking performance in recent years, where the essence is the efficient matching operator cross-correlation and its variants. Besides the remarkable success, it is important to note that the heuristic matching network design relies heavily on expert experience. Moreover, we experimentally find that one sole matching operator is difficult to guarantee stable tracking in all challenging environments. Thus, in this work, we introduce six novel matching operators from the perspective of feature fusion instead of explicit similarity learning, namely Concatenation, Pointwise-Addition, Pairwise-Relation, FiLM, Simple-Transformer and Transductive-Guidance, to explore more feasibility on matching operator selection. The analyses reveal these operators’ selective adaptability on different environment degradation types, which inspires us to combine them to explore complementary features. To this end, we propose binary channel manipulation (BCM) to search for the optimal combination of these operators. BCM determines to retrain or discard one operator by learning its contribution to other tracking steps. By inserting the learned matching networks to a strong baseline tracker Ocean [47], our model achieves favorable gains by 67.2 → 71.4, 52.6 → 58.3, 70.3 → 76.0 success on OTB100, LaSOT, and TrackingNet, respectively. Notably, Our tracker, dubbed AutoMatch, uses less than half of training data/time than the baseline tracker, and runs at 50 FPS using PyTorch. Code and model are released at https://github.com/JudasDie/SOTS. Yihao Liu 0001, Xiao Wang 0014, Bing Li 0001, Weiming Hu 0004 |
ICCV | 5 |
| 2021 | DSIC: Dynamic Sample-Individualized Connector for Multi-Scale Object DetectionabstractAlthough object detection has reached a milestone recently, the scale variation is still the key challenge. Integrating multilevel features is presented to alleviate the problems, like Feature Pyramid Network (FPN) and its improvements. However, the specifically designed architectures and fixed data flow paths of these methods are not flexible for feature fusion, especially when fed with various samples. To overcome the limitations, we propose a Dynamic Sample-Individualized Connector (DSIC) for multi-scale object detection, which dynamically adjusts network connections to fit different samples. In particular, DSIC consists of two components: Intra-scale Selection Gate (ISG) and Cross-scale Selection Gate (CSG). With the help of the presented gate operator, ISG adaptively extracts proper multi-level features from backbone as the inputs of feature integration. CSG automatically activates informative data flow paths based on the extracted multi-level features. These two components are both plug-and-play and can be embedded in any backbone. Experimental results demonstrate that the proposed method outperforms the state-of-the- arts. Zekun Li 0006, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Yanan Miao |
ICME | 4 |
| 2021 | Adaptive Coarse-to-Fine Interactor for Multi-Scale Object DetectionabstractScale variation is one of the key challenges of object detection. Multi-level feature fusion is presented to alleviate the problems, e.g., Feature Pyramid Network (FPN) and its extended methods. However, the input features fed into these methods and the interaction among features from different levels are insufficient and rigid. To fully exploit the features of multi-scale objects and enhance the feature interaction, we propose a novel and effective framework called Adaptive Coarse-to-Fine Interactor (ACFI). Specifically, ACFI consists of three cascaded components: Multi-Resolution Fusion (MRF), Fine-Grained Interaction (FGI), and Edge-aware Enhancement (EAE). MRF adaptively extracts multi-level features from multi-resolution images and multi-stage features, and then these features are fed into FGI to have a fine-grained interaction utilizing bottom-up guidance. After that, EAE further refines the features obtained by FGI, and enhances the detailed edge information and suppresses the redundant noise. After the coarse-to-fine process, we can obtain powerful multiscale representations of various objects. Each component can be embedded into any backbones, separately. Experimental results show the superiority of our method and verify the effectiveness of each proposed module. Zekun Li 0006, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
IJCNN | 4 |
| 2021 | Web Objectionable Video Recognition Based on Deep Multi-Instance Learning With Representative Prototypes SelectionabstractTo protect underage people from accessing objectionable videos in the Internet, an effective objectionable video recognition algorithm is necessary for web filtering. Recently, the multi-instance learning has been introduced for objectionable video recognition and achieves impressive results. However, hand-crafted features as well as redundant and noisy frames in objectionable videos become an intractable problem that inevitably degrades the recognition performance. In this paper, we propose a novel representative prototype selection algorithm embedding deep multi-instance representation learning. In the proposed method, an improved convolutional neural network is designed for multimodal multi-instance feature learning and a self-expressive dictionary learning model based on sparse and low rank constraint is designed to select the representative prototypes from each subspace of instances. Then the bag-level feature is constructed via mapping the bag to the selected prototypes. Experiments on three objectionable video sets show the effectiveness of our method for objectionable video recognition. Xinmiao Ding, Bing Li 0001, Yangxi Li, Weihua Xiong, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2021 | Robust Texture-Aware Computer-Generated Image Forensic: Benchmark and AlgorithmabstractWith advances in rendering techniques and generative adversarial networks, computer-generated (CG) images tend to be indistinguishable from photographic (PG) images. Revisiting previous works towards CG image forensic, we observed that existing datasets are constructed years ago and limited in both quantity and diversity. Besides, current algorithms only consider the global visual features for forensic, ignoring finer differences between CG and PG images. To mitigate these problems, we first contribute a Large-Scale CG images Benchmark (LSCGB), and then propose a simple yet strong baseline model to address the forensic task. On the one hand, the introduced benchmark has three superior properties, 1) large-scale: the benchmark contains 71168 CG and 71168 PG images with the corresponding expert-annotated labels. It is orders of magnitude bigger than previous datasets. 2) high diversity: we collect CG images from 4 different scenes generated by various rendering techniques. The PG images are varied in terms of image content, camera models, and photographer styles. 3) small bias: we carefully filter the collected images to ensure that the distributions of color, brightness, tone and saturation between CG and PG images are close. Furthermore, inspired by an empirical study on texture difference between CG and PG images, an effective texture-aware network is proposed to improve forensic accuracy. Concretely, we first strengthen texture information of multilevel features extracted from a backbone. Then, the relations among feature channels are explored by learning its gram matrix. Each feature channel represents a specific texture pattern. The gram matrix is thus able to embed the finer texture differences. Experimental results demonstrate that this baseline surpasses the existing methods. The benchmark is publically available at https://github.com/wmbai/LSCGB. Weiming Bai, Bing Li 0001, Yangxi Li, Congxuan Zhang, Weiming Hu 0004 |
IEEE Trans. Image Process. | 7 |
| 2021 | Multi-Scale Low-Discriminative Feature Reactivation for Weakly Supervised Object LocalizationabstractFor weakly supervised object localization (WSOL), how to avoid the network focusing only on some small discriminative parts is a main challenge needed to solve. The widely-used Class Activation Mapping (CAM) based paradigm usually employs Adversarial Learning (AL) strategy to search more object parts by constantly hiding discovered object features, but the adversarial process is difficult to control. In this paper, we propose a novel CAM-based framework with Multi-scale Low-Discriminative Feature Reactivation (mLDFR) for WSOL. The mLDFR framework reactivates the low-discriminative object parts via bottom-up continuous feature maps recalibration and multi-scale object category mapping. Compared with the AL-based methods, our method fully improves the localization power of the network without damaging the classification power and can perform multi-instance localization, which are hard to achieve under the AL-based framework. Moreover, the mLDFR framework is flexible, and can be built on the top of various classical CNN backbones. Experimental results demonstrate the superiority of our method. With VGG16 as backbone, we achieve 46.96% Cls-Loc top1 err and 66.12% CorLoc on ILSVRC2014, 38.07% Cls-Loc top1 err and 75.04% CorLoc on CUB200-2011, surpassing the state-of-the-arts by a large margin. Bo Wang 0147, Chunfeng Yuan, Bing Li 0001, Xinmiao Ding, Zeya Li, Ying Wu 0001, Weiming Hu 0004 |
IEEE Trans. Image Process. | 7 |
| 2021 | Toward Accurate Pixelwise Object Tracking via Attention RetrievalabstractPixelwise single object tracking is challenging due to the competition of running speeds and segmentation accuracy. Current state-of-the-art real-time approaches seamlessly connect tracking and segmentation by sharing computation of the backbone network, e.g., SiamMask and D3S fork a light branch from the tracking model to predict segmentation mask. Although efficient, directly reusing features from tracking networks may harm the segmentation accuracy, since background clutter in the backbone feature tends to introduce false positives in segmentation. To mitigate this problem, we propose a unified tracking-retrieval-segmentation framework consisting of an attention retrieval network (ARN) and an iterative feedback network (IFN). Instead of segmenting the target inside the bounding box, the proposed framework performs soft spatial constraints on backbone features to obtain an accurate global segmentation map. Concretely, in ARN, a look-up-table (LUT) is first built by sufficiently using the information of the first frame. By retrieving it, a target-aware attention map is generated to suppress the negative influence of background clutter. To ulteriorly refine the contour of the segmentation, IFN iteratively enhances the features at different resolutions by taking the predicted mask as feedback guidance. Our framework sets a new state of the art on the recent pixelwise tracking benchmark VOT2020 and runs at 40 fps. Notably, the proposed model surpasses SiamMask by 11.7/4.2/5.5 points on VOT2020, DAVIS2016, and DAVIS2017, respectively. Code is available at https://github.com/JudasDie/SOTS. Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Houwen Peng |
IEEE Trans. Image Process. | 4 |
| 2021 | EDP: An Efficient Decomposition and Pruning Scheme for Convolutional Neural Network CompressionabstractModel compression methods have become popular in recent years, which aim to alleviate the heavy load of deep neural networks (DNNs) in real-world applications. However, most of the existing compression methods have two limitations: 1) they usually adopt a cumbersome process, including pretraining, training with a sparsity constraint, pruning/decomposition, and fine-tuning. Moreover, the last three stages are usually iterated multiple times. 2) The models are pretrained under explicit sparsity or low-rank assumptions, which are difficult to guarantee wide appropriateness. In this article, we propose an efficient decomposition and pruning (EDP) scheme via constructing a compressed-aware block that can automatically minimize the rank of the weight matrix and identify the redundant channels. Specifically, we embed the compressed-aware block by decomposing one network layer into two layers: a new weight matrix layer and a coefficient matrix layer. By imposing regularizers on the coefficient matrix, the new weight matrix learns to become a low-rank basis weight, and its corresponding channels become sparse. In this way, the proposed compressed-aware block simultaneously achieves low-rank decomposition and channel pruning by only one single data-driven training stage. Moreover, the network of architecture is further compressed and optimized by a novel Pruning & Merging (PM) module which prunes redundant channels and merges redundant decomposed layers. Experimental results (17 competitors) on different data sets and networks demonstrate that the proposed EDP achieves a high compression ratio with acceptable accuracy degradation and outperforms state-of-the-arts on compression rate, accuracy, inference time, and run-time memory. Xiaofeng Ruan, Yufan Liu 0001, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Yangxi Li, Stephen J. Maybank |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | RDSNet: A New Deep Architecture forReciprocal Object Detection and Instance SegmentationabstractObject detection and instance segmentation are two fundamental computer vision tasks. They are closely correlated but their relations have not yet been fully explored in most previous work. This paper presents RDSNet, a novel deep architecture for reciprocal object detection and instance segmentation. To reciprocate these two tasks, we design a two-stream structure to learn features on both the object level (i.e., bounding boxes) and the pixel level (i.e., instance masks) jointly. Within this structure, information from the two streams is fused alternately, namely information on the object level introduces the awareness of instance and translation variance to the pixel level, and information on the pixel level refines the localization accuracy of objects on the object level in return. Specifically, a correlation module and a cropping module are proposed to yield instance masks, as well as a mask based boundary refinement module for more accurate bounding boxes. Extensive experimental analyses and comparisons on the COCO dataset demonstrate the effectiveness and efficiency of RDSNet. The source code is available at https://github.com/wangsr126/RDSNet. Shaoru Wang, Yongchao Gong, Junliang Xing, Lichao Huang, Chang Huang, Weiming Hu 0004 |
AAAI | 6 |
| 2020 | Recursive Least-Squares Estimator-Aided Online Learning for Visual TrackingabstractOnline learning is crucial to robust visual object tracking as it can provide high discrimination power in the presence of background distractors. However, there are two contradictory factors affecting its successful deployment on the real visual tracking platform: the discrimination issue due to the challenges in vanilla gradient descent, which does not guarantee good convergence; the robustness issue due to over-fitting resulting from excessive update with limited memory size (the oldest samples are discarded). Despite many dedicated techniques proposed to somehow treat those issues, in this paper we take a new way to strike a compromise between them based on the recursive least-squares estimation (LSE) algorithm. After connecting each fully-connected layer with LSE separately via normal equations, we further propose an improved mini-batch stochastic gradient descent algorithm for fully-connected network learning with memory retention in a recursive fashion. This characteristic can spontaneously reduce the risk of over-fitting resulting from catastrophic forgetting in excessive online learning. Meanwhile, it can effectively improve convergence though the cost function is computed over all the training samples that the algorithm has ever seen. We realize this recursive LSE-aided online learning technique in the state-of-the-art RT-MDNet tracker, and the consistent improvements on four challenging benchmarks prove its efficiency without additional offline training and too much tedious work on parameter adjusting. Weiming Hu 0004, Yan Lu 0001 |
CVPR | 2 |
| 2020 | Object Relational Graph With Teacher-Recommended Learning for Video CaptioningabstractTaking full advantage of the information from both vision and language is critical for the video captioning task. Existing models lack adequate visual representation due to the neglect of interaction between object, and sufficient training for content-related words due to long-tailed problems. In this paper, we propose a complete video captioning system including both a novel model and an effective training strategy. Specifically, we propose an object relational graph (ORG) based encoder, which captures more detailed interaction features to enrich visual representation. Meanwhile, we design a teacher-recommended learning (TRL) method to make full use of the successful external language model (ELM) to integrate the abundant linguistic knowledge into the caption model. The ELM generates more semantically similar word proposals which extend the groundtruth words used for training to deal with the long-tailed problem. Experimental evaluations on three benchmarks: MSVD, MSR-VTT and VATEX show the proposed ORG-TRL system achieves state-of-the-art performance. Extensive ablation studies and visualizations illustrate the effectiveness of our system. Ziqi Zhang 0010, Yaya Shi, Chunfeng Yuan, Bing Li 0001, Peijin Wang, Weiming Hu 0004, Zhengjun Zha |
CVPR | 6 |
| 2020 | Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model
Yufan Liu 0001, Minglang Qiao, Mai Xu, Bing Li 0001, Weiming Hu 0004, Ali Borji |
ECCV (20) | 5 |
| 2020 | Ocean: Object-Aware Anchor-Free Tracking
Houwen Peng, Jianlong Fu, Bing Li 0001, Weiming Hu 0004 |
ECCV (21) | 5 |
| 2020 | End-to-End Temporal Feature Aggregation for Siamese TrackersabstractWhile siamese networks have demonstrated the significant improvement on object tracking performances, how to utilize the temporal information in siamese trackers has not been widely studied yet. In this paper, we introduce a novel siamese tracking architecture equipped with a temporal aggregation module, which improves the per-frame features by aggregating temporal information from adjacent frames. This temporal fusion strategy enables the siamese trackers to handle poor object appearance like motion blur, occlusion, etc. Furthermore, we incorporate the adversarial dropout module in the siamese network for computing discriminative target features in an end-to-end-fashion. Comprehensive experiments demonstrate that the proposed tracker performs favorably against state-of-the-art trackers. Zhenbang Li, Qiang Wang 0051, Bing Li 0001, Weiming Hu 0004 |
ICIP | 5 |
| 2020 | Globally Spatial-Temporal Perception: a Long-Term Tracking SystemabstractAlthough siamese trackers have achieved superior performance, these kinds of approaches tend to favour the local search mechanism and are thus prone to accumulating inaccuracies of predicted positions, leading to tracking drift over time, especially in long-term tracking scenario. To solve these problems, we propose a siamese tracker in the spirit of the faster RCNN's two-stage detection paradigm. This new tracker is dedicated to reducing cumulative inaccuracies and improving robustness based on a global perception mechanism, which allows the target to be retrieved in time spatially over the whole image plane. Since the very deep network can be enabled for feature learning in this two-stage tracking framework, the power of discrimination is guaranteed. What's more, we also add a CNN-based trajectory prediction module exploiting the target's temporal motion information to mitigate the interference of distractors. These two spatial and temporal modules exploit both the high-level appearance information and complementary trajectory information to improve the tracking robustness. Comprehensive experiments demonstrate that the proposed Globally Spatial-Temporal Perception-based tracking system performs favorably against state-of-the-art trackers. Zhenbang Li, Qiang Wang 0051, Bing Li 0001, Weiming Hu 0004 |
ICIP | 5 |
| 2020 | Anchor-Free One-Stage Online Multi-object Tracking
Zongwei Zhou, Yangxi Li, Junliang Xing, Liang Li 0003, Weiming Hu 0004 |
PRCV (2) | 6 |
| 2020 | Dual L1-Normalized Context Aware Tensor Power Iteration and Its Applications to Multi-object Tracking and Multi-graph MatchingabstractAbstract The multi-dimensional assignment problem is universal for data association analysis such as data association-based visual multi-object tracking and multi-graph matching. In this paper, multi-dimensional assignment is formulated as a rank-1 tensor approximation problem. A dualL1-normalized context/hyper-context aware tensor power iteration optimization method is proposed. The method is applied to multi-object tracking and multi-graph matching. In the optimization method, tensor power iteration with the dual unit norm enables the capture of information across multiple sample sets. Interactions between sample associations are modeled as contexts or hyper-contexts which are combined with the global affinity into a unified optimization. The optimization is flexible for accommodating various types of contextual models. In multi-object tracking, the global affinity is defined according to the appearance similarity between objects detected in different frames. Interactions between objects are modeled as motion contexts which are encoded into the global association optimization. The tracking method integrates high order motion information and high order appearance variation. The multi-graph matching method carries out matching over graph vertices and structure matching over graph edges simultaneously. The matching consistency across multi-graphs is based on the high-order tensor optimization. Various types of vertex affinities and edge/hyper-edge affinities are flexibly integrated. Experiments on several public datasets, such as the MOT16 challenge benchmark, validate the effectiveness of the proposed methods. Weiming Hu 0004, Xinchu Shi, Zongwei Zhou, Junliang Xing, Haibin Ling, Stephen J. Maybank |
Int. J. Comput. Vis. | 1 |
| 2020 | Tracking-by-Fusion via Gaussian Process Regression Extended to Transfer LearningabstractThis paper presents a new Gaussian Processes (GPs)-based particle filter tracking framework. The framework non-trivially extends Gaussian process regression (GPR) to transfer learning, and, following the tracking-by-fusion strategy, integrates closely two tracking components, namely a GPs component and a CFs one. First, the GPs component analyzes and models the probability distribution of the object appearance by exploiting GPs. It categorizes the labeled samples into auxiliary and target ones, and explores unlabeled samples in transfer learning. The GPs component thus captures rich appearance information over object samples across time. On the other hand, to sample an initial particle set in regions of high likelihood through the direct simulation method in particle filtering, the powerful yet efficient correlation filters (CFs) are integrated, leading to the CFs component. In fact, the CFs component not only boosts the sampling quality, but also benefits from the GPs component, which provides re-weighted knowledge as latent variables for determining the impact of each correlation filter template from the auxiliary samples. In this way, the transfer learning based fusion enables effective interactions between the two components. Superior performance on four object tracking benchmarks (OTB-2015, Temple-Color, and VOT2015/2016), and in comparison with baselines and recent state-of-the-art trackers, has demonstrated clearly the effectiveness of the proposed framework. Qiang Wang 0051, Junliang Xing, Haibin Ling, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Graph convolutional network with structure pooling and joint-wise channel attention for action recognition
Gaoqun Ma, Chunfeng Yuan, Bing Li 0001, Fangshi Wang, Weiming Hu 0004 |
Pattern Recognit. | 7 |
| 2020 | Distractor-aware discrimination learning for online multiple object tracking
Zongwei Zhou, Wenhan Luo, Qiang Wang 0051, Junliang Xing, Weiming Hu 0004 |
Pattern Recognit. | 5 |
| 2020 | Manipulating Template Pixels for Model Adaptation of Siamese Visual TrackingabstractIn this letter, we show that the challenging model adaptation task in visual object tracking can be handled by simply manipulating pixels of the template image in Siamese networks. For a target that is not included in the offline training set, a slight modification of the template image pixels will improve the prediction result of the offline trained Siamese network. The popular adversarial example generation methods can be used to perform template pixel manipulation for model adaptation. Different from current template update methods, which aim to combine the target features from previous frames, we focus on the initial adaptation using target ground-truth in the first frame. Our model adaptation method is pluggable, in the sense that it does not alter the overall architecture of its base tracker. To our knowledge, this work is the first attempt to directly manipulating template pixels for model adaptation in Siamese-based trackers. Extensive experiments on recent benchmarks demonstrate that our method achieves better performance than some other state-of-the-art trackers. Our code is available at https://github.com/lizhenbang56/MTP. Zhenbang Li, Bing Li 0001, Liang Li 0006, Weiming Hu 0004 |
IEEE Signal Process. Lett. | 5 |
| 2020 | Multi-Cue Semi-Supervised Color Constancy With Limited Training SamplesabstractColor constancy is one of the fundamental tasks in computer vision. Many supervised methods, including recently proposed Convolutional Neural Networks (CNN)-based methods, have been proved to work well on this problem, but they often require a sufficient number of labeled data. However, it is expensive and time-consuming to collect a large number of labeled training images with accurately measured illumination. In order to reduce the dependence on labeled images and leverage unlabeled ones without measured illumination, we propose a novel semi-supervised framework with limited training samples for illumination estimation. Our key insight is that the images with similar features from different cues will share similar lighting conditions. Consequently, three graphs based on three visual cues, low-level RGB color distribution, mid-level initial illuminant estimates and high-level scene content, are constructed to represent the relationship among different images. Then a multi-cue semi-supervised color constancy method (MSCC) is proposed after integrating these three graphs into a unified model. Extensive experiments on benchmark datasets demonstrate that our proposed MSCC method outperforms nearly all the existing supervised methods with limited labeled samples. Even with no unlabeled samples, MSCC still obtains better performance and stableness than most supervised methods. Xinwei Huang, Bing Li 0001, Shuai Li 0001, Weihua Xiong, Xuanwu Yin, Weiming Hu 0004, Hong Qin 0001 |
IEEE Trans. Image Process. | 7 |
| 2020 | Anisotropic Convolution for Image ClassificationabstractConvolutional neural networks are built upon simple but useful convolution modules. The traditional convolution has a limitation on feature extraction and object localization due to its fixed scale and geometric structure. Besides, the loss of spatial information also restricts the networks' performance and depth. To overcome these limitations, this paper proposes a novel anisotropic convolution by adding a scale factor and a shape factor into the traditional convolution. The anisotropic convolution augments the receptive fields flexibly and dynamically depending on the valid sizes of objects. In addition, the anisotropic convolution is a generalized convolution. The traditional convolution, dilated convolution and deformable convolution can be viewed as its special cases. Furthermore, in order to improve the training efficiency and avoid falling into a local optimum, this paper introduces a simplified implementation of the anisotropic convolution. The anisotropic convolution can be applied to arbitrary convolutional networks and the enhanced networks are called ACNs (anisotropic convolutional networks). Experimental results show that ACNs achieve better performance than many state-of-the-art methods and the baseline networks in tasks of image classification and object localization, especially in classification task of tiny images. Bing Li 0001, Chunfeng Yuan, Yangxi Li, Haohao Wu, Weiming Hu 0004, Fangshi Wang |
IEEE Trans. Image Process. | 6 |
| 2020 | Tangent Fisher Vector on Matrix Manifolds for Action RecognitionabstractIn this paper, we address the problem of representing and recognizing human actions from videos on matrix manifolds. For this purpose, we propose a new vector representation method, named tangent Fisher vector, to describe video sequences in the Fisher kernel framework. We first extract dense curved spatio-temporal cuboids from each video sequence. Compared with the traditional 'straight cuboids', the dense curved spatio-temporal cuboids contain much more local motion information. Each cuboid is then described using a linear dynamical system (LDS) to simultaneously capture the local appearance and dynamics. Furthermore, a simple yet efficient algorithm is proposed to learn the LDS parameters and approximate the observability matrix at the same time. Each video sequence is thus represented by a set of LDSs. Considering that each LDS can be viewed as a point in a Grassmann manifold, we propose to learn an intrinsic GMM on the manifold to cluster the LDS points. Finally a tangent Fisher vector is computed by first accumulating all the tangent vectors in each Gaussian component, and then concatenating the normalized results across all the Gaussian components. A kernel is defined to measure the similarity between tangent Fisher vectors for classification and recognition of a video sequence. This approach is evaluated on the state-of-the-art human action benchmark datasets. The recognition performance is competitive when compared with current state-of-the-art results. Guan Luo, Jiutong Wei, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Image Process. | 3 |
| 2020 | STA-CNN: Convolutional Spatial-Temporal Attention Learning for Action RecognitionabstractConvolutional Neural Networks have achieved excellent successes for object recognition in still images. However, the improvement of Convolutional Neural Networks over the traditional methods for recognizing actions in videos is not so significant, because the raw videos usually have much more redundant or irrelevant information than still images. In this paper, we propose a Spatial-Temporal Attentive Convolutional Neural Network (STA-CNN) which selects the discriminative temporal segments and focuses on the informative spatial regions automatically. The STA-CNN model incorporates a Temporal Attention Mechanism and a Spatial Attention Mechanism into a unified convolutional network to recognize actions in videos. The novel Temporal Attention Mechanism automatically mines the discriminative temporal segments from long and noisy videos. The Spatial Attention Mechanism firstly exploits the instantaneous motion information in optical flow features to locate the motion salient regions and it is then trained by an auxiliary classification loss with a Global Average Pooling layer to focus on the discriminative non-motion regions in the video frame. The STA-CNN model achieves the state-of-the-art performance on two of the most challenging datasets, UCF-101 (95.8%) and HMDB-51 (71.5%). Hao Yang 0010, Chunfeng Yuan, Li Zhang 0050, Yunda Sun, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Image Process. | 5 |
| 2020 | Anomaly Detection Using Local Kernel Density Estimation and Context-Based RegressionabstractCurrent local density-based anomaly detection methods are limited in that the local density estimation and the neighborhood density estimation are not accurate enough for complex and large databases, and the detection performance depends on the size parameter of the neighborhood. In this paper, we propose a new kernel function to estimate samples' local densities and propose a weighted neighborhood density estimation to increase the robustness to changes in the neighborhood size. We further propose a local kernel regression estimator and a hierarchical strategy for combining information from the multiple scale neighborhoods to refine anomaly factors of samples. We apply our general anomaly detection method to image saliency detection by regarding salient pixels in objects as anomalies to the background regions. Local density estimation in the visual feature space and kernel-based saliency score propagation in the image enable the assignment of similar saliency values to homogenous object regions. Experimental results on several benchmark datasets demonstrate that our anomaly detection methods overall outperform several state-of-art anomaly detection methods. The effectiveness of our image saliency detection method is validated by comparison with several state-of-art saliency detection methods. Weiming Hu 0004, Bing Li 0001, Ou Wu 0001, Junping Du 0001, Stephen J. Maybank |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2019 | Knowledge Distillation via Instance Relationship GraphabstractThe key challenge of knowledge distillation is to extract general, moderate and sufficient knowledge from a teacher network to guide a student network. In this paper, a novel Instance Relationship Graph (IRG) is proposed for knowledge distillation. It models three kinds of knowledge, including instance features, instance relationships and feature space transformation, while the latter two kinds of knowledge are neglected by previous methods. Firstly, the IRG is constructed to model the distilled knowledge of one network layer, by considering instance features and instance relationships as vertexes and edges respectively. Secondly, an IRG transformation is proposed to models the feature space transformation across layers. It is more moderate than directly mimicking the features at intermediate layers. Finally, hint loss functions are designed to force a student's IRGs to mimic the structures of a teacher's IRGs. The proposed method effectively captures the knowledge along the whole network via IRGs, and thus shows stable convergence and strong robustness to different network architectures. In addition, the proposed method shows superior performance over existing methods on datasets of various scales. Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Chunfeng Yuan, Weiming Hu 0004, Yangxi Li, Yunqiang Duan |
CVPR | 5 |
| 2019 | Fast Online Object Tracking and Segmentation: A Unifying ApproachabstractIn this paper we illustrate how to perform both visual object tracking and semi-supervised video object segmentation, in real-time, with a single simple approach. Our method, dubbed SiamMask, improves the offline training procedure of popular fully-convolutional Siamese approaches for object tracking by augmenting their loss with a binary segmentation task. Once trained, SiamMask solely relies on a single bounding box initialisation and operates online, producing class-agnostic object segmentation masks and rotated bounding boxes at 55 frames per second. Despite its simplicity, versatility and fast speed, our strategy allows us to establish a new state-of-the-art among real-time trackers on VOT-2018, while at the same time demonstrating competitive performance and the best speed for the semi-supervised video object segmentation task on DAVIS-2016 and DAVIS-2017. Qiang Wang 0051, Li Zhang 0040, Luca Bertinetto, Weiming Hu 0004, Philip Torr 0001 |
CVPR | 4 |
| 2019 | Anchor Diffusion for Unsupervised Video Object SegmentationabstractUnsupervised video object segmentation has often been tackled by methods based on recurrent neural networks and optical flow. Despite their complexity, these kinds of approach tend to favour short-term temporal dependencies and are thus prone to accumulating inaccuracies, which cause drift over time. Moreover, simple (static) image segmentation models, alone, can perform competitively against these methods, which further suggests that the way temporal dependencies are modelled should be reconsidered. Motivated by these observations, in this paper we explore simple yet effective strategies to model long-term temporal dependencies. Inspired by the non-local operators, we introduce a technique to establish dense correspondences between pixel embeddings of a reference "anchor" frame and the current one. This allows the learning of pairwise dependencies at arbitrarily long distances without conditioning on intermediate frames. Without online supervision, our approach can suppress the background and precisely segment the foreground object even in challenging scenarios, while maintaining consistent performance over time. With a mean IoU of 81.7%, our method ranks first on the DAVIS-2016 leaderboard of unsupervised methods, while still being competitive against state-of-the-art online semi-supervised approaches. We further evaluate our method on the FBMS dataset and the video saliency dataset ViSal, showing results competitive with the state of the art. Zhao Yang 0002, Qiang Wang 0051, Luca Bertinetto, Song Bai 0001, Weiming Hu 0004, Philip Torr 0001 |
ICCV | 5 |
| 2019 | Multimodal Semantic Attention Network for Video CaptioningabstractInspired by the fact that different modalities in videos carry complementary information, we propose a Multimodal Semantic Attention Network(MSAN), which is a new encoder-decoder framework incorporating multimodal semantic attributes for video captioning. In the encoding phase, we detect and generate multimodal semantic attributes by formulating it as a multi-label classification problem. Moreover, we add auxiliary classification loss to our model that can obtain more effective visual features and high-level multimodal semantic attribute distributions for sufficient video encoding. In the decoding phase, we extend each weight matrix of the conventional LSTM to an ensemble of attribute-dependent weight matrices, and employ attention mechanism to pay attention to different attributes at each time of the captioning process. We evaluate algorithm on two popular public benchmarks: MSVD and MSR-VTT, achieving competitive results with current state-of-the-art across six evaluation metrics. Bing Li 0001, Chunfeng Yuan, Zhengjun Zha, Weiming Hu 0004 |
ICME | 5 |
| 2019 | Rank-1 Tensor Approximation for High-Order Association in Multi-target Tracking
Xinchu Shi, Haibin Ling, Weiming Hu 0004, Peng Chu, Junliang Xing |
Int. J. Comput. Vis. | 4 |
| 2019 | Asymmetric 3D Convolutional Neural Networks for action recognition
Hao Yang 0010, Chunfeng Yuan, Bing Li 0001, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank |
Pattern Recognit. | 6 |
| 2018 | Hierarchical Nonlinear Orthogonal Adaptive-Subspace Self-Organizing Map Based Feature Extraction for Human Action RecognitionabstractFeature extraction is a critical step in the task of action recognition. Hand-crafted features are often restricted because of their fixed forms and deep learning features are more effective but need large-scale labeled data for training. In this paper, we propose a new hierarchical Nonlinear Orthogonal Adaptive-Subspace Self-Organizing Map(NOASSOM) to adaptively and learn effective features from data without supervision. NOASSOM is extended from Adaptive-Subspace Self-Organizing Map (ASSOM) which only deals with linear data and is trained with supervision by the labeled data. Firstly, by adding a nonlinear orthogonal map layer, NOASSOM is able to handle the nonlinear input data and it avoids defining the specific form of the nonlinear orthogonal map by a kernel trick. Secondly, we modify loss function of ASSOM such that every input sample is used to train model individually. In this way, NOASSOM effectively learns the statistic patterns from data without supervision. Thirdly, we propose a hierarchical NOASSOM to extract more representative features. Finally, we apply the proposed hierarchical NOASSOM to efficiently describe the appearance and motion information around trajectories for action recognition. Experimental results on widely used datasets show that our method has superior performance than many state-of-the-art hand-crafted features and deep learning features based methods. Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Hao Yang 0010, Zhikang Fu |
AAAI | 4 |
| 2018 | Deep Cost-Sensitive and Order-Preserving Feature Learning for Cross-Population Age EstimationabstractFacial age estimation from a face image is an important yet very challenging task in computer vision, since humans with different races and/or genders, exhibit quite different patterns in their facial aging processes. To deal with the influence of race and gender, previous methods perform age estimation within each population separately. In practice, however, it is often very difficult to collect and label sufficient data for each population. Therefore, it would be helpful to exploit an existing large labeled dataset of one (source) population to improve the age estimation performance on another (target) population with only a small labeled dataset available. In this work, we propose a Deep Cross-Population (DCP) age estimation model to achieve this goal. In particular, our DCP model develops a two-stage training strategy. First, a novel cost-sensitive multitask loss function is designed to learn transferable aging features by training on the source population. Second, a novel order-preserving pair-wise loss function is designed to align the aging features of the two populations. By doing so, our DCP model can transfer the knowledge encoded in the source population to the target population. Extensive experiments on the two of the largest benchmark datasets show that our DCP model outperforms several strong baseline methods and many state-of-the-art methods. Kai Li 0022, Junliang Xing, Chi Su, Weiming Hu 0004, Stephen J. Maybank |
CVPR | 4 |
| 2018 | Learning Attentions: Residual Attentional Siamese Network for High Performance Online Visual TrackingabstractOffline training for object tracking has recently shown great potentials in balancing tracking accuracy and speed. However, it is still difficult to adapt an offline trained model to a target tracked online. This work presents a Residual Attentional Siamese Network (RASNet) for high performance object tracking. The RASNet model reformulates the correlation filter within a Siamese tracking framework, and introduces different kinds of the attention mechanisms to adapt the model without updating the model online. In particular, by exploiting the offline trained general attention, the target adapted residual attention, and the channel favored feature attention, the RASNet not only mitigates the over-fitting problem in deep network training, but also enhances its discriminative capacity and adaptability due to the separation of representation learning and discriminator learning. The proposed deep architecture is trained from end to end and takes full advantage of the rich spatial temporal information to achieve robust visual tracking. Experimental results on two latest benchmarks, OTB-2015 and VOT2017, show that the RASNet tracker has the state-of-the-art tracking accuracy while runs at more than 80 frames per second. Qiang Wang 0051, Zhu Teng, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank |
CVPR | 5 |
| 2018 | Interaction-Aware Spatio-Temporal Pyramid Attention Networks for Action Classification
Chunfeng Yuan, Bing Li 0001, Yangxi Li, Weiming Hu 0004 |
ECCV (16) | 6 |
| 2018 | Visual Tracking via Spatially Aligned Correlation Filters Network
Mengdan Zhang, Qiang Wang 0051, Junliang Xing, Peixi Peng, Weiming Hu 0004, Stephen J. Maybank |
ECCV (3) | 6 |
| 2018 | Distractor-Aware Siamese Networks for Visual Object Tracking
Qiang Wang 0051, Bo Li 0114, Wei Wu 0021, Weiming Hu 0004 |
ECCV (9) | 6 |
| 2018 | SPCNet: Scale Position Correlation Network for End-to-End Visual TrackingabstractWe present a novel Scale Position Correlation Network (SPCNet) for learning to track objects robustly and efficiently. Different from most previous Correlation Filter (CF) based tracking models, SPCNet unifies the feature representation learning and CF based appearance modeling within one end-to-end learnable framework. In particular, SPCNet learns to track objects within a joint scale-position space, and is very effective in learning features for the accurate prediction of object scale and position. To learn our model from end to end, the SPCNet introduces a differentiable correlation filter layer into a Siamese architecture. Therefore, the localization error can be effectively back-propagated through the whole network, enabling fast adaptation of feature learning and appearance modeling for the objects to be tracked. Such task driven feature learning admits a very lightweight design that can be efficiently pre-trained. In addition, the dense appearance modeling in the joint scale-position space is also efficient. It benefits from the computation of gradients within the Fourier frequency domain. Such careful architecture design ensures that SPCNet is effective and efficient with a small model size. Extensive experimental analyses and evaluations on three largest benchmarks, OTB-2013, OTB-2015, and VOT2015, demonstrate its superiority over many state-of-the-art algorithms. Qiang Wang 0051, Mengdan Zhang, Junliang Xing, Weiming Hu 0004 |
ICPR | 5 |
| 2018 | Online Multi-Target Tracking with Tensor-Based High-Order Graph MatchingabstractIn this paper we formulate multi-target tracking (MTT) as a high-order graph matching problem and propose a l1-norm tensor power iteration solution. Concretely, the search for trajectory-observation correspondences in MTT task is cast as a hypergraph matching problem to maximize a multi-linear objective function over all permutations of the associations. This function is defined by a tensor representing the affinity between association tuples where pair-wise similarities, motion consistency and spatial structural information can be embedded expediently. To solve the matching problem, a dual-direction unit l1-norm constrained tensor power iteration algorithm is proposed. Additionally, as measuring the appearance affinity with features extracted from the rectangle patch, which is adopted in most methods, has a weak discrimination when bounding boxes overlap each other heavily, we present a deep pair-wise appearance similarity metric based on object mask in this paper where just the features from true target region are utilized. Experimental evaluation shows that our approach achieves an accuracy comparable to state-of-the-art online trackers. The source code of the proposed approach will be released to facilitate further studies on the MTT problem. Zongwei Zhou, Junliang Xing, Mengdan Zhang, Weiming Hu 0004 |
ICPR | 4 |
| 2018 | Do not Lose the Details: Reinforced Representation Learning for High Performance Visual TrackingabstractThis work presents a novel end-to-end trainable CNN model for high performance visual object tracking. It learns both low-level fine-grained representations and a high-level semantic embedding space in a mutual reinforced way, and a multi-task learning strategy is proposed to perform the correlation analysis on representations from both levels. In particular, a fully convolutional encoder-decoder network is designed to reconstruct the original visual features from the semantic projections to preserve all the geometric information. Moreover, the correlation filter layer working on the fine-grained representations leverages a global context constraint for accurate object appearance modeling. The correlation filter in this layer is updated online efficiently without network fine-tuning. Therefore, the proposed tracker benefits from two complementary effects: the adaptability of the fine-grained correlation analysis and the generalization capability of the semantic embedding. Extensive experimental evaluations on four popular benchmarks demonstrate its state-of-the-art performance. Qiang Wang 0051, Mengdan Zhang, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank |
IJCAI | 5 |
| 2018 | Dual Sticky Hierarchical Dirichlet Process Hidden Markov Model and Its Application to Natural Language Description of MotionsabstractIn this paper, a new nonparametric Bayesian model called the dual sticky hierarchical Dirichlet process hidden Markov model (HDP-HMM) is proposed for mining activities from a collection of time series data such as trajectories. All the time series data are clustered. Each cluster of time series data, corresponding to a motion pattern, is modeled by an HMM. Our model postulates a set of HMMs that share a common set of states (topics in an analogy with topic models for document processing), but have unique transition distributions. The number of HMMs and the number of topics are both automatically determined. The sticky prior avoids redundant states and makes our HDP-HMM more effective to model multimodal observations. For the application to motion trajectory modeling, topics correspond to motion activities. The learnt topics are clustered into atomic activities which are assigned predicates. We propose a Bayesian inference method to decompose a given trajectory into a sequence of atomic activities. The sources and sinks in the scene are learnt by clustering endpoints (origins and destinations) of trajectories. The semantic motion regions are learnt using the points in trajectories. On combining the learnt sources and sinks, the learnt semantic motion regions, and the learnt sequence of atomic activities, the action represented by a trajectory can be described in natural language in as automatic a way as possible. The effectiveness of our dual sticky HDP-HMM is validated on several trajectory datasets. The effectiveness of the natural language descriptions for motions is demonstrated on the vehicle trajectories extracted from a traffic scene. Weiming Hu 0004, Guodong Tian, Yongxin Kang, Chunfeng Yuan, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Towards Robust and Accurate Multi-View and Partially-Occluded Face AlignmentabstractFace alignment acts as an important task in computer vision. Regression-based methods currently dominate the approach to solving this problem, which generally employ a series of mapping functions from the face appearance to iteratively update the face shape hypothesis. One keypoint here is thus how to perform the regression procedure. In this work, we formulate this regression procedure as a sparse coding problem. We learn two relational dictionaries, one for the face appearance and the other one for the face shape, with coupled reconstruction coefficient to capture their underlying relationships. To deploy this model for face alignment, we derive the relational dictionaries in a stage-wised manner to perform close-loop refinement of themselves, i.e., the face appearance dictionary is first learned from the face shape dictionary and then used to update the face shape hypothesis, and the updated face shape dictionary from the shape hypothesis is in return used to refine the face appearance dictionary. To improve the model accuracy, we extend this model hierarchically from the whole face shape to face part shapes, thus both the global and local view variations of a face are captured. To locate facial landmarks under occlusions, we further introduce an occlusion dictionary into the face appearance dictionary to recover face shape from partially occluded face appearance. The occlusion dictionary is learned in a data driven manner from background images to represent a set of elemental occlusion patterns, a sparse combination of which models various practical partial face occlusions. By integrating all these technical innovations, we obtain a robust and accurate approach to locate facial landmarks under different face views and possibly severe occlusions for face images in the wild. Extensive experimental analyses and evaluations on different benchmark datasets, as well as two new datasets built by ourselves, have demonstrated the robustness and accuracy of our proposed model, especially for face images with large view variations and/or severe occlusions. Junliang Xing, Zhiheng Niu, Junshi Huang, Weiming Hu 0004, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | FatRegion: A Fast Adaptive Tree-Structured Region Extraction ApproachabstractCoherent image regions can be used as good features for many computer vision tasks, such as object tracking, segmentation, and recognition. Most of previous region extraction methods, however, are not suitable for online applications because of their either heavy computations or unsatisfactory results. We propose a seed-based region growing and merging approach to generate simultaneously coherent and discriminative image regions. We present a quadtree-based seed initialization algorithm to adaptively place seeds into different image areas and then grow them into regions by a color- and edge-guided growing procedure. To merge these regions in different levels, we propose to use the generalized boundary strength to measure the quality of region merging result. In addition, we present a region merging algorithm of linear time complexity to perform efficient and effective region merging. Overall, our new approach simultaneously holds these advantages: 1) it is extremely fast with linear complexity in both time and space, which takes less than 50 ms to process an HVGA image; 2) it can give a direct control of the region number and well adapt to image regions with various sizes and shapes; and 3) it provides a tree-structured representation of the regions and thus can model the image from multiple scales. We evaluate the proposed approach on the standard benchmarks with extensive comparisons with the state-of-the-art methods. The experimental results demonstrate its good comprehensive performances. Example applications using the extracted regions as features for online object tracking and multiclass object segmentation also exhibit its potential for many computer vision tasks. Junliang Xing, Weiming Hu 0004, Haizhou Ai, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Deep Constrained Siamese Hash Coding Network and Load-Balanced Locality-Sensitive Hashing for Near Duplicate Image DetectionabstractWe construct a new efficient near duplicate image detection method using a hierarchical hash code learning neural network and load-balanced locality-sensitive hashing (LSH) indexing. We propose a deep constrained siamese hash coding neural network combined with deep feature learning. Our neural network is able to extract effective features for near duplicate image detection. The extracted features are used to construct a LSH-based index. We propose a load-balanced LSH method to produce load-balanced buckets in the hashing process. The load-balanced LSH significantly reduces the query time. Based on the proposed load-balanced LSH, we design an effective and feasible algorithm for near duplicate image detection. Extensive experiments on three benchmark data sets demonstrate the effectiveness of our deep siamese hash encoding network and load-balanced LSH. Weiming Hu 0004, Yabo Fan, Junliang Xing, Zhaoquan Cai 0001, Stephen J. Maybank |
IEEE Trans. Image Process. | 1 |
| 2018 | Context-Dependent Random Walk Graph Kernels and Tree Pattern Graph Matching Kernels With Applications to Action RecognitionabstractGraphs are effective tools for modeling complex data. Setting out from two basic substructures, random walks and trees, we propose a new family of context-dependent random walk graph kernels and a new family of tree pattern graph matching kernels. In our context-dependent graph kernels, context information is incorporated into primary random walk groups. A multiple kernel learning algorithm with a proposed l1,2-norm regularization is applied to combine context-dependent graph kernels of different orders. This improves the similarity measurement between graphs. In our tree-pattern graph matching kernel, a quadratic optimization with a sparse constraint is proposed to select the correctly matched tree-pattern groups. This augments the discriminative power of the tree-pattern graph matching. We apply the proposed kernels to human action recognition, where each action is represented by two graphs which record the spatiotemporal relations between local feature vectors. Experimental comparisons with state-of-the-art algorithms on several benchmark datasets demonstrate the effectiveness of the proposed kernels for recognizing human actions. It is shown that our kernel based on tree-pattern groups, which have more complex structures and exploit more local topologies of graphs than random walks, yields more accurate results but requires more runtime than the context-dependent walk graph kernel. Weiming Hu 0004, Baoxin Wu, Chunfeng Yuan, Yangxi Li, Stephen J. Maybank |
IEEE Trans. Image Process. | 1 |
| 2018 | Iteratively Divide-and-Conquer Learning for Nonlinear Classification and RankingabstractNonlinear classifiers (i.e., kernel support vector machines (SVMs)) are effective for nonlinear data classification. However, nonlinear classifiers are usually prohibitively expensive when dealing with large nonlinear data. Ensembles of linear classifiers have been proposed to address this inefficiency, which is called the ensemble linear classifiers for nonlinear data problem. In this article, a new iterative learning approach is introduced that involves two steps at each iteration: partitioning the data into clusters according to Gaussian mixture models with local consistency and then training basic classifiers (i.e., linear SVMs) for each cluster. The two divide-and-conquer steps are combined into a graphical model. Meanwhile, with training, each classifier is regarded as a task; clustered multitask learning is employed to capture the relatedness among different tasks and avoid overfitting in each task. In addition, two novel extensions are introduced based on the proposed approach. First, the approach is extended for quality-aware web data classification. In this problem, the types of web data vary in terms of information quality. The ignorance of the variations of information quality of web data leads to poor classification models. The proposed approach can effectively integrate quality-aware factors into web data classification. Second, the approach is extended for listwise learning to rank to construct an ensemble of linear ranking models, whereas most existing listwise ranking methods construct a solely linear ranking model. Experimental results on benchmark datasets show that our approach outperforms state-of-the-art algorithms. During prediction for nonlinear classification, it also obtains comparable classification performance to kernel SVMs, with much higher efficiency. Ou Wu 0001, Weiming Hu 0004 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2017 | Spatio-Temporal Self-Organizing Map Deep Network for Dynamic Object Detection from Videos
Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
CVPR | 4 |
| 2017 | SCNN: Sequential convolutional neural network for human action recognition in videosabstractConvolutional Neural Network (CNN) and Recurrent Neural Network (RNN) are two typical kinds of neural networks. While CNN models have achieved great success on image recognition due to their strong abilities in abstracting spatial information from multiple levels, RNN models have not achieved significant progress in video analyzing tasks (e.g. action recognition), although RNN can inherently model temporal dependencies from videos. In this work, we propose a Sequential Convolutional Neural Network, denoted as SCNN, to extract effective spatial-temporal features from videos, thus incorporating the strengths of both convolutional operation and recurrent operation. Our SCNN model extends RNN to directly process feature maps, rather than vectors flattened from feature maps, to keep spatial structures of the inputs. It replaces the full connections of RNN with convolutional connections to decrease parameter numbers, computational cost, and over-fitting risk. Moreover, we introduce asymmetric convolutional layers into SCNN to reduce parameter numbers and computational cost further. Our final SCNN deep architecture used for action recognition achieves very good performances on two challenging benchmarks, UCF-101 and HMDB-51, outperforming many state-of-the-art methods. Hao Yang 0010, Chunfeng Yuan, Junliang Xing, Weiming Hu 0004 |
ICIP | 4 |
| 2017 | Diversity encouraging ensemble of convolutional networks for high performance action recognitionabstractWe present a simple and effective ensemble method, Diversity Encouraging Ensemble (DEE), for deep convolutional networks to boost their performances. By training the convolutional network in two stages, we generate multiple component networks without adding any training cost. On the one hand, we modify the structure parameters of component networks in the training process to enlarge the diversities of the networks, which is found to be beneficial to improving the ensemble performance. On the other hand, we exploit monotonous decreasing learning rate schedule to accelerate the speed of deep network converging to different local minima, and we decrease the training time of integrating multiple networks to that of training a single network from traditional multi-step learning policy. We evaluate our ensemble method on two challenging action datasets, UCF-101 and HMDB-51, and obtain performance improvements from single deep network and other ensemble methods. Our results also outperform many state-of-the-art action recognition methods. Hao Yang 0010, Chunfeng Yuan, Junliang Xing, Weiming Hu 0004 |
ICIP | 4 |
| 2017 | Human activity prediction using temporally-weighted generalized time warping
Haoran Wang 0001, Wankou Yang, Chunfeng Yuan, Haibin Ling, Weiming Hu 0004 |
Neurocomputing | 5 |
| 2017 | Towards human-like and transhuman perception in AI 2.0: a reviewabstractPerception is the interaction interface between an intelligent system and the real world. Without sophisticated and flexible perceptual capabilities, it is impossible to create advanced artificial intelligence (AI) systems. For the next-generation AI, called ‘AI 2.0’, one of the most significant features will be that AI is empowered with intelligent perceptual capabilities, which can simulate human brain’s mechanisms and are likely to surpass human brain in terms of performance. In this paper, we briefly review the state-of-the-art advances across different areas of perception, including visual perception, auditory perception, speech perception, and perceptual information processing and learning engines. On this basis, we envision several R&D trends in intelligent perception for the forthcoming era of AI 2.0, including: (1) human-like and transhuman active vision; (2) auditory perception and computation in an actual auditory setting; (3) speech perception and computation in a natural interaction setting; (4) autonomous learning of perceptual information; (5) large-scale perceptual information processing and learning platforms; and (6) urban omnidirectional intelligent perception and reasoning engines. We believe these research directions should be highlighted in the future plans for AI 2.0. Yonghong Tian 0001, Xilin Chen 0001, Hongkai Xiong, Li-Rong Dai 0001, Jing Chen 0002, Junliang Xing, Jing Chen 0003, Xihong Wu, Weiming Hu 0004, Yu Hu 0003, Tiejun Huang 0001, Wen Gao 0001 |
Frontiers Inf. Technol. Electron. Eng. | 10 |
| 2017 | Semi-Supervised Tensor-Based Graph Embedding Learning and Its Application to Visual Discriminant TrackingabstractAn appearance model adaptable to changes in object appearance is critical in visual object tracking. In this paper, we treat an image patch as a two-order tensor which preserves the original image structure. We design two graphs for characterizing the intrinsic local geometrical structure of the tensor samples of the object and the background. Graph embedding is used to reduce the dimensions of the tensors while preserving the structure of the graphs. Then, a discriminant embedding space is constructed. We prove two propositions for finding the transformation matrices which are used to map the original tensor samples to the tensor-based graph embedding space. In order to encode more discriminant information in the embedding space, we propose a transfer-learning- based semi-supervised strategy to iteratively adjust the embedding space into which discriminative information obtained from earlier times is transferred. We apply the proposed semi-supervised tensor-based graph embedding learning algorithm to visual tracking. The new tracking algorithm captures an object's appearance characteristics during tracking and uses a particle filter to estimate the optimal object state. Experimental results on the CVPR 2013 benchmark dataset demonstrate the effectiveness of the proposed tracking algorithm. Weiming Hu 0004, Junliang Xing, Chao Zhang 0089, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Multi-View Multi-Instance Learning Based on Joint Sparse Representation and Multi-View Dictionary LearningabstractIn multi-instance learning (MIL), the relations among instances in a bag convey important contextual information in many applications. Previous studies on MIL either ignore such relations or simply model them with a fixed graph structure so that the overall performance inevitably degrades in complex environments. To address this problem, this paper proposes a novel multi-view multi-instance learning algorithm (MIL) that combines multiple context structures in a bag into a unified framework. The novel aspects are: (i) we propose a sparse -graph model that can generate different graphs with different parameters to represent various context relations in a bag, (ii) we propose a multi-view joint sparse representation that integrates these graphs into a unified framework for bag classification, and (iii) we propose a multi-view dictionary learning algorithm to obtain a multi-view graph dictionary that considers cues from all views simultaneously to improve the discrimination of the MIL. Experiments and analyses in many practical applications prove the effectiveness of the M IL. Bing Li 0001, Chunfeng Yuan, Weihua Xiong, Weiming Hu 0004, Houwen Peng, Xinmiao Ding, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Salient Object Detection via Structured Matrix DecompositionabstractLow-rank recovery models have shown potential for salient object detection, where a matrix is decomposed into a low-rank matrix representing image background and a sparse matrix identifying salient objects. Two deficiencies, however, still exist. First, previous work typically assumes the elements in the sparse matrix are mutually independent, ignoring the spatial and pattern relations of image regions. Second, when the low-rank and sparse matrices are relatively coherent, e.g., when there are similarities between the salient objects and background or when the background is complicated, it is difficult for previous models to disentangle them. To address these problems, we propose a novel structured matrix decomposition model with two structural regularizations: (1) a tree-structured sparsity-inducing regularization that captures the image structure and enforces patches from the same object to have similar saliency values, and (2) a Laplacian regularization that enlarges the gaps between salient objects and the background in feature space. Furthermore, high-level priors are integrated to guide the matrix decomposition and boost the detection. We evaluate our model for salient object detection on five challenging datasets including single object, multiple objects and complex scene images, and show competitive results as compared with 24 state-of-the-art methods in terms of seven performance metrics. Houwen Peng, Bing Li 0001, Haibin Ling, Weiming Hu 0004, Weihua Xiong, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | D2C: Deep cumulatively and comparatively learning for human age estimation
Kai Li 0022, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank |
Pattern Recognit. | 3 |
| 2017 | Diagnosing deep learning models for high accuracy age estimation from a single image
Junliang Xing, Kai Li 0022, Weiming Hu 0004, Chunfeng Yuan, Haibin Ling |
Pattern Recognit. | 3 |
| 2016 | Tensor Power Iteration for Multi-graph MatchingabstractDue to its wide range of applications, matching between two graphs has been extensively studied and remains an active topic. By contrast, it is still under-exploited on how to jointly match multiple graphs, partly due to its intrinsic combinatorial intractability. In this work, we address this challenging problem in a principled way under the rank-1 tensor approximation framework. In particular, we formulate multi-graph matching as a combinational optimization problem with two main ingredients: unary matching over graph vertices and structure matching over graph edges, both of which across multiple graphs. Then we propose an efficient power iteration solution for the resulting NP-hard optimization problem. The proposed algorithm has several advantages: 1) the intrinsic matching consistency across multiple graphs based on the high-order tensor optimization, 2) the free employment of powerful high-order node affinity, 3) the flexible integration between various types of node affinities and edge/hyper-edge affinities. Experiments on diverse and challenging datasets validate the effectiveness of the proposed approach in comparison with state-of the-arts. Xinchu Shi, Haibin Ling, Weiming Hu 0004, Junliang Xing |
CVPR | 3 |
| 2016 | Graph Based Skeleton Motion Representation and Similarity Measurement for Action Recognition
Chunfeng Yuan, Weiming Hu 0004, Bing Li 0001, Yanning Zhang 0001 |
ECCV (7) | 3 |
| 2016 | Bootstrapping deep feature hierarchy for pornographic image recognitionabstractAutomatically recognizing pornographic images from the Web is a vital step to purify Internet environment. Inspired by the rapid developments of deep learning models, we present a deep architecture of convolutional neural network (CNN) for high accuracy pornographic image recognition. The proposed architecture is built upon existing CNNs which accepts input images of different sizes and incorporates features from different hierarchy to perform prediction. To effectively train the model, we propose a two-stage training strategy to learn the model parameters from scratch and end-to-end. During the training procedure, we also employ a hard negative sampling strategy to further reduce the false positive rate of the model. Experimental results on a large dataset demonstrate good performance of the proposed model and the effectiveness of our training strategies, with a considerable improvement over some traditional methods using hand-crafted features and deep learning method using mainstream CNN architecture. Kai Li 0022, Junliang Xing, Bing Li 0001, Weiming Hu 0004 |
ICIP | 4 |
| 2016 | Fast kernel SVM training via support vector identificationabstractTraining kernel SVM on large datasets suffers from high computational complexity and requires a large amount of memory. However, a desirable property of SVM is that its decision function is solely determined by the support vectors, a subset of training examples with non-vanishing weights. This motivates a novel efficient algorithm for training kernel SVM via support vector identification. The efficient training algorithm involves two steps. In the first step, we randomly sample the training data without replacement several times, each time a small subset of training data is sampled. Then a kernel SVM is trained on each subset, and the resulting kernel SVM models are used to identify the support vectors on the margin. In the second step, an optimization problem is solved to estimate the Lagrange multipliers corresponding to these support vectors. After obtaining the support vectors and Lagrange multipliers, we can approximate the decision function of kernel SVM. Due to the cubic complexity of standard kernel SVM training algorithm, training many kernel SVMs on small subsets of training data is much more efficient than training a single kernel SVM on the whole training data especially for large datasets. Therefore, our algorithm has better scalability than kernel SVM. Besides, training SVMs on each subset can be done independently, and hence our algorithm can be easily parallelized for further speedup. Since our algorithm only identifies the support vectors on the margin, it produces less number of support vectors as compared to that produced by standard kernel SVM. This makes our algorithm more efficient in prediction too. Experimental results show that our method outperforms state-of-the-art methods and achieves performance on par with the kernel SVM albeit with much improved efficiency. Zhouyu Fu, Ou Wu 0001, Weiming Hu 0004 |
ICPR | 4 |
| 2016 | Multi-Cue Illumination Estimation via a Tree-Structured Group Joint Sparse Representation
Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Brian V. Funt, Junliang Xing |
Int. J. Comput. Vis. | 3 |
| 2016 | Fusing ℝ Features and Local Features with Context-Aware Kernels for Action Recognition
Chunfeng Yuan, Baoxin Wu, Xi Li 0001, Weiming Hu 0004, Stephen J. Maybank, Fangshi Wang |
Int. J. Comput. Vis. | 4 |
| 2016 | Learning A Superpixel-Driven Speed Function for Level Set TrackingabstractA key problem in level set tracking is to construct a discriminative speed function for effective contour evolution. In this paper, we propose a level set tracking method based on a discriminative speed function, which produces a superpixel-driven force for effective level set evolution. Based on kernel density estimation and metric learning, the speed function is capable of effectively encoding the discriminative information on object appearance within a feasible metric space. Furthermore, we introduce adaptive object shape modeling into the level set evolution process, which leads to the tracking robustness in complex scenarios. To ensure the efficiency of adaptive object shape modeling, we develop a simple but efficient weighted non-negative matrix factorization method that can online learn an object shape dictionary. Experimental results on a number of challenging video sequences demonstrate the effectiveness and robustness of the proposed tracking method. Xi Li 0001, Weiming Hu 0004 |
IEEE Trans. Cybern. | 3 |
| 2016 | Listwise Learning to Rank from CrowdsabstractLearning to rank has received great attention in recent years as it plays a crucial role in many applications such as information retrieval and data mining. The existing concept of learning to rank assumes that each training instance is associated with a reliable label. However, in practice, this assumption does not necessarily hold true as it may be infeasible or remarkably expensive to obtain reliable labels for many learning to rank applications. Therefore, a feasible approach is to collect labels from crowds and then learn a ranking function from crowdsourcing labels. This study explores the listwise learning to rank with crowdsourcing labels obtained from multiple annotators, who may be unreliable. A new probabilistic ranking model is first proposed by combining two existing models. Subsequently, a ranking function is trained by proposing a maximum likelihood learning approach, which estimates ground-truth labels and annotator expertise, and trains the ranking function iteratively. In practical crowdsourcing machine learning, valuable side information (e.g., professional grades) about involved annotators is normally attainable. Therefore, this study also investigates learning to rank from crowd labels when side information on the expertise of involved annotators is available. In particular, three basic types of side information are investigated, and corresponding learning algorithms are consequently introduced. Further, the top-k learning to rank from crowdsourcing labels are explored to deal with long training ranking lists. The proposed algorithms are tested on both synthetic and real-world data. Results reveal that the maximum likelihood estimation approach significantly outperforms the average approach and existing crowdsourcing regression methods. The performances of the proposed algorithms are comparable to those of the learning model in consideration reliable labels. The results of the investigation further indicate that side information is helpful in inferring both ranking functions and expertise degrees of annotators. Ou Wu 0001, Qiang You, Fen Xia, Weiming Hu 0004 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2016 | Listwise Learning to Rank by Exploring Structure of ObjectsabstractListwise learning to rank (LTR) is aimed at constructing a ranking model from listwise training data to order objects. In most existing studies, each training instance consists of a set of objects described by preference features. In a preference feature space for the objects in training, the structure of the objects is associated with the absolute preference degrees for the objects. The degrees significantly influence the ordering of the objects. Nevertheless, the structure of the training objects in their preference feature space has rarely been studied. In addition, most listwise LTR algorithms yield a single linear ranking model for all objects, but this ranking model may be insufficient to capture the underlying nonlinear ranking mechanism among all objects. This study proposes a divide-and-train method to learn a nonlinear ranking model from listwise training data. First, a rank-preserving clustering approach is used to infer the structure of objects in their preference feature space and all the objects in training data are divided into several clusters. Each cluster is assumed to correspond to a preference degree and an ordinal regression function is then learned. Second, considering that relations exist among the clusters, a multi-task listwise ranking approach is then employed to train linear ranking functions for all the clusters (or preference degrees) simultaneously. Our proposed method utilizes both the (relative) preferences among objects and the intrinsic structure of objects. Experimental results on benchmark data sets suggest that the proposed method outperforms state-oft-the-art listwise LTR algorithms. Ou Wu 0001, Qiang You, Fen Xia, Fei Yuan 0003, Weiming Hu 0004 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2016 | Multi-Instance Multi-Label Learning Combining Hierarchical Context and its Application to Image AnnotationabstractIn image annotation, one image is often modeled as a bag of regions (“instances”) associated with multiple labels, which is a typical application of multi-instance multi-label learning (MIML). Although lots of research has shown that the interplay embedded among instances and labels can largely boost the image annotation accuracy, most existing MIML methods consider none or partial context cues. In this paper, we propose a novel context-aware MIML model to integrate the instance context and label context into a general framework. Specially, the instance context is constructed with multiple graphs, while the label context is built up through a linear combination of several common latent conceptions that link low level features and high level semantic labels. Comparison with other leading methods on several benchmark datasets in terms of image annotation shows that our proposed method can get better performance than the state-of-the-art approaches. Xinmiao Ding, Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Bo Wang 0147 |
IEEE Trans. Multim. | 5 |
| 2016 | Multi-Perspective Cost-Sensitive Context-Aware Multi-Instance Sparse Coding and Its Application to Sensitive Video RecognitionabstractWith the development of video-sharing websites, P2P, micro-blog, mobile WAP websites, and so on, sensitive videos can be more easily accessed. Effective sensitive video recognition is necessary for web content security. Among web sensitive videos, this paper focuses on violent and horror videos. Based on color emotion and color harmony theories, we extract visual emotional features from videos. A video is viewed as a bag and each shot in the video is represented by a key frame which is treated as an instance in the bag. Then, we combine multi-instance learning (MIL) with sparse coding to recognize violent and horror videos. The resulting MIL-based model can be updated online to adapt to changing web environments. We propose a cost-sensitive context-aware multi- instance sparse coding (MI-SC) method, in which the contextual structure of the key frames is modeled using a graph, and fusion between audio and visual features is carried out by extending the classic sparse coding into cost-sensitive sparse coding. We then propose a multi-perspective multi- instance joint sparse coding (MI-J-SC) method that handles each bag of instances from an independent perspective, a contextual perspective, and a holistic perspective. The experiments demonstrate that the features with an emotional meaning are effective for violent and horror video recognition, and our cost-sensitive context-aware MI-SC and multi-perspective MI-J-SC methods outperform the traditional MIL methods and the traditional SVM and KNN-based methods. Weiming Hu 0004, Xinmiao Ding, Bing Li 0001, Fangshi Wang, Stephen J. Maybank |
IEEE Trans. Multim. | 1 |
| 2016 | Multimodal Web Aesthetics Assessment Based on Structural SVM and Multitask Fusion LearningabstractThe overall visual attributes (e.g., aesthetics) of Web pages significantly influence user experience. A beautiful and well laid out Web page greatly facilitates user access and enhances the browsing experience. In this paper, a new method is proposed to learn an assessment model for the (visual) aesthetics of Web pages. First, multimodal features (structural, local visual, global visual, and functional) of a Web page that are known to significantly affect the aesthetics of a Web page are extracted to construct a feature vector. Second, the interuser disagreement of aesthetics is analyzed and novel aesthetic representations are obtained from the multiuser ratings of a page. A structural learning algorithm is proposed for the new aesthetic representations. Third, as a Web page's functional purpose also affects the perceived aesthetics, we divide Web pages into different types using functional features, and a soft multitask fusion learning strategy is introduced to train assessment models for pages with functional purposes. Experimental results show the effectiveness of our method: 1) the combination of structural, local, and global visual features outperforms existing state-of-the-art Web aesthetic features; 2) the proposed structural learning algorithm achieves good results for the new aesthetic representations; and 3) the proposed soft multitask fusion learning strategy improves the performances of aesthetics assessment models. Ou Wu 0001, Haiqiang Zuo, Weiming Hu 0004, Bing Li 0001 |
IEEE Trans. Multim. | 3 |
| 2015 | Local Subspace Collaborative TrackingabstractSubspace models have been widely used for appearance based object tracking. Most existing subspace based trackers employ a linear subspace to represent object appearances, which are not accurate enough to model large variations of objects. To address this, this paper presents a local subspace collaborative tracking method for robust visual tracking, where multiple linear and nonlinear subspaces are learned to better model the nonlinear relationship of object appearances. First, we retain a set of key samples and compute a set of local subspaces for each key sample. Then, we construct a hyper sphere to represent the local nonlinear subspace for each key sample. The hyper sphere of one key sample passes the local key samples and also is tangent to the local linear subspace of the specific key sample. In this way, we are able to represent the nonlinear distribution of the key samples and also approximate the local linear subspace near the specific key sample, so that local distributions of the samples can be represented more accurately. Experimental results on challenging video sequences demonstrate the effectiveness of our method. Xiaoqin Zhang 0002, Weiming Hu 0004, Junliang Xing, Jiwen Lu, Jie Zhou 0001 |
ICCV | 3 |
| 2015 | Load-balanced locality-sensitive hashing: A new method for efficient near duplicate image detectionabstractLocality-Sensitive Hashing (LSH) is a mainstream method for the Near Duplicate Image Detection (NDID) problem. Previous LSH based methods, however, do not have a principled way to make the indexing structure generate the buckets of similar sizes, which will inevitably degrade the detection effectiveness and efficiency. In this work, we propose a Load-Balanced Locality-Sensitive Hashing (LBLSH) method with a new indexing structure to produce load-balanced buckets for the hashing process. As proved in the paper, the proposed LBLSH can guarantee load-balanced buckets in the hashing process and significantly reduce the query time and the storage space. Based on the proposed LBLSH method, we design an effective and feasible algorithm for the NDID problem. Extensive experiments on two benchmark datasets demonstrate the effectiveness and efficiency of our method. Yabo Fan, Junliang Xing, Weiming Hu 0004 |
ICIP | 3 |
| 2015 | Robust visual tracking using joint scale-spatial correlation filtersabstractScale adaptation is crucial to object tracking as the visual size of the target changes continuously. Many existing tracking algorithms, however, simply ignore scale changes either for the consideration of tracking efficiency or the lack of principle ways to scale estimation. In this work, we present an efficient and effective scale adaptive tracking algorithm by proposing a correlation filter based tracker in the joint spatial and scale space. We find that the exhaustive template searching in this joint space can be well modeled by a block-circulant matrix. With the properties of the block-circulant matrices, we prove that the expensive template matching can be transformed to efficient dot product in frequency domain by fast Fourier Transform. Based on these findings, our new tracker significantly improves the robustness and adaptability of previous competitive spatial correlation trackers. On the latest single object tracking benchmark, our tracker advances the state-of-the-art tracking results with a very large margin. Mengdan Zhang, Junliang Xing, Weiming Hu 0004 |
ICIP | 4 |
| 2015 | Optimizing Locally Linear Classifiers with Supervised Anchor Point Learning
Zhouyu Fu, Ou Wu 0001, Weiming Hu 0004 |
IJCAI | 4 |
| 2015 | Predicting Image Memorability by Multi-view Adaptive RegressionabstractThe images we encounter throughout our lives make different impressions on us: Some are remembered at first glance, while others are forgotten. This phenomenon is caused by the intrinsic memorability of images revealed by recent studies [5,6]. In this paper, we address the issue of automatically estimating the memorability of images by proposing a novel multi-view adaptive regression (MAR) model. The MAR model provides an effective mapping of visual features to memorability scores by taking advantage of robust feature selection and multiple feature integration. It consists of three major components: an adaptive loss function, an adaptive regularization and a multi-view modeling strategy. Moreover, we design an alternating direction method (ADM) optimization algorithm to solve the proposed objective function. Experimental results on the MIT benchmark dataset show the superiority of the proposed model compared with existing image memorability prediction methods. Houwen Peng, Kai Li 0022, Bing Li 0001, Haibin Ling, Weihua Xiong, Weiming Hu 0004 |
ACM Multimedia | 6 |
| 2015 | A Robust Tracking System for Low Frame Rate Video
Xiaoqin Zhang 0002, Weiming Hu 0004, Nianhua Xie, Hujun Bao, Stephen J. Maybank |
Int. J. Comput. Vis. | 2 |
| 2015 | Erratum to: A Robust Tracking System for Low Frame Rate Video
Xiaoqin Zhang 0002, Weiming Hu 0004, Nianhua Xie, Hujun Bao, Stephen J. Maybank |
Int. J. Comput. Vis. | 2 |
| 2015 | Single and Multiple Object Tracking Using a Multi-Feature Joint Sparse RepresentationabstractIn this paper, we propose a tracking algorithm based on a multi-feature joint sparse representation. The templates for the sparse representation can include pixel values, textures, and edges. In the multi-feature joint optimization, noise or occlusion is dealt with using a set of trivial templates. A sparse weight constraint is introduced to dynamically select the relevant templates from the full set of templates. A variance ratio measure is adopted to adaptively adjust the weights of different features. The multi-feature template set is updated adaptively. We further propose an algorithm for tracking multi-objects with occlusion handling based on the multi-feature joint sparse reconstruction. The observation model based on sparse reconstruction automatically focuses on the visible parts of an occluded object by using the information in the trivial templates. The multi-object tracking is simplified into a joint Bayesian inference. The experimental results show the superiority of our algorithm over several state-of-the-art tracking algorithms. Weiming Hu 0004, Wei Li 0034, Xiaoqin Zhang 0002, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Horror Image Recognition Based on Context-Aware Multi-Instance LearningabstractHorror content sharing on the Web is a growing phenomenon that can interfere with our daily life and affect the mental health of those involved. As an important form of expression, horror images have their own characteristics that can evoke extreme emotions. In this paper, we present a novel context-aware multi-instance learning (CMIL) algorithm for horror image recognition. The CMIL algorithm identifies horror images and picks out the regions that cause the sensation of horror in these horror images. It obtains contextual cues among adjacent regions in an image using a random walk on a contextual graph. Borrowing the strength of the fuzzy support vector machine (FSVM), we define a heuristic optimization procedure based on the FSVM to search for the optimal classifier for the CMIL. To improve the initialization of the CMIL, we propose a novel visual saliency model based on the tensor analysis. The average saliency value of each segmented region is set as its initial fuzzy membership in the CMIL. The advantage of the tensor-based visual saliency model is that it not only adaptively selects features, but also dynamically determines fusion weights for saliency value combination from different feature subspaces. The effectiveness of the proposed CMIL model is demonstrated by its use in horror image recognition on two large-scale image sets collected from the Internet. Bing Li 0001, Weihua Xiong, Ou Wu 0001, Weiming Hu 0004, Stephen J. Maybank, Shuicheng Yan |
IEEE Trans. Image Process. | 4 |
| 2014 | Quality-Based Learning for Web Data ClassificationabstractThe types of web data vary in terms of information quantity and quality. For example, some pages contain numerous texts, whereas some others contain few texts; some web videos are in high resolution, whereas some other web videos are in low resolution. As a consequence, the quality of extracted features from different web data may also vary greatly. Existing learning algorithms on web data classification usually ignore the variations of information quality or quantity. In this paper, the information quantity and quality of web data are described by quality-related factors such as text length and image quantity, and a new learning method is proposed to train classifiers based on quality-related factors. The method divides training data into subsets according to the clustering results of quality-related factors and then trains classifiers by using a multi-task learning strategy for each subset. Experimental results indicate that the quality-related factors are useful in web data classification, and the proposed method outperforms conventional algorithms that do not consider information quantity and quality. Ou Wu 0001, Ruiguang Hu, Weiming Hu 0004 |
AAAI | 4 |
| 2014 | Nonlinear Classification via Linear SVMs and Multi-Task LearningabstractKernel SVM is prohibitively expensive when dealing with large nonlinear data. While ensembles of linear classifiers have been proposed to address this inefficiency, these methods are time-consuming or lack robustness. We propose an efficient classifier for nonlinear data using a new iterative learning algorithm, which partitions the data into clusters, and then trains a linear SVM for each cluster. These two steps are combined into a graphical model, with the parameters estimated efficiently using the EM algorithm. During training, clustered multi-task learning is used to capture the relatedness among the multiple linear SVMs and avoid overfitting. Experimental results on benchmark datasets show that our method outperforms state-of-the-art methods. During prediction, it also obtains comparable classification performance to kernel SVM, with much higher efficiency. Ou Wu 0001, Weiming Hu 0004, Peter O'Donovan |
CIKM | 3 |
| 2014 | Multi-target Tracking with Motion Context in Tensor Power IterationabstractInteractions between moving targets often provide discriminative clues for multiple target tracking (MTT), though many existing approaches ignore such interactions due to difficulty in effectively handling them. In this paper, we model interactions between neighbor targets by pair-wise motion context, and further encode such context into the global association optimization. To solve the resulting global non-convex maximization, we propose an effective and efficient power iteration framework. This solution enjoys two advantages for MTT: First, it allows us to combine the global energy accumulated from individual trajectories and the between-trajectory interaction energy into a united optimization, which can be solved by the proposed power iteration algorithm. Second, the framework is flexible to accommodate various types of pairwise context models and we in fact studied two different context models in this paper. For evaluation, we apply the proposed methods to four public datasets involving different challenging scenarios such as dense aerial borne traffic tracking, dense point set tracking, and semi-crowded pedestrian tracking. In all the experiments, our approaches demonstrate very promising results in comparison with state-of-the-art trackers. Xinchu Shi, Haibin Ling, Weiming Hu 0004, Chunfeng Yuan, Junliang Xing |
CVPR | 3 |
| 2014 | Towards Multi-view and Partially-Occluded Face AlignmentabstractWe present a robust model to locate facial landmarks under different views and possibly severe occlusions. To build reliable relationships between face appearance and shape with large view variations, we propose to formulate face alignment as an l1-induced Stagewise Relational Dictionary (SRD) learning problem. During each training stage, the SRD model learns a relational dictionary to capture consistent relationships between face appearance and shape, which are respectively modeled by the pose-indexed image features and the shape displacements for current estimated landmarks. During testing, the SRD model automatically selects a sparse set of the most related shape displacements for the testing face and uses them to refine its shape iteratively. To locate facial landmarks under occlusions, we further propose to learn an occlusion dictionary to model different kinds of partial face occlusions. By deploying the occlusion dictionary into the SRD model, the alignment performance for occluded faces can be further improved. Our algorithm is simple, effective, and easy to implement. Extensive experiments on two benchmark datasets and two newly built datasets have demonstrated its superior performances over the state-of-the-art methods, especially for faces with large view variations and/or occlusions. Junliang Xing, Zhiheng Niu, Junshi Huang, Weiming Hu 0004, Shuicheng Yan |
CVPR | 4 |
| 2014 | Transfer Learning Based Visual Tracking with Gaussian Processes Regression
Haibin Ling, Weiming Hu 0004, Junliang Xing |
ECCV (3) | 3 |
| 2014 | RGBD Salient Object Detection: A Benchmark and Algorithms
Houwen Peng, Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Rongrong Ji |
ECCV (3) | 4 |
| 2014 | Learning from Multi-User Multi-Attribute AnnotationsabstractMining the data source from a crowd of people has elicited increasing attention in recent years. In existing studies, multiple users are utilized, in which each user is generally required to annotate only one attribute for each sample. However, there are cases in numerous annotation tasks wherein despite of the presence of multiple users, each user should classify or rate multiple attributes for each sample. This situation is referred to as multi-user multi-attribute annotations in this paper. This work deals with the learning problem under multi-user multi-attribute annotations. A generative model is introduced to describe the human labeling process for multi-user multi-attribute annotations. Subsequently, a maximum likelihood approach is leveraged to infer the parameters in the generative model, namely, ground-truth labels, user expertise, and annotation difficulties. The classifiers for each attribute are also learned simultaneously. Furthermore, the correlations among attributes are taken into account during inference and learning using conditional random field. The experimental results reveal that compared with existing methods that ignore the characteristics of multi-user multi-attribute annotations, our approach can obtain better estimation of the ground truth labels, user experts, annotation difficulties as well as attribute classifiers. Ou Wu 0001, Shuxiao Li, Honghui Dong, Ying Chen 0018, Weiming Hu 0004 |
SDM | 5 |
| 2014 | Bin Ratio-Based Histogram Distances and Their Application to Image ClassificationabstractLarge variations in image background may cause partial matching and normalization problems for histogram-based representations, i.e., the histograms of the same category may have bins which are significantly different, and normalization may produce large changes in the differences between corresponding bins. In this paper, we deal with this problem by using the ratios between bin values of histograms, rather than bin values' differences which are used in the traditional histogram distances. We propose a bin ratio-based histogram distance (BRD), which is an intra-cross-bin distance, in contrast with previous bin-to-bin distances and cross-bin distances. The BRD is robust to partial matching and histogram normalization, and captures correlations between bins with only a linear computational complexity. We combine the BRD with the ℓ1 histogram distance and the χ(2) histogram distance to generate the ℓ1 BRD and the χ(2) BRD, respectively. These combinations exploit and benefit from the robustness of the BRD under partial matching and the robustness of the ℓ1 and χ(2) distances to small noise. We propose a method for assessing the robustness of histogram distances to partial matching. The BRDs and logistic regression-based histogram fusion are applied to image classification. The experimental results on synthetic data sets show the robustness of the BRDs to partial matching, and the experiments on seven benchmark data sets demonstrate promising results of the BRDs for image classification. Weiming Hu 0004, Nianhua Xie, Ruiguang Hu, Haibin Ling, Qiang Chen 0007, Shuicheng Yan, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Learning Human Actions by Combining Global Dynamics and Local AppearanceabstractIn this paper, we address the problem of human action recognition through combining global temporal dynamics and local visual spatio-temporal appearance features. For this purpose, in the global temporal dimension, we propose to model the motion dynamics with robust linear dynamical systems (LDSs) and use the model parameters as motion descriptors. Since LDSs live in a non-Euclidean space and the descriptors are in non-vector form, we propose a shift invariant subspace angles based distance to measure the similarity between LDSs. In the local visual dimension, we construct curved spatio-temporal cuboids along the trajectories of densely sampled feature points and describe them using histograms of oriented gradients (HOG). The distance between motion sequences is computed with the Chi-Squared histogram distance in the bag-of-words framework. Finally we perform classification using the maximum margin distance learning method by combining the global dynamic distances and the local visual distances. We evaluate our approach for action recognition on five short clips data sets, namely Weizmann, KTH, UCF sports, Hollywood2 and UCF50, as well as three long continuous data sets, namely VIRAT, ADL and CRIM13. We show competitive results as compared with current state-of-the-art methods. Guan Luo, Guodong Tian, Chunfeng Yuan, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2014 | Online Adaboost-Based Parameterized Methods for Dynamic Distributed Network Intrusion DetectionabstractCurrent network intrusion detection systems lack adaptability to the frequently changing network environments. Furthermore, intrusion detection in the new distributed architectures is now a major requirement. In this paper, we propose two online Adaboost-based intrusion detection algorithms. In the first algorithm, a traditional online Adaboost process is used where decision stumps are used as weak classifiers. In the second algorithm, an improved online Adaboost process is proposed, and online Gaussian mixture models (GMMs) are used as weak classifiers. We further propose a distributed intrusion detection framework, in which a local parameterized detection model is constructed in each node using the online Adaboost algorithm. A global detection model is constructed in each node by combining the local parametric models using a small number of samples in the node. This combination is achieved using an algorithm based on particle swarm optimization (PSO) and support vector machines. The global model in each node is used to detect intrusions. Experimental results show that the improved online Adaboost process with GMMs obtains a higher detection rate and a lower false alarm rate than the traditional online Adaboost process that uses decision stumps. Both the algorithms outperform existing intrusion detection algorithms. It is also shown that our PSO, and SVM-based algorithm effectively combines the local detection models into the global model in each node; the global model in a node can handle the intrusion types that are found in other nodes, without sharing the samples of these intrusion types. Weiming Hu 0004, Yanguo Wang, Ou Wu 0001, Stephen J. Maybank |
IEEE Trans. Cybern. | 1 |
| 2014 | Image Classification Using Multiscale Information Fusion Based on Saliency Driven Nonlinear Diffusion FilteringabstractIn this paper, we propose saliency driven image multiscale nonlinear diffusion filtering. The resulting scale space in general preserves or even enhances semantically important structures such as edges, lines, or flow-like structures in the foreground, and inhibits and smoothes clutter in the background. The image is classified using multiscale information fusion based on the original image, the image at the final scale at which the diffusion process converges, and the image at a midscale. Our algorithm emphasizes the foreground features, which are important for image classification. The background image regions, whether considered as contexts of the foreground or noise to the foreground, can be globally handled by fusing information from different scales. Experimental tests of the effectiveness of the multiscale space for the image classification are conducted on the following publicly available datasets: 1) the PASCAL 2005 dataset; 2) the Oxford 102 flowers dataset; and 3) the Oxford 17 flowers dataset, with high classification rates. Weiming Hu 0004, Ruiguang Hu, Nianhua Xie, Haibin Ling, Stephen J. Maybank |
IEEE Trans. Image Process. | 1 |
| 2014 | Evaluating Combinational Illumination Estimation Methods on Real-World ImagesabstractIllumination estimation is an important component of color constancy and automatic white balancing. A number of methods of combining illumination estimates obtained from multiple subordinate illumination estimation methods now appear in the literature. These combinational methods aim to provide better illumination estimates by fusing the information embedded in the subordinate solutions. The existing combinational methods are surveyed and analyzed here with the goals of determining: 1) the effectiveness of fusing illumination estimates from multiple subordinate methods; 2) the best method of combination; 3) the underlying factors that affect the performance of a combinational method; and 4) the effectiveness of combination for illumination estimation in multiple-illuminant scenes. The various combinational methods are categorized in terms of whether or not they require supervised training and whether or not they rely on high-level scene content cues (e.g., indoor versus outdoor). Extensive tests and enhanced analyzes using three data sets of real-world images are conducted. For consistency in testing, the images were labeled according to their high-level features (3D stages, indoor/outdoor) and this label data is made available on-line. The tests reveal that the trained combinational methods (direct combination by support vector regression in particular) clearly outperform both the non-combinational methods and those combinational methods based on scene content cues. Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Brian V. Funt |
IEEE Trans. Image Process. | 3 |
| 2014 | Action Recognition Using Nonnegative Action Component Representation and Sparse Basis SelectionabstractIn this paper, we propose using high-level action units to represent human actions in videos and, based on such units, a novel sparse model is developed for human action recognition. There are three interconnected components in our approach. First, we propose a new context-aware spatial-temporal descriptor, named locally weighted word context, to improve the discriminability of the traditionally used local spatial-temporal descriptors. Second, from the statistics of the context-aware descriptors, we learn action units using the graph regularized nonnegative matrix factorization, which leads to a part-based representation and encodes the geometrical information. These units effectively bridge the semantic gap in action recognition. Third, we propose a sparse model based on a joint l2,1-norm to preserve the representative items and suppress noise in the action units. Intuitively, when learning the dictionary for action representation, the sparse model captures the fact that actions from the same class share similar units. The proposed approach is evaluated on several publicly available data sets. The experimental results and analysis clearly demonstrate the effectiveness of the proposed approach. Haoran Wang 0001, Chunfeng Yuan, Weiming Hu 0004, Haibin Ling, Wankou Yang, Changyin Sun 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | Modeling Geometric-Temporal Context With Directional Pyramid Co-Occurrence for Action RecognitionabstractIn this paper, we present a new geometric-temporal representation for visual action recognition based on local spatio-temporal features. First, we propose a modified covariance descriptor under the log-Euclidean Riemannian metric to represent the spatio-temporal cuboids detected in the video sequences. Compared with previously proposed covariance descriptors, our descriptor can be measured and clustered in Euclidian space. Second, to capture the geometric-temporal contextual information, we construct a directional pyramid co-occurrence matrix (DPCM) to describe the spatio-temporal distribution of the vector-quantized local feature descriptors extracted from a video. DPCM characterizes the co-occurrence statistics of local features as well as the spatio-temporal positional relationships among the concurrent features. These statistics provide strong descriptive power for action recognition. To use DPCM for action recognition, we propose a directional pyramid co-occurrence matching kernel to measure the similarity of videos. The proposed method achieves the state-of-the-art performance and improves on the recognition performance of the bag-of-visual-words (BOVWs) models by a large margin on six public data sets. For example, on the KTH data set, it achieves 98.78% accuracy while the BOVW approach only achieves 88.06%. On both Weizmann and UCF CIL data sets, the highest possible accuracy of 100% is achieved. Chunfeng Yuan, Xi Li 0001, Weiming Hu 0004, Haibin Ling, Stephen J. Maybank |
IEEE Trans. Image Process. | 3 |
| 2014 | Context-Aware Hypergraph Construction for Robust Spectral ClusteringabstractSpectral clustering is a powerful tool for unsupervised data analysis. In this paper, we propose a context-aware hypergraph similarity measure (CAHSM), which leads to robust spectral clustering in the case of noisy data. We construct three types of hypergraphs-the pairwise hypergraph, the k-nearest-neighbor (kNN) hypergraph, and the high-order over-clustering hypergraph. The pairwise hypergraph captures the pairwise similarity of data points; the kNNhypergraph captures the neighborhood of each point; and the clustering hypergraph encodes high-order contexts within the dataset. By combining the affinity information from these three hypergraphs, the CAHSM algorithm is able to explore the intrinsic topological information of the dataset. Therefore, data clustering using CAHSM tends to be more robust. Considering the intra-cluster compactness and the inter-cluster separability of vertices, we further design a discriminative hypergraph partitioning criterion (DHPC). Using both CAHSM and DHPC, a robust spectral clustering algorithm is developed. Theoretical analysis and experimental evaluation demonstrate the effectiveness and robustness of the proposed algorithm. Xi Li 0001, Weiming Hu 0004, Chunhua Shen, Anthony R. Dick, Zhongfei Zhang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | Human Pose Estimation and Tracking via Parsing a Tree Structure Based Human ModelabstractHuman pose estimation and tracking is the task of determining the states (location, orientation, and scale) of each body part over time. It is important for many vision understanding applications, such as visual interactive gaming, immersive virtual reality, visual surveillance, and content-based image retrieval. However, it remains a challenging task due to unknown image background, presence of clutter and especially the high dimensional state space (usually 30+ dimensions). In this paper, we contribute to human pose estimation and tracking in two aspects. First, we design two efficient Markov Chain dynamics under the data-driven Markov Chain Monte Carlo framework to effectively explore the high dimensional state space. Second, we parse the tree structure state space into a lexicographic order according to the image observations and body topology, and the optimization process is conducted in this order. This realizes a much more efficient exploration of the state space than the sampling based search or exhaustive search, and thus achieves a tremendous speed-up. Experimental results demonstrate the efficiency and effectiveness of the proposed method in estimating and tracking various kinds of human poses, even against cluttered backgrounds, in poor illumination or under partial self-occlusion. Xiaoqin Zhang 0002, Weiming Hu 0004, Xiaofeng Tong, Stephen J. Maybank, Yimin Zhang 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2013 | Salient Object Detection via Low-Rank and Structured Sparse Matrix DecompositionabstractSalient object detection provides an alternative solution to various image semantic understanding tasks such as object recognition, adaptive compression and image retrieval. Recently, low-rank matrix recovery (LR) theory has been introduced into saliency detection, and achieves impressed results. However, the existing LR-based models neglect the underlying structure of images, and inevitably degrade the associated performance. In this paper, we propose a Low-rank and Structured sparse Matrix Decomposition (LSMD) model for salient object detection. In the model, a tree-structured sparsity-inducing norm regularization is firstly introduced to provide a hierarchical description of the image structure to ensure the completeness of the extracted salient object. The similarity of saliency values within the salient object is then guaranteed by the $\ell _\infty$-norm. Finally, high-level priors are integrated to guide the matrix decomposition and enhance the saliency detection. Experimental results on the largest public benchmark database show that our model outperforms existing LR-based approaches and other state-of-the-art methods, which verifies the effectiveness and robustness of the structure cues in our model. Houwen Peng, Bing Li 0001, Rongrong Ji, Weiming Hu 0004, Weihua Xiong, Congyan Lang |
AAAI | 4 |
| 2013 | Illumination Estimation Based on Bilayer Sparse CodingabstractComputational color constancy is a very important topic in computer vision and has attracted many researchers' attention. Recently, lots of research has shown the effects of using high level visual content cues for improving illumination estimation. However, nearly all the existing methods are essentially combinational strategies in which image's content analysis is only used to guide the combination or selection from a variety of individual illumination estimation methods. In this paper, we propose a novel bilayer sparse coding model for illumination estimation that considers image similarity in terms of both low level color distribution and high level image scene content simultaneously. For the purpose, the image's scene content information is integrated with its color distribution to obtain optimal illumination estimation model. The experimental results on real-world image sets show that our algorithm is superior to some prevailing illumination estimation methods, even better than some combinational methods. Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Houwen Peng |
CVPR | 3 |
| 2013 | Multi-target Tracking by Rank-1 Tensor ApproximationabstractIn this paper we formulate multi-target tracking (MTT) as a rank-1 tensor approximation problem and propose an ℓ1norm tensor power iteration solution. In particular, a high order tensor is constructed based on trajectories in the time window, with each tensor element as the affinity of the corresponding trajectory candidate. The local assignment variables are the ℓ1normalized vectors, which are used to approximate the rank-1 tensor. Our approach provides a flexible and effective formulation where both pairwise and high-order association energies can be used expediently. We also show the close relation between our formulation and the multi-dimensional assignment (MDA) model. To solve the optimization in the rank-1 tensor approximation, we propose an algorithm that iteratively powers the intermediate solution followed by an ℓ1normalization. Aside from effectively capturing high-order motion information, the proposed solver runs efficiently with proved convergence. The experimental validations are conducted on two challenging datasets and our method demonstrates promising performances on both. Xinchu Shi, Haibin Ling, Junliang Xing, Weiming Hu 0004 |
CVPR | 4 |
| 2013 | Multi-task Sparse Learning with Beta Process Prior for Action RecognitionabstractIn this paper, we formulate human action recognition as a novel Multi-Task Sparse Learning(MTSL) framework which aims to construct a test sample with multiple features from as few bases as possible. Learning the sparse representation under each feature modality is considered as a single task in MTSL. Since the tasks are generated from multiple features associated with the same visual input, they are not independent but inter-related. We introduce a Beta process(BP) prior to the hierarchical MTSL model, which efficiently learns a compact dictionary and infers the sparse structure shared across all the tasks. The MTSL model enforces the robustness in coefficient estimation compared with performing each task independently. Besides, the sparseness is achieved via the Beta process formulation rather than the computationally expensive L1 norm penalty. In terms of non-informative gamma hyper-priors, the sparsity level is totally decided by the data. Finally, the learning problem is solved by Gibbs sampling inference which estimates the full posterior on the model parameters. Experimental results on the KTH and UCF sports datasets demonstrate the effectiveness of the proposed MTSL approach for action recognition. Chunfeng Yuan, Weiming Hu 0004, Guodong Tian, Haoran Wang 0001 |
CVPR | 2 |
| 2013 | 3D R Transform on Spatio-temporal Interest Points for Action RecognitionabstractSpatio-temporal interest points serve as an elementary building block in many modern action recognition algorithms, and most of them exploit the local spatio-temporal volume features using a Bag of Visual Words (BOVW) representation. Such representation, however, ignores potentially valuable information about the global spatio-temporal distribution of interest points. In this paper, we propose a new global feature to capture the detailed geometrical distribution of interest points. It is calculated by using the R transform which is defined as an extended 3D discrete Radon transform, followed by applying a two-directional two-dimensional principal component analysis. Such R feature captures the geometrical information of the interest points and keeps invariant to geometry transformation and robust to noise. In addition, we propose a new fusion strategy to combine the R feature with the BOVW representation for further improving recognition accuracy. We utilize a context-aware fusion method to capture both the pairwise similarities and higher-order contextual interactions of the videos. Experimental results on several publicly available datasets demonstrate the effectiveness of the proposed approach for action recognition. Chunfeng Yuan, Xi Li 0001, Weiming Hu 0004, Haibin Ling, Stephen J. Maybank |
CVPR | 3 |
| 2013 | Distance Map of Various Weights: A new feature for adaptive object trackingabstractIn this paper, we propose a new feature, Distance Map of Various Weights (DMVW) based on distances between rows' textures, to perform tracking. The proposed new feature provides an effective object appearance model which is both illumination-invariant and robust to occlusion. We also develop a 2D PCA based method to effectively evaluate the new feature. We demonstrate the validity of the rows' or column's weights in computing 2D PCA subspaces. To balance the importance of local and global information, we define a coefficient to revise the locality extent of the proposed feature. A new method based on entropy of candidate state evaluation is proposed to select the most discriminative coefficient. Experimental results on challenging video sequences demonstrated the effectiveness of our method. Junliang Xing, Xiaoqin Zhang 0002, Weiming Hu 0004 |
ICASSP | 4 |
| 2013 | Adaptive cooperative tracking based on multi-graph embedding and Markov Random FieldabstractAppearance model is of fundamental importance in a tracking algorithm. In this paper, we propose a new tracking method based on a cooperative object appearance model which incorporates both the discriminative and generative information. We represent the discriminative information with graph embedding (GE). To represent the local object appearance effectively, we divide the object and nearby background into patches. As the discriminative conditions around the 4 object boundaries are different, we divide the patches into 4 groups and perform GE for each group. Markov Random Filed (MRF) is designed to represent the generative information. We propose a novel MRF based method which not only considers the single patch's appearance but also the appearance relations between neighbor patches (not the relations between neighbor patches' states). The proposed cooperative appearance model can represent the object appearance's variation effectively and meanwhile discriminate the object from background robustly. Experimental results on challenging test sequences demonstrated the effectiveness of our method. Junliang Xing, Xiaoqin Zhang 0002, Weiming Hu 0004 |
ICASSP | 4 |
| 2013 | Combining sparse appearance features and dense motion features via random forest for action detectionabstractThis paper presents a new method to detect human actions in video by combining sparse appearance features and dense motion features in the unified random forest framework. We compute sparse appearance features to capture the main appearance changes and dense motion features to capture the tiny motion changes in the video. We take advantage of the randomization of channel selection in random trees to combine these two complementary types of features. In addition, linear classification is applied to grow each tree with high efficiency. Each leaf in these trees stores the class distribution and location information of the training samples and action detection for the test video is accomplished by Hough voting of the leaves in each tree. Experimental results demonstrate that our method achieves the state-of-the-art performance on two datasets. Chunfeng Yuan, Haoran Wang 0001, Weiming Hu 0004 |
ICASSP | 4 |
| 2013 | Discriminant Tracking Using Tensor Representation with Semi-supervised ImprovementabstractVisual tracking has witnessed growing methods in object representation, which is crucial to robust tracking. The dominant mechanism in object representation is using image features encoded in a vector as observations to perform tracking, without considering that an image is intrinsically a matrix, or a 2^nd-order tensor. Thus approaches following this mechanism inevitably lose a lot of useful information, and therefore cannot fully exploit the spatial correlations within the 2D image ensembles. In this paper, we address an image as a 2^nd-order tensor in its original form, and find a discriminative linear embedding space approximation to the original nonlinear sub manifold embedded in the tensor space based on the graph embedding framework. We specially design two graphs for characterizing the intrinsic local geometrical structure of the tensor space, so as to retain more discriminant information when reducing the dimension along certain tensor dimensions. However, spatial correlations within a tensor are not limited to the elements along these dimensions. This means that some part of the discriminant information may not be encoded in the embedding space. We introduce a novel technique called semi-supervised improvement to iteratively adjust the embedding space to compensate for the loss of discriminant information, hence improving the performance of our tracker. Experimental results on challenging videos demonstrate the effectiveness and robustness of the proposed tracker. Junliang Xing, Weiming Hu 0004, Stephen J. Maybank |
ICCV | 3 |
| 2013 | Robust Object Tracking with Online Multi-lifespan Dictionary LearningabstractRecently, sparse representation has been introduced for robust object tracking. By representing the object sparsely, i.e., using only a few templates via L1-norm minimization, these so-called L1-trackers exhibit promising tracking results. In this work, we address the object template building and updating problem in these L1-tracking approaches, which has not been fully studied. We propose to perform template updating, in a new perspective, as an online incremental dictionary learning problem, which is efficiently solved through an online optimization procedure. To guarantee the robustness and adaptability of the tracking algorithm, we also propose to build a multi-lifespan dictionary model. By building target dictionaries of different life spans, effective object observations can be obtained to deal with the well-known drifting problem in tracking and thus improve the tracking accuracy. We derive effective observation models both generatively and discriminatively based on the online multi-lifespan dictionary learning model and deploy them to the Bayesian sequential estimation framework to perform tracking. The proposed approach has been extensively evaluated on ten challenging video sequences. Experimental results demonstrate the effectiveness of the online learned templates, as well as the state-of-the-art tracking performance of the proposed approach. Junliang Xing, Bing Li 0001, Weiming Hu 0004, Shuicheng Yan |
ICCV | 4 |
| 2013 | Learning silhouette dynamics for human action recognitionabstractIn this paper, we address the problem of recognizing human actions with motion dynamics alone. For this purpose, we propose to use silhouette sequences to represent the human actions by discarding the appearance information, and then model the sequences with linear dynamical systems (LDSs). Recognition is achieved by directly comparing the distance between LDSs, rather than resorting to complex Bayesian learning and inference. In particular, we introduce an efficient optimization method to learn robust LDSs, and develop a shift invariant distance metric to measure the similarity on the LDSs space. We evaluate our approach on the human action data set and achieve comparable results. Guan Luo, Weiming Hu 0004 |
ICIP | 2 |
| 2013 | Mining activities using sticky multimodal dual hierarchical Dirichlet process hidden Markov modelabstractIn this paper, a new nonparametric Bayesian model called Sticky Multimodal Dual Hierarchical Dirichlet Process Hidden Markov Model (SMD-HDP-HMM) is proposed for mining activities from a collection of time series. An activity is modeled as an HMM where each state corresponds to an atomic activity. By extensively using Dirichlet Process (DP), multiple HMMs sharing a common set of states are learned and the numbers of HMMs and states are both automatically determined. Each time series is modeled to be generated by one of the HMMs such that all time series are clustered into activities. Simultaneously state sequences for time series are learned and each of them is decomposed into a sequence of atomic activities. Experimental results on KTH activity dataset demonstrate the advantage of our method. Guodong Tian, Chunfeng Yuan, Weiming Hu 0004, Zhaoquan Cai 0001 |
ICIP | 3 |
| 2013 | An Improved Hierarchical Dirichlet Process-Hidden Markov Model and Its Application to Trajectory Modeling and Retrieval
Weiming Hu 0004, Guodong Tian, Xi Li 0001, Stephen J. Maybank |
Int. J. Comput. Vis. | 1 |
| 2013 | An Incremental DPMM-Based Method for Trajectory Clustering, Modeling, and RetrievalabstractTrajectory analysis is the basis for many applications, such as indexing of motion events in videos, activity recognition, and surveillance. In this paper, the Dirichlet process mixture model (DPMM) is applied to trajectory clustering, modeling, and retrieval. We propose an incremental version of a DPMM-based clustering algorithm and apply it to cluster trajectories. An appropriate number of trajectory clusters is determined automatically. When trajectories belonging to new clusters arrive, the new clusters can be identified online and added to the model without any retraining using the previous data. A time-sensitive Dirichlet process mixture model (tDPMM) is applied to each trajectory cluster for learning the trajectory pattern which represents the time-series characteristics of the trajectories in the cluster. Then, a parameterized index is constructed for each cluster. A novel likelihood estimation algorithm for the tDPMM is proposed, and a trajectory-based video retrieval model is developed. The tDPMM-based probabilistic matching method and the DPMM-based model growing method are combined to make the retrieval model scalable and adaptable. Experimental comparisons with state-of-the-art algorithms demonstrate the effectiveness of our algorithm. Weiming Hu 0004, Xi Li 0001, Guodong Tian, Stephen J. Maybank, Zhongfei Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Action recognition using linear dynamic systems
Haoran Wang 0001, Chunfeng Yuan, Guan Luo, Weiming Hu 0004, Changyin Sun 0001 |
Pattern Recognit. | 4 |
| 2013 | Block covariance based l1 tracker with a subtle template dictionary
Xiaoqin Zhang 0002, Wei Li 0034, Weiming Hu 0004, Haibin Ling, Stephen J. Maybank |
Pattern Recognit. | 3 |
| 2013 | Robust Head Tracking Based on Multiple Cues Fusion in the Kernel-Bayesian FrameworkabstractThis paper presents a robust head tracking algorithm based on multiple cues fusion in a kernel-Bayesian framework. In this algorithm, the object to be tracked is characterized using a spatial-constraint mixture of the Gaussians-based appearance model and a multichannel chamfer matching-based shape model. These two models complement each other and their combination is discriminative in distinguishing the object from the background. A selective updating technique for the appearance model is employed to accommodate appearance and illumination changes. Meantime, the kernel method-mean shift algorithm is embedded into the Bayesian framework to give a heuristic prediction in the hypotheses generation process. This alleviates the great computational load suffered by conventional Bayesian trackers. Experimental results demonstrate that the proposed algorithm is effective. Xiaoqin Zhang 0002, Weiming Hu 0004, Hujun Bao, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Active Contour-Based Visual Tracking by Integrating Colors, Shapes, and MotionsabstractIn this paper, we present a framework for active contour-based visual tracking using level sets. The main components of our framework include contour-based tracking initialization, color-based contour evolution, adaptive shape-based contour evolution for non-periodic motions, dynamic shape-based contour evolution for periodic motions, and the handling of abrupt motions. For the initialization of contour-based tracking, we develop an optical flow-based algorithm for automatically initializing contours at the first frame. For the color-based contour evolution, Markov random field theory is used to measure correlations between values of neighboring pixels for posterior probability estimation. For adaptive shape-based contour evolution, the global shape information and the local color information are combined to hierarchically evolve the contour, and a flexible shape updating model is constructed. For the dynamic shape-based contour evolution, a shape mode transition matrix is learnt to characterize the temporal correlations of object shapes. For the handling of abrupt motions, particle swarm optimization is adopted to capture the global motion which is applied to the contour in the current frame to produce an initial contour in the next frame. Weiming Hu 0004, Wei Li 0034, Wenhan Luo, Xiaoqin Zhang 0002, Stephen J. Maybank |
IEEE Trans. Image Process. | 1 |
| 2013 | A survey of appearance models in visual object trackingabstractVisual object tracking is a significant computer vision task which can be applied to many domains, such as visual surveillance, human computer interaction, and video compression. Despite extensive research on this topic, it still suffers from difficulties in handling complex object appearance changes caused by factors such as illumination variation, partial occlusion, shape deformation, and camera motion. Therefore, effective modeling of the 2D appearance of tracked objects is a key issue for the success of a visual tracker. In the literature, researchers have proposed a variety of 2D appearance models. To help readers swiftly learn the recent advances in 2D appearance models for visual object tracking, we contribute this survey, which provides a detailed review of the existing 2D appearance models. In particular, this survey takes a module-based architecture that enables readers to easily grasp the key points of visual object tracking. In this survey, we first decompose the problem of appearance modeling into two different processing stages: visual representation and statistical modeling. Then, different 2D appearance models are categorized and discussed with respect to their composition modules. Finally, we address several issues of interest as well as the remaining challenges for future research on this topic. The contributions of this survey are fourfold. First, we review the literature of visual representations according to their feature-construction mechanisms (i.e., local and global). Second, the existing statistical modeling schemes for tracking-by-detection are reviewed according to their model-construction mechanisms: generative, discriminative, and hybrid generative-discriminative. Third, each type of visual representations or statistical modeling techniques is analyzed and discussed from a theoretical or practical viewpoint. Fourth, the existing benchmark resources (e.g., source codes and video datasets) are examined in this survey. Xi Li 0001, Weiming Hu 0004, Chunhua Shen, Zhongfei Zhang, Anthony R. Dick, Anton van den Hengel |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2013 | Measuring the Visual Complexities of Web PagesabstractVisual complexities (VisComs) of Web pages significantly affect user experience, and automatic evaluation can facilitate a large number of Web-based applications. The construction of a model for measuring the VisComs of Web pages requires the extraction of typical features and learning based on labeled Web pages. However, as far as the authors are aware, little headway has been made on measuring VisCom in Web mining and machine learning. The present article provides a new approach combining Web mining techniques and machine learning algorithms for measuring the VisComs of Web pages. The structure of a Web page is first analyzed, and the layout is then extracted. Using a Web page as a semistructured image, three classes of features are extracted to construct a feature vector. The feature vector is fed into a learned measuring function to calculate the VisCom of the page. In the proposed approach of the present study, the type of the measuring function and its learning depend on the quantification strategy for VisCom. Aside from using a category and a score to represent VisCom as existing work, this study presents a new strategy utilizing a distribution to quantify the VisCom of a Web page. Empirical evaluation suggests the effectiveness of the proposed approach in terms of both features and learning algorithms. Ou Wu 0001, Weiming Hu 0004 |
ACM Trans. Web | 2 |
| 2012 | Visual Saliency Map from Tensor AnalysisabstractModeling visual saliency map of an image provides important information for image semantic understanding in many applications. Most existing computational visual saliency models follow a bottom-up framework that generates independent saliency map in each selected visual feature space and combines them in a proper way. Two big challenges to be addressed explicitly in these methods are (1) which features should be extracted for all pixels of the input image and (2) how to dynamically determine importance of the saliency map generated in each feature space. In order to address these problems, we present a novel saliency map computational model based on tensor decomposition and reconstruction. Tensor representation and analysis not only explicitly represent image's color values but also imply two important relationships inherent to color image. One is reflecting spatial correlations between pixels and the other one is representing interplay between color channels. Therefore, saliency map generator based on the proposed model can adaptively find the most suitable features and their combinational coefficients for each pixel. Experiments on a synthetic image set and a real image set show that our method is superior or comparable to other prevailing saliency map models. Bing Li 0001, Weihua Xiong, Weiming Hu 0004 |
AAAI | 3 |
| 2012 | Horror Video Scene Recognition Based on Multi-view Multi-instance Learning
Xinmiao Ding, Bing Li 0001, Weiming Hu 0004, Weihua Xiong, Zhenchong Wang |
ACCV (3) | 3 |
| 2012 | Multiple sample group pairs' graph embedding for trackingabstractThis paper presents a new method which uses graph embedding and foreground-background patch pairs to perform object tracking. We first use particle filter to sample some particles. Then we evaluate each particle based on graph embedding and foreground-background patch pairs. For each particle, we use a two-layer model to represent the object, i.e. the inner layer (object layer) and the outer layer (background layer). Both the two layers are divided into patches. We cluster the foreground patches to several classes. Each class forms one sample group pair with the background patches. We perform graph embedding on multiple sample group pairs to discriminate the foreground and the background. Experimental results showed that our method tracked the objects efficiently. Weiming Hu 0004, Xiaoqin Zhang 0002 |
ICIP | 2 |
| 2012 | Context-aware horror video scene recognition via cost-sensitive sparse coding
Xinmiao Ding, Bing Li 0001, Weiming Hu 0004, Weihua Xiong, Zhenchong Wang |
ICPR | 3 |
| 2012 | Scaring or pleasing: exploit emotional impact of an imageabstractAutomatic image emotion analysis has emerged as a hot topic due to its potential application on high-level image understanding. Considering the fact that the emotion evoked by an image is not only from its global appearance but also interplays among local regions, we propose a novel affective image classification system based on bilayer sparse representation (BSR). The BSR model contains two layers: The global sparse representation (GSR) is to define global similarities between a test image and all the training images; and the local sparse representation (LSR) is to define similarities of local regions' appearances and their co-occurrence between a test image and all the training images. The experiments on real data sets demonstrate that our system is effective on image emotion recognition. Bing Li 0001, Songhe Feng, Weihua Xiong, Weiming Hu 0004 |
ACM Multimedia | 4 |
| 2012 | Context-aware affective images classification based on bilayer sparse representationabstractIn image understanding, the automatic recognition of emotion in an image is becoming important from an applicative viewpoint. Considering the fact that the emotion evoked by an image is not only from its global appearance but also interplays among local regions, we propose a novel context-aware classification model based on bilayer sparse representation (BSR) that simultaneously takes the local context and global-local context into account. The BSR model contains two layers: global sparse representation (GSR) and local sparse representation (LSR). The GSR is to define global similarities between a test image and all training images; while the LSR is to define similarities of local regions' appearances and their co-occurrence between a test image and all training images. The experiments on two data sets demonstrate that our method is effective on affective images classification. Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Xinmiao Ding |
ACM Multimedia | 3 |
| 2012 | Unsupervised Ensemble Learning for Mining Top-n Outliers
Weiming Hu 0004, Zhongfei Zhang, Ou Wu 0001 |
PAKDD (1) | 2 |
| 2012 | Single and Multiple Object Tracking Using Log-Euclidean Riemannian Subspace and Block-Division Appearance ModelabstractObject appearance modeling is crucial for tracking objects, especially in videos captured by nonstationary cameras and for reasoning about occlusions between multiple moving objects. Based on the log-euclidean Riemannian metric on symmetric positive definite matrices, we propose an incremental log-euclidean Riemannian subspace learning algorithm in which covariance matrices of image features are mapped into a vector space with the log-euclidean Riemannian metric. Based on the subspace learning algorithm, we develop a log-euclidean block-division appearance model which captures both the global and local spatial layout information about object appearances. Single object tracking and multi-object tracking with occlusion reasoning are then achieved by particle filtering-based Bayesian state inference. During tracking, incremental updating of the log-euclidean block-division appearance model captures changes in object appearance. For multi-object tracking, the appearance models of the objects can be updated even in the presence of occlusions. Experimental results demonstrate that the proposed tracking algorithm obtains more accurate results than six state-of-the-art tracking algorithms. Weiming Hu 0004, Xi Li 0001, Wenhan Luo, Xiaoqin Zhang 0002, Stephen J. Maybank, Zhongfei Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Supervised class-specific dictionary learning for sparse modeling in action recognition
Haoran Wang 0001, Chunfeng Yuan, Weiming Hu 0004, Changyin Sun 0001 |
Pattern Recognit. | 3 |
| 2012 | Efficient Clustering Aggregation Based on Data FragmentsabstractClustering aggregation, known as clustering ensembles, has emerged as a powerful technique for combining different clustering results to obtain a single better clustering. Existing clustering aggregation algorithms are applied directly to data points, in what is referred to as the point-based approach. The algorithms are inefficient if the number of data points is large. We define an efficient approach for clustering aggregation based on data fragments. In this fragment-based approach, a data fragment is any subset of the data that is not split by any of the clustering results. To establish the theoretical bases of the proposed approach, we prove that clustering aggregation can be performed directly on data fragments under two widely used goodness measures for clustering aggregation taken from the literature. Three new clustering aggregation algorithms are described. The experimental results obtained using several public data sets show that the new algorithms have lower computational complexity than three well-known existing point-based clustering aggregation algorithms (Agglomerative, Furthest, and LocalSearch); nevertheless, the new algorithms do not sacrifice the accuracy. Ou Wu 0001, Weiming Hu 0004, Stephen J. Maybank, Mingliang Zhu, Bing Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2011 | Evaluating combinational color constancy methods on real-world imagesabstractLight color estimation is crucial to the color constancy problem. Past decades have witnessed great progress in solving this problem. Contrary to traditional methods, many researchers propose a variety of combinational color constancy methods through applying different color constancy mathematical models on an image simultaneously and then give out a final estimation in diverse ways. Although many comprehensive evaluations or reviews about color constancy methods are available, few focus on combinational strategies. In this paper, we survey some prevailing combinational strategies systematically; divide them into three categories and compare them qualitatively on three real-world image data sets in terms of the angular error and the perceptual Euclidean distance. The experimental results show that combinational strategies with training procedure always produces better performance. Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Ou Wu 0001 |
CVPR | 3 |
| 2011 | Efficient block-division model for robust multiple object trackingabstractTracking multiple objects under occlusion is one of the most challenging issues in computer vision. Occlusion results in mistaken match when finding the most similar candidate. Adapting to the change of objects is essential for tracking as objects often undergo intrinsic changes, but noise is unavoidably introduced during updating of the object, and this further confuses the tracker. In order to address these problems, a block-division appearance model is introduced to efficiently handle occlusion. In this model, spatial information is introduced to avoid the mistaken match between object and candidate. Based on this model, a selective updating strategy is proposed to incrementally learn the change of the object, avoiding introducing noise when updating. At the same time occlusion is deduced by monitoring the variation of each block. Experimental results in various videos validate the effectiveness of our algorithm in tracking multiple objects under occlusion. Wenhan Luo, Xiaoqin Zhang 0002, Yang Liu 0020, Xi Li 0001, Weiming Hu 0004, Wei Li 0034 |
ICASSP | 5 |
| 2011 | Multi-cue based multi-target tracking using online random forestsabstractDiscriminative tracking has become popular tracking methods due to their descriptive power for foreground/background separation. Among these methods, online random forest is recently proposed and received a large amount of research attention due to its advantages such as efficiency and robust ness to noise, etc. However, the fact that only one kind of features is used limits the discriminative performance of this tracker. Additionally, the standard online forest tracker works only for a single target object. In this paper, we introduce a novel tracking method that integrates multiple cues capturing both geometric structures and edge-based shape information. Compared with the current online random forest based tracking algorithm, the proposed multi-cue tracker is more robust thanks to the complimentary information provided from these hybrid cues. Furthermore, the new tracker can track multiple targets as well as single target object. The effectiveness of the proposed tracker is validated using five public sequences. Xinchu Shi, Xiaoqin Zhang 0002, Yang Liu 0020, Weiming Hu 0004, Haibin Ling |
ICASSP | 4 |
| 2011 | Horror video scene recognition via Multiple-Instance learningabstractAlong with the ever-growing Web comes the proliferation of objectionable content, such as pornography, violence, horror information, etc. Horror videos, whose threat to childrens health is no less than pornographic video, are sometimes neglected by existing Web filtering tools. Consequently, an effective horror video filtering tool is necessary for preventing children from accessing these harmful horror videos. In this paper, by introducing color emotion and color harmony theories, we propose a horror video scenes recognition algorithm. Firstly, the video scenes are decomposed into a set of shots. Then we extract the visual features, audio features and emotional features of each shot, the video scene is viewed as a bag and each shot is treated as an instance of the corresponding bag. Finally, by combining the three features, the horror video scenes are recognized by the Multiple-Instance learning(MIL). According to the experimental results on diverse video scenes, the proposed scheme based on the emotional perception could effectively deal with the horror video scene recognition and promising results are achieved. Bing Li 0001, Weiming Hu 0004, Ou Wu 0001 |
ICASSP | 3 |
| 2011 | Learning to predict the perceived visual quality of photosabstractVisual quality (VisQ) representation is a fundamental step in the learning of a VisQ prediction model for photos. It not only reflects how we understand VisQ but also determines the label type. Existing studies apply a scalar value (i.e., a categorical label or a score) to represent VisQ. As VisQ is a subjective property, only a scalar value is insufficient to represent human's perceived VisQ of a photo. This study represents VisQ by a distribution on pre-defined ordinal basic ratings in order to capture the subjectivity of VisQ better. When using the new representation, the label type is structural instead of scalar. Conventional learning algorithms cannot be directly applied in model learning. Meanwhile, for many photos, the numbers of users involved in the evaluation are limited, making some labels unreliable. In this study, a new algorithm called support vector distribution regression (SVDR) is presented to deal with the structural output learning. Two independent learning strategies (reliability-sensitive learning and label refinement) are proposed to alleviate the difficulty of insufficient involved users for rating. Combining SVDR with the two learning strategies, two separate structural-output regression algorithms (i.e., reliability-sensitive SVDR and label refinement-based SVDR) are produced. Experimental results demonstrate the effectiveness of our introduced learning strategies and learning algorithms. Ou Wu 0001, Weiming Hu 0004 |
ICCV | 2 |
| 2011 | Context-Aware Multi-instance Learning Based on Hierarchical Sparse RepresentationabstractMulti-instance learning (MIL), a variant of supervised learning framework, has been applied in many applications. More recently, researchers focus on two important issues for MIL: Instances' contextual structures representation in the same bag and online MIL schemes. In this paper, we present an effective context-aware multi-instance learning technique using a hierarchical sparse representation (HSR-MIL) that addresses the two challenges simultaneously. We firstly construct the inner contextual structure among instances in the same bag based on a novel sparse ε-graph. We then propose a graph kernel based sparse bag classifier through a modified kernel sparse coding in higher-dimension feature space. At last, the HSR-MIL approach is extended to achieve online learning manner with an incremental kernel matrix update scheme. The experiments on several data sets demonstrate that our method has better performances and online learning ability. Bing Li 0001, Weihua Xiong, Weiming Hu 0004 |
ICDM | 3 |
| 2011 | Web Horror Image Recognition Based on Context-Aware Multi-instance LearningabstractAlong with the ever-growing Web, horror contents sharing in the Internet has interfered with our daily life and affected our, especially children's, health. Therefore horror image recognition is becoming more important for web objectionable content filtering. This paper presents a novel context-aware multi-instance learning (CMIL) model for this task. This work is distinguished by three key contributions. Firstly, the traditional multi-instance learning is extended to context-aware multi-instance learning model through integrating an undirected graph in each bag that represents contextual relationships among instances. Secondly, by introducing a novel energy function, a heuristic optimization algorithm based on Fuzzy Support Vector Machine (FSVM) is given out to find the optimal classifier on CMIL. Finally, the CMIL is applied to recognize horror images. Experimental results on an image set collected from the Internet show that the proposed method is effective on horror image recognition. Bing Li 0001, Weihua Xiong, Weiming Hu 0004 |
ICDM | 3 |
| 2011 | Robust visual tracking via transfer learningabstractIn this paper, we propose a boosting based tracking framework using transfer learning. To deal with complex appearance variations, the proposed tracking framework tries to utilize discriminative information from previous frames to conduct the tracking task in the current frame, and thus transfers some prior knowledge from the previous source data domain to the current target data domain, resulting in a high discriminative tracker for distinguishing the object from the background. The proposed tracking system has been tested on several challenging sequences. Experimental results demonstrate the effectiveness of the proposed tracking framework. Wenhan Luo, Xi Li 0001, Wei Li 0034, Weiming Hu 0004 |
ICIP | 4 |
| 2011 | Learning to Rank under Multiple AnnotatorsabstractLearning to rank has received great attention in recent years as it plays a crucial role in information retrieval. The existing concept of learning to rank assumes that each training sample is associated with an instance and a reliable label. However, in practice, this assumption does not necessarily hold true. This study focuses on the learning to rank when each training instance is labeled by multiple annotators that may be unreliable. In such a scenario, no accurate labels can be obtained. This study proposes two learning approaches. One is to simply estimate the ground truth first and then to learn a ranking model with it. The second approach is a maximum likelihood learning approach which estimates the ground truth and learns the ranking model iteratively. The two approaches have been tested on both synthetic and real-world data. The results reveal that the maximum likelihood approach outperforms the first approach significantly and is comparable of achieving results with the learning model considering reliable labels. Further more, both the approaches have been applied for ranking the Web visual clutter. Ou Wu 0001, Weiming Hu 0004 |
IJCAI | 2 |
| 2011 | RKOF: Robust Kernel-Based Local Outlier Detection
Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Ou Wu 0001 |
PAKDD (2) | 2 |
| 2011 | Evaluating the visual quality of web pages using a computational aesthetic approachabstractCurrent Web mining explores useful and valuable information (content) online for users. However, there is scant research on the overall visual aspect of Web pages, even though visual elements such as aesthetics significantly influence user experience. A beautiful and well-laid out Web page greatly facilitates users' accessing and enhances browsing experiences.We use "visual quality (VisQ)" to denote the aesthetics of Web pages. In this paper, a computational aesthetics approach is proposed to learn the evaluation model for the visual quality of Web pages. First, a Web page layout extraction algorithm (V-LBE) is introduced to partition a Web page into major layout blocks. Then, regarding a Web page as a semi-structured image, features (e.g., layout,visual complexity, colorfulness) known to significantly affect the visual quality of a Web page are extracted to construct a feature vector. We present a multi-cost-sensitive learning for visual quality classification and a multi-value regression for visual quality score assignment. Our experiments compare the extracted features and conclude that the Web page's layout visual features (LV) and text visual features (TV) are the primary affecting factors toward Web page's visual quality. The performance of the learned visual quality classifier is close to some persons'. The learned regression function also achieves promising results. Ou Wu 0001, Yunfei Chen 0002, Bing Li 0001, Weiming Hu 0004 |
WSDM | 4 |
| 2011 | Incremental Tensor Subspace Learning and Its Applications to Foreground Segmentation and TrackingabstractAppearance modeling is very important for background modeling and object tracking. Subspace learning-based algorithms have been used to model the appearances of objects or scenes. Current vector subspace-based algorithms cannot effectively represent spatial correlations between pixel values. Current tensor subspace-based algorithms construct an offline representation of image ensembles, and current online tensor subspace learning algorithms cannot be applied to background modeling and object tracking. In this paper, we propose an online tensor subspace learning algorithm which models appearance changes by incrementally learning a tensor subspace representation through adaptively updating the sample mean and an eigenbasis for each unfolding matrix of the tensor. The proposed incremental tensor subspace learning algorithm is applied to foreground segmentation and object tracking for grayscale and color image sequences. The new background models capture the intrinsic spatiotemporal characteristics of scenes. The new tracking algorithm captures the appearance characteristics of an object during tracking and uses a particle filter to estimate the optimal object state. Experimental evaluations against state-of-the-art algorithms demonstrate the promise and effectiveness of the proposed incremental tensor subspace learning algorithm, and its applications to foreground segmentation and object tracking. Weiming Hu 0004, Xi Li 0001, Xiaoqin Zhang 0002, Xinchu Shi, Stephen J. Maybank, Zhongfei Zhang |
Int. J. Comput. Vis. | 1 |
| 2011 | Visual tracking via dynamic tensor analysis with mean update
Xiaoqin Zhang 0002, Xinchu Shi, Weiming Hu 0004, Xi Li 0001, Stephen J. Maybank |
Neurocomputing | 3 |
| 2011 | Adaptive learning codebook for action recognition
Yu Kong 0001, Xiaoqin Zhang 0002, Weiming Hu 0004, Yunde Jia |
Pattern Recognit. Lett. | 3 |
| 2011 | Recognition of adult images, videos, and web page bagsabstractIn this article, we develop an integrated adult-content recognition system which can detect adult images, adult videos, and adult Web page bags, where a Web page bag consists of a Web page and a predefined number of Web pages linked to it through hyperlinks. In our adult image-recognition algorithm, we model skin patches rather than skin pixels, resulting in better results than state-of-the-art algorithms which model skin pixels. In our adult video-recognition algorithm, information from the accompanying audio section around an image in an adult video is used to obtain a prior classification of the image. The algorithm achieves a better performance than the ones which use image information alone or audio information alone. The adult Web page bag recognition is carried out using multi-instance learning based on the combination of classifying texts, images and videos in Web pages. Both the speed and the accuracy for recognizing the Web adult content are increased, in contrast to recognizing Web pages one-by-one. Weiming Hu 0004, Haiqiang Zuo, Ou Wu 0001, Yunfei Chen 0002, Zhongfei Zhang, David Suter |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2011 | A Survey on Visual Content-Based Video Indexing and RetrievalabstractVideo indexing and retrieval have a wide spectrum of promising applications, motivating the interest of researchers worldwide. This paper offers a tutorial and an overview of the landscape of general strategies in visual content-based video indexing and retrieval, focusing on methods for video structure analysis, including shot boundary detection, key frame extraction and scene segmentation, extraction of features including static key frame features, object features and motion features, video data mining, video annotation, video retrieval including query interfaces, similarity measure and relevance feedback, and video browsing. Finally, we analyze future research directions. Weiming Hu 0004, Nianhua Xie, Li Li 0010, Xianglin Zeng, Stephen J. Maybank |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2010 | Horror Image Recognition Based on Emotional Attention
Bing Li 0001, Weiming Hu 0004, Weihua Xiong, Ou Wu 0001, Wei Li 0034 |
ACCV (2) | 2 |
| 2010 | Occlusion Handling with ℓ1-Regularized Sparse Reconstruction
Wei Li 0034, Bing Li 0001, Xiaoqin Zhang 0002, Weiming Hu 0004, Hanzi Wang, Guan Luo |
ACCV (4) | 4 |
| 2010 | Top-Down Cues for Event Recognition
Li Li 0010, Chunfeng Yuan, Weiming Hu 0004, Bing Li 0001 |
ACCV (3) | 3 |
| 2010 | Group ranking with application to image retrievalabstractMany existing ranking-related information processing applications can be summarized into one theoretical problem called group ranking (GR). A simple average-ranking approach is usually applied to GR. Although the approach seems reasonable, no theoretical analysis about its intrinsic mechanism has been presented, increasing the difficulty of evaluating the ranking results. This study provides a formal analysis for GR. We first construct an objective function for the GR problem, and discover that each GR problem can be transformed into a rank aggregation problem whose objective function is proved to be equal to the objective function of GR. As a consequence, the average-ranking approach can be explained by two well-known rank aggregation techniques. We incorporate two other effective rank aggregation methods into the GR problem and obtain two new GR algorithms. We apply the GR algorithms into image retrieval to diversify the image search results returned by search engines. Experimental results show the effectiveness of the proposed GR algorithms. Ou Wu 0001, Weiming Hu 0004, Bing Li 0001 |
CIKM | 2 |
| 2010 | Use bin-ratio information for category and scene classificationabstractIn this paper we propose using bin-ratio information, which is collected from the ratios between bin values of histograms, for scene and category classification. To use such information, a new histogram dissimilarity, bin-ratio dissimilarity (BRD), is designed. We show that BRD provides several attractive advantages for category and scene classification tasks: First, BRD is robust to cluttering, partial occlusion and histogram normalization; Second, BRD captures rich co-occurrence information while enjoying a linear computational complexity; Third, BRD can be easily combined with other dissimilarity measures, such as L1and χ2, to gather complimentary information. We apply the proposed methods to category and scene classification tasks in the bag-of-words framework. The experiments are conducted on several widely tested datasets including PASCAL 2005, PASCAL 2008, Oxford flowers, and Scene-15 dataset. In all experiments, the proposed methods demonstrate excellent performance in comparison with previously reported solutions. Nianhua Xie, Haibin Ling, Weiming Hu 0004, Xiaoqin Zhang 0002 |
CVPR | 3 |
| 2010 | Spatio-Temporal Proximity Distribution Kernels for Action Recognition
Chunfeng Yuan, Weiming Hu 0004, Hanzi Wang, Xi Li 0001, Nianhua Xie |
ICASSP | 2 |
| 2010 | Horror movie scene recognition based on emotional perceptionabstractThe number of video clips available online is growing at a tremendous pace. Meanwhile, the video scenes of pornography, violence and horror permeate the whole Web. Horror videos, whose threat to children's health is no less than pornographic video, are sometimes neglected by existing Web filtering tools. Consequently, an effective horror video filtering tool is necessary for preventing children from accessing these horror videos. In this paper, by introducing color emotion and color harmony theories, we propose a horror video scene recognition algorithm. Firstly, the video scenes are decomposed into a set of shots. Then we extract the visual features, audio features and color emotion features of each shot. Finally, by combining the three features, the horror video scenes are recognized by the Support Vector Machine (SVM) classifier. According to the experimental results on diverse video scenes, the proposed scheme based on the emotional perception could deal effectively with the horror video scene recognition and promising results are achieved. Bing Li 0001, Weiming Hu 0004, Ou Wu 0001 |
ICIP | 3 |
| 2010 | Compact visual codebook for action recognitionabstractVisual codebook has been popular in object classification as well as action analysis. However, its performance is often sensitive to the codebook size that is usually predefined. Moreover, the codebook generated by unsupervised methods, e.g., K-means, often suffers from the problem of ambiguity and weak efficiency. In other words, the visual codebook contains a lot of noisy and/or ambiguous words. In this paper, we propose a novel method to address these issues by constructing a compact but effective visual codebook using sparse reconstruction. Given a large codebook generated by K-means, we reformulate it in a sparse manner, and learn the weight of each word in the original visual codebook. Since the weights are sparse, they naturally introduce a new compact codebook. We apply this compact codebook to action recognition tasks and verify it on the widely used Weizmann action database. The experimental results show clearly the benefits of the proposed solution. Qingdi Wei, Xiaoqin Zhang 0002, Yu Kong 0001, Weiming Hu 0004, Haibin Ling |
ICIP | 4 |
| 2010 | Local Outlier Detection Based on Kernel RegressionabstractOutlier detection keeps an important and attractive task of the knowledge discovery in databases. In this paper, a novel approach named Multi-scale Local Kernel Regression is proposed. It transfers the unsupervised learning of outlier detection to the classic non-parameter regression learning. Through preprocessing the original data by the basic local density-based method, it adopts the local kernel regression estimator in the multiple scale neighborhoods to determine outliers. Experiments on several real life data sets demonstrate that this approach is promising in detection performance. Weiming Hu 0004, Wei Li 0034, Zhongfei Zhang, Ou Wu 0001 |
ICPR | 2 |
| 2010 | Event Recognition Based on Top-Down Motion AttentionabstractHow to fuse static and dynamic information is a key issue in event analysis. In this paper, a top-down motion guided fusing method is proposed for recognizing events in an unconstrained news video. In the method, the static information is represented as a Bag-of-SIFT-features and motion information is employed to generate event specific attention map to direct the sampling of the interest points. We build class-specific motion histograms for each event so as to give more weight on the interest points that are discriminative to the corresponding event. Experimental results on TRECVID 2005 video corpus demonstrate that the proposed method can improve the mean average accuracy of recognition. Li Li 0010, Weiming Hu 0004, Bing Li 0001, Chunfeng Yuan, Pengfei Zhu 0001, Wanqing Li 0001 |
ICPR | 2 |
| 2010 | Discriminative Level Set for Contour TrackingabstractConventional contour tracking algorithms with level set often use generative models to construct the energy function. For tracking through cluttered and noisy background, however, a generative model may not be discriminative enough. In this paper we integrate the discriminative methods into a level set framework when constructing the level set energy function. We train a set of weak classifiers to distinguish the object from the background. Each weak classifier is designed to select the most discriminative feature space and integrated via AdaBoost according to their training errors. We also introduce a novel interaction term to explore the correlation between pixels near the object edge. This term together with the discriminative model both enhance the discriminative power of the level set. The experimental results show that the contour tracked by our approach is more accurate than the conventional algorithms with the generative model. Our algorithm successfully tracks the object contour even in a cluttered environment. Wei Li 0034, Xiaoqin Zhang 0002, Weiming Hu 0004, Haibin Ling |
ICPR | 4 |
| 2010 | Semi-supervised Trajectory Learning Using a Multi-Scale Key Point Based Trajectory RepresentationabstractMotion trajectories contain rich high-level semantic information such as object behaviors and gestures, which can be effectively captured by supervised trajectory learning. However, it is usually a tough task to obtain a large number of high-quality manually labeled samples in real applications. Thus, how to perform trajectory learning in small training sample size situations is an important research topic. In this paper, we propose a trajectory learning framework using graph-based semi-supervised transductive learning, which propagates training sample labels along a particular graph. Furthermore, a novel trajectory descriptor based on multi-scale key points is proposed to characterize the spatial structural information. Experimental results demonstrate effectiveness of our framework. Yang Liu 0020, Xi Li 0001, Weiming Hu 0004 |
ICPR | 3 |
| 2010 | Prototype Learning Using Metric Learning Based Behavior RecognitionabstractBehavior recognition is an attractive direction in the computer vision domain. In this paper, we propose a novel behavior recognition method based on prototype learning using metric learning. Prototype learning algorithm can improve the classification performance of nearest-neighbor classifier, reduce the storage and computation requirements. And the metric learning algorithm is used to advance the performance of the prototype learning. In this paper, We use a kind of compound feature including local feature and motion feature to recognize human behaviors. The experimental results show the effectiveness of our method. Pengfei Zhu 0001, Weiming Hu 0004, Chunfeng Yuan, Li Li 0010 |
ICPR | 2 |
| 2010 | Video Scene Segmentation Using Time Constraint Dominant-Set Clustering
Xianglin Zeng, Xiaoqin Zhang 0002, Weiming Hu 0004, Wanqing Li 0001 |
MMM | 3 |
| 2010 | Identifying Multi-instance OutliersabstractThis paper studies a new data mining problem called multi-instance outlier identification. This problem arises in tasks where each sample consists of many alternative feature vectors (instances) that describe it. This paper defines the multi-instance outliers and analyzes the basic types of multi-instance outliers. Two general identification approaches are proposed based on the state-of-the-art (single-instance) outlier detector LOF (local outlier factor). One approach utilizes the underlying mechanism of the kernel method and plunges the set distance into LOF to detect the multi-instance outliers. The other approach takes each instance's neighborhood into account. Based on the two approaches, four concrete multi-instance outlier detectors are then introduced. We conduct experiments over four synthetic data collections and three real-world data collections (two Musk data sets [22, 23] and a hard-drive inspection data set [24]). The experimental results show that the proposed multi-instance outlier detectors are effective while the algorithms that ignore the multi-instance settings perform poorly. Especially, the results on the two Musk sets are consistent with the multi-instance learning results; the results on the hard-drive inspection data set demonstrate that multi-instance outlier identification is promising for real applications. Ou Wu 0001, Weiming Hu 0004, Bing Li 0001, Mingliang Zhu |
SDM | 3 |
| 2010 | Topic Detection for Discussion Threads with Domain KnowledgeabstractThe online communities are becoming so popular along with the development of the web but indexing and searching for the discussion data are big challenges to current applications. Topic detection was proposed to solve the problem but the accuracy is still not satisfactory, mainly because key elements are usually implicit or ambiguous which literal content comparison cannot handle. In this paper, we propose to improve the basic topic detection model by combining domain knowledge. The domain knowledge can be automatically extracted from a collection of external knowledge sources and applied to the content analysis of the threads. Two approaches, i.e. the LDA and the Concept Mapping, are proposed to implement the knowledge extraction and integration. Experimental results show that both approaches make the detection accuracy outperform the previous model. The LDA approach achieves better overall performance while the Concept Mapping is more suitable for dynamic knowledge sources. Mingliang Zhu, Weiming Hu 0004, Ou Wu 0001 |
Web Intelligence | 2 |
| 2010 | Learning to evaluate the visual quality of web pagesabstractA beautiful and well-laid out Web page greatly facilitates users' accessing and enhances browsing experiences. We use "visual quality (VQ)" to denote the aesthetics of Web pages. In this paper, a computational aesthetics approach is proposed to learn the evaluation model for the visual quality of Web pages. First, a Web page layout extraction algorithm (V-LBE) is introduced to partition a Web page into major layout blocks. Then, regarding a Web page as a semi-structured image, features known to significantly affect the visual quality of a Web page are extracted to construct a feature vector. The experimental results show the initial success of our approach. Potential applications include Web search and Web design. Ou Wu 0001, Yunfei Chen 0002, Bing Li 0001, Weiming Hu 0004 |
WWW | 4 |
| 2010 | Patch-based skin color detection and its application to pornography image filteringabstractAlong with the explosive growth of the World Wide Web, an immense industry for the production and consumption of pornography has grown. Though the censorship and legal restraints on pornography are discriminating in different historical, cultural and national contexts, selling pornography to minors is not allowed in most cases. Detecting human skin tone is of utmost importance in pornography image filtering algorithms. In this paper, we propose two patch-based skin color detection algorithms: regular patch and irregular patch skin color detection algorithms. On the basis of skin detection, we extract 31-dimensional features from the input image, and these features are fed into a random forest classifier. Our algorithm has been incorporated into an adult-content filtering infrastructure, and is now in active use for preventing minors from accessing pornographic images via mobile phones. Haiqiang Zuo, Weiming Hu 0004, Ou Wu 0001 |
WWW | 2 |
| 2010 | Linear discriminant analysis using rotational invariant L1 norm
Xi Li 0001, Weiming Hu 0004, Hanzi Wang, Zhongfei Zhang |
Neurocomputing | 2 |
| 2010 | Robust object tracking using a spatial pyramid heat kernel structural information representation
Xi Li 0001, Weiming Hu 0004, Hanzi Wang, Zhongfei Zhang |
Neurocomputing | 2 |
| 2010 | Heat Kernel Based Local Binary Pattern for Face RepresentationabstractFace classification has recently become a very hot research topic in computer vision and multimedia information processing. It has many potential applications, in which face representation is the most fundamental task. Most existing face representation methods perform poorly in capturing the intrinsic structural information of face appearance. To address this problem, we propose a novel multiscale heat kernel based face representation, for heat kernels perform well in characterizing the topological structural information of face appearance. Further, the local binary pattern (LBP) descriptor is incorporated into the multiscale heat kernel face representation for the purpose of capturing texture information of face appearance. As a result, we have the heat kernel based local binary pattern (HKLBP) descriptor. Finally, a Support Vector Machine (SVM) classifier is learned in theHKLBPfeature space for face classification. Experimental results demonstrate the effectiveness and superiority of our face classification framework. Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Hanzi Wang |
IEEE Signal Process. Lett. | 2 |
| 2010 | Multiple Object Tracking Via Species-Based Particle Swarm OptimizationabstractMultiple object tracking is particularly challenging when many objects with similar appearances occlude one another. Most existing approaches concatenate the states of different objects, view the multi-object tracking as a joint motion estimation problem and search for the best state of the joint motion in a rather high dimensional space. However, this centralized framework suffers from a high computational load. We bring a new view to the tracking problem from a swarm intelligence perspective. In analogy with the foraging behavior of bird flocks, we propose a species-based particle swarm optimization algorithm for multiple object tracking, in which the global swarm is divided into many species according to the number of objects, and each species searches for its object and maintains track of it. The interaction between different objects is modeled as species competition and repulsion, and the occlusion relationship is implicitly deduced from the “power” of each species, which is a function of the image observations. Therefore, our approach decentralizes the joint tracker to a set of individual trackers, each of which tries to maximize its visual evidence. Experimental results demonstrate the efficiency and effectiveness of our method. Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Learning Group Activity in Soccer Videos from Local Motion
Yu Kong 0001, Weiming Hu 0004, Xiaoqin Zhang 0002, Hanzi Wang, Yunde Jia |
ACCV (1) | 2 |
| 2009 | Spectral Graph Partitioning Based on a Random Walk Diffusion Similarity Measure
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Yang Liu 0020 |
ACCV (2) | 2 |
| 2009 | Human Action Recognition under Log-Euclidean Riemannian Metric
Chunfeng Yuan, Weiming Hu 0004, Xi Li 0001, Stephen J. Maybank, Guan Luo |
ACCV (1) | 2 |
| 2009 | Human Action Recognition Using Pyramid Vocabulary Tree
Chunfeng Yuan, Xi Li 0001, Weiming Hu 0004, Hanzi Wang |
ACCV (3) | 3 |
| 2009 | A Smarter Particle Filter
Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank |
ACCV (2) | 2 |
| 2009 | Segment Model Based Vehicle Motion AnalysisabstractMotion analysis is a very attractive research direction in computer vision field. In this paper, we propose a framework for analyzing real vehicle motion in visual traffic surveillance by using Segment Model (SM), which is a kind of probabilistic model. SM can grasp the underlying information of observation sequence by using segment distribution. It has been proved to be more precise than that of HMM. In the experiments, we compare our approach with the template matching method based on the Hausdorff distance and the state space method based on the Hidden Markov Model (HMM). The experimental results show the effectiveness of our approach. Pengfei Zhu 0001, Weiming Hu 0004, Xi Li 0001, Li Li 0010 |
AVSS | 2 |
| 2009 | Fragment-based clustering ensemblesabstractClustering ensembles combine different clustering solutions into a single robust and stable one. Most of existing methods become highly time-consuming when the data size turns to large. In this paper, we study the properties of the defined 'clustering fragment' and put forward a useful proposition. Solid proofs are presented with two widely used goodness measures for clustering ensembles. Finally, a new ensemble framework termed as fragment-based clustering ensembles is proposed. Theoretically, most of existing methods can be improved by adopting this framework. To evaluate the proposed framework, three new methods are introduced by bring three popular clustering ensemble methods into our framework. The experimental results on several public data sets show that the three introduced methods are greatly improved in computational complexity and also achieved better or similar accurate results than the original methods. Ou Wu 0001, Mingliang Zhu, Weiming Hu 0004 |
CIKM | 3 |
| 2009 | Image spam filtering using Fourier-Mellin invariant featuresabstractImage spam is a new obfuscating method which spammers invented to more effectively bypass conventional text based spam filters. In this paper, a framework for filtering image spams by using the Fourier-Mellin invariant features is described. Fourier-Mellin features are robust for most kinds of image spam variations. A one-class classifier, the support vector data description (SVDD), is exploited to model the boundary of image spam class in the feature space without using information of legitimate emails. Experimental results demonstrate that our framework is effective for fighting image spam. Haiqiang Zuo, Xi Li 0001, Ou Wu 0001, Weiming Hu 0004, Guan Luo |
ICASSP | 4 |
| 2009 | Efficient human pose estimation via parsing a tree structure based human modelabstractHuman pose estimation is the task of determining the states (location, orientation and scale) of each body part. It is important for many vision understanding applications, e.g. visual interactive gaming, immersive virtual reality, content-based image retrieval, etc. However, it remains a challenging task because of unknown image background, presence of clutter, partial occlusion and especially the high dimensional state space (usually 30+ dimensions). In this paper, we contribute to human pose estimation in two aspects. First, we design two efficient Markov Chain dynamics under the data-driven Markov Chain Monte Carlo (DDMCMC) framework to effectively explore the complex solution space. Second, we parse the tree structure state space into a lexicographic order according to the image observations and body topology, and the optimization process is conducted in this order. This realizes a much more efficient exploration than the sampling based search and exhaustive search, and thus achieves a tremendous speed-up. Experimental results demonstrate the efficiency and effectiveness of the proposed method in estimating various kinds of human poses, even with cluttered background , poor illumination or partial self-occlusion. Xiaoqin Zhang 0002, Xiaofeng Tong, Weiming Hu 0004, Stephen J. Maybank, Yimin Zhang 0002 |
ICCV | 4 |
| 2009 | Contour tracking with abrupt motionabstractTraditional contour tracking methods can not handle abrupt motion or low frame rate video. This is because the basis of the traditional tracking lies in the assumption that the motion is smooth between consecutive frames. However, the abrupt motion destroys the foundation of this assumption. In this paper, we integrate the stochastic search into the level set evolution to reinstitute the continuity. Our approach can be viewed as a two-layer hierarchical level set-based tracking framework in which Particle Swarm Optimization (PSO) and level set evolution are fused seamlessly. In the first layer, the PSO is adopted to capture the global motion of the object. The coarse contour is obtained by applying the global motion to the contour in the previous frame. For the second layer, the level set evolution based on the coarse contour is carried out to track the local deformation, which results in the actual contour. The promising experimental results for numerous real videos reveal the effectiveness of our approach. Wei Li 0034, Xiaoqin Zhang 0002, Weiming Hu 0004 |
ICIP | 3 |
| 2009 | A Boosted Semi-supervised Learning Framework for Web Page FilteringabstractThe World Wide Web provides great convenience for users to obtain information. However, there exists much harmful information on the Internet, such as pornographic content and prohibited drugs' information. Thus, how to filter harmful Web pages on the Internet is quite an important issue. In general, the problem of harmful Web page filtering is converted to that of Web page classification, which needs plenty of well labeled training samples. However, the cost of labeling a large set of Web pages is very expensive. To address this problem, we adopt a semi-supervised framework for Web page filtering. In this framework, each Web page is represented by bags of different features, extracted using its HTML structure. Then a semi-supervised learning strategy is taken for efficiently obtaining well labeled training samples. Finally, a boosting classifier is utilized for harmful Web page filtering. Experiments have demonstrated the effectiveness of our framework. Zhu He, Xi Li 0001, Weiming Hu 0004 |
SMC | 3 |
| 2009 | Adaptive Distributed Intrusion Detection Using Parametric ModelabstractDue to the increasing demands for network security, distributed intrusion detection has become a hot research topic in computer science. However, the design and maintenance of the intrusion detection system (IDS) is still a challenging task due to its dynamic, scalability, and privacy properties. In this paper, we propose a distributed IDS framework which consists of the individual and global models. Specifically, the individual model for the local unit derives from Gaussian Mixture Model based on online Adaboost algorithm, while the global model is constructed through the PSO-SVM fusion algorithm. Experimental results demonstrate that our approach can achieve a good detection performance while being trained online and consuming little traffic to communicate between local units. Weiming Hu 0004, Xiaoqin Zhang 0002, Xi Li 0001 |
Web Intelligence | 2 |
| 2009 | Rank Aggregation Based Text Feature SelectionabstractFiltering feature selection method (filtering method, for short) is a well-known feature selection strategy in pattern recognition and data mining. Filtering method outperforms other feature selection methods in many cases when the dimension of features is large. There are so many filtering methods proposed in previous work leading to the “selection trouble” that how to select an appropriate filtering method for a given text data set. Since to find the best filtering method is usually intractable in real application, this paper takes an alternative path. We propose a feature selection framework that fuses the results obtained by different filtering methods. In fact, deriving a better rank list from different rank lists, known as rank aggregation, is a hot topic studied in many disciplines. Based on the proposed framework and Markov chains rank aggregation techniques, in this paper, we present two new feature selection methods: FR-MC1 and FR-MC4. We also introduce a perturbation algorithm to alleviate the drawbacks of Markov chains rank aggregation techniques. Empirical evaluation on two public text data sets shows that the two new feature selection methods achieve better or comparable results than classical filtering methods, which also demonstrate the effectiveness of our framework. Ou Wu 0001, Haiqiang Zuo, Mingliang Zhu, Weiming Hu 0004, Hanzi Wang |
Web Intelligence | 4 |
| 2009 | Detecting image spam using local invariant features and pyramid match kernelabstractImage spam is a new obfuscating method which spammers invented to more effectively bypass conventional text based spam filters. In this paper, we extract local invariant features of images and run a one-class SVM classifier which uses the pyramid match kernel as the kernel function to detect image spam. Experimental results demonstrate that our algorithm is effective for fighting image spam. Haiqiang Zuo, Weiming Hu 0004, Ou Wu 0001, Yunfei Chen 0002, Guan Luo |
WWW | 2 |
| 2009 | Occlusion Reasoning for Tracking Multiple PeopleabstractOcclusion reasoning is one of the most challenging issues in visual surveillance. In this letter, we propose a new approach for reasoning about occlusions between multiple people. In our approach, occlusion relationships between people are explicitly defined and deduction of the occlusion relationships is integrated into the whole tracking framework. The prior knowledge is supplied by a set of models which include a 2-D elliptical shape model, a spatial-color mixture of Gaussians appearance model, and a motion model with constant velocity. An observation likelihood function is constructed based on the similarity between the observations and the object appearance models with given states. The occlusion relationships are deduced from the current states of the objects and the current observations, using the observation likelihood function. The previous occlusion relationships are not required for deducing the current occlusion relationships. The problem of tracking and occlusion reasoning for more than two people is formulated mathematically, and a solution is proposed based on particle filtering. Experimental results on several real video sequences from indoor and outdoor scenes show the effectiveness of our approach. Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Unsupervised Active Learning Based on Hierarchical Graph-Theoretic ClusteringabstractMost existing active learning approaches are supervised. Supervised active learning has the following problems: inefficiency in dealing with the semantic gap between the distribution of samples in the feature space and their labels, lack of ability in selecting new samples that belong to new categories that have not yet appeared in the training samples, and lack of adaptability to changes in the semantic interpretation of sample categories. To tackle these problems, we propose an unsupervised active learning framework based on hierarchical graph-theoretic clustering. In the framework, two promising graph-theoretic clustering algorithms, namely, dominant-set clustering and spectral clustering, are combined in a hierarchical fashion. Our framework has some advantages, such as ease of implementation, flexibility in architecture, and adaptability to changes in the labeling. Evaluations on data sets for network intrusion detection, image classification, and video classification have demonstrated that our active learning framework can effectively reduce the workload of manual classification while maintaining a high accuracy of automatic classification. It is shown that, overall, our framework outperforms the support-vector-machine-based supervised active learning, particularly in terms of dealing much more efficiently with new samples whose categories have not yet appeared in the training samples. Weiming Hu 0004, Nianhua Xie, Stephen J. Maybank |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2008 | Trajectory-Based Video Retrieval Using Dirichlet Process Mixture ModelsabstractIn this paper, we present a trajectory-based video retrieval framework using Dirichlet process mixture models. The main contribution of this framework is four-fold. (1) We apply a Dirichlet process mixture model (DPMM) to unsupervised trajectory learning. DPMM is a countably infinite mixture model with its components growing by itself. (2) We employ a time-sensitive Dirichlet process mixture model (tDPMM) to learn trajectories ’ time-series characteristics. Furthermore, a novel likelihood estimation algorithm for tDPMM is proposed for the first time. (3) We develop a tDPMM-based probabilistic model matching scheme, which is empirically shown to be more error-tolerating and is able to deliver higher retrieval accuracy than the peer methods in the literature. (4) The framework has a nice scalability and adaptability in the sense that when new cluster data are presented, the framework automatically identifies the new cluster information without having to redo the training. Theoretic analysis and experimental evaluations against the state-of-the-art methods demonstrate the promise and effectiveness of the framework. 1 Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Guan Luo |
BMVC | 2 |
| 2008 | Visual tracking via incremental Log-Euclidean Riemannian subspace learningabstractRecently, a novel Log-Euclidean Riemannian metric is proposed for statistics on symmetric positive definite (SPD) matrices. Under this metric, distances and Riemannian means take a much simpler form than the widely used affine-invariant Riemannian metric. Based on the Log-Euclidean Riemannian metric, we develop a tracking framework in this paper. In the framework, the covariance matrices of image features in the five modes are used to represent object appearance. Since a nonsingular covariance matrix is a SPD matrix lying on a connected Riemannian manifold, the Log-Euclidean Riemannian metric is used for statistics on the covariance matrices of image features. Further, we present an effective online Log-Euclidean Riemannian subspace learning algorithm which models the appearance changes of an object by incrementally learning a low-order Log-Euclidean eigenspace representation through adaptively updating the sample mean and eigenbasis. Tracking is then led by the Bayesian state inference framework in which a particle filter is used for propagating sample distributions over the time. Theoretic analysis and experimental evaluations demonstrate the promise and effectiveness of the proposed framework. Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Mingliang Zhu, Jian Cheng 0002 |
CVPR | 2 |
| 2008 | Sequential particle swarm optimization for visual trackingabstractVisual tracking usually involves an optimization process for estimating the motion of an object from measured images in a video sequence. In this paper, a new evolutionary approach, PSO (particle swarm optimization), is adopted for visual tracking. Since the tracking process is a dynamic optimization problem which is simultaneously influenced by the object state and the time, we propose a sequential particle swarm optimization framework by incorporating the temporal continuity information into the traditional PSO algorithm. In addition, the parameters in PSO are changed adaptively according to the fitness values of particles and the predicted motion of the tracked object, leading to a favourable performance in tracking applications. Furthermore, we show theoretically that, in a Bayesian inference view, the sequential PSO framework is in essence a multilayer importance sampling based particle filter. Experimental results demonstrate that, compared with the state-of-the-art particle filter and its variation - the unscented particle filter, the proposed tracking algorithm is more robust and effective, especially when the object has an arbitrary motion or undergoes large appearance changes. Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank, Xi Li 0001, Mingliang Zhu |
CVPR | 2 |
| 2008 | Robust Visual Tracking Based on an Effective Appearance Model
Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002 |
ECCV (4) | 2 |
| 2008 | Level set tracking with dynamical shape priorsabstractDynamical shape priors are curical for level set-based non- rigid object tracking with noise, occlusions or background clutter. In this paper, we propose a level set tracking framework using dynamical shape priors to capture contours changes of an object in a periodic action sequence. The framework consists of two stages - off-line training and on-line tracking. During the off-line training stage, a graph- based dominant set clustering (DSC) method is applied to learn a shape codebook with each codeword representing a certain shape mode. Then a codeword transition matrix is learnt to characterize the temporal correlations of contours of an object. During the on-line tracking stage, we fuse the knowledge of shape priors and current observations, and adopt maximum a posteriori (MAP) estimation to predict the current shape mode. The experimental results on synthetic and real video sequences demonstrate the effectiveness of our method. Xi Li 0001, Weiming Hu 0004 |
ICIP | 3 |
| 2008 | Key-frame extraction using dominant-set clusteringabstractKey frames play an important role in video abstraction. Clustering is a popular approach for key-frame extraction. In this paper, we propose a novel method for key-frame extraction based on dominant-set clustering. Compared with the existing clustering-based methods, the proposed method dynamically decides the number of key frames depending on the complexity of video shots, produces key frames in a progressive manner and requires less computation. Experimental results on different types of video shots have verified the effectiveness of the method. Xianglin Zeng, Weiming Hu 0004, Wanqing Li 0001, Xiaoqin Zhang 0002 |
ICME | 2 |
| 2008 | Recognition of blue movies by fusion of audio and videoabstractAlong with the explosive growth of the Internet, comes the proliferation of pornography. Compared with the pornographic texts and images, blue movies can do much harm to children, due to the greater realism and voyeurism of blue movies. In this paper, a framework for recognizing blue movies by fusing the audio and video information is described. A one-class Gaussian mixture model (GMM) is used to recognize porno-sounds. A generalized contour-based pornographic image recognition algorithm is used to detect pornographic image frames of a video shot. Then a fusion algorithm based on the Bayes theory is employed to combine the recognition results from audio and video. Experimental results demonstrate that our framework which exploits both audio and video modalities is more robust and achieves better performance than one which uses either one alone. Haiqiang Zuo, Ou Wu 0001, Weiming Hu 0004 |
ICME | 3 |
| 2008 | Group action recognition in soccer videosabstractGroup action recognition in soccer videos is a challenging problem due to the difficulties of group action representation and camera motion estimation. This paper presents a novel approach for recognizing group action with a moving camera. In our approach, ego-motion is estimated by the Kanade-Lucas-Tomasi feature sets on successive frames. The optical flow is then computed on compensated frames. Due to the inaccurate ego-motion estimation, the optical flow can not reflect accurate motion of objects. In this paper, we propose a new motion descriptor which treats the optical flow as spatial patterns and extracts accurate global motion from the noisy optical flow. The latent-dynamic conditional random field model is employed to recognize group action. Experimental results show that our approach is promising. Yu Kong 0001, Xiaoqin Zhang 0002, Qingdi Wei, Weiming Hu 0004, Yunde Jia |
ICPR | 4 |
| 2008 | Multiclass spectral clustering based on discriminant analysisabstractMany existing spectral clustering algorithms share a conventional graph partitioning criterion: normalized cuts (NC). However, one problem with NC is that it poorly captures the graph¿s local marginal information which is very important to graph-based clustering. In this paper, we present a discriminant analysis based graph partitioning criterion (DAC), which is designed to effectively capture the graph¿s local marginal information characterized by the intra-class compactness and the inter-class separability. DAC preserves the intrinsic topological structures of the similarity graph on data points by constructing a k-nearest neighboring subgraph for each data point. Consequently, the clustering results generated by the DAC-based clustering algorithm (DACA) are robust to the outlier disturbance. Theoretic analysis and experimental evaluations demonstrate the promise and effectiveness of DACA. Xi Li 0001, Zhongfei Zhang, Yanguo Wang, Weiming Hu 0004 |
ICPR | 4 |
| 2008 | Boosted cannabis image recognitionabstractWith the large number of Web sites promoting the use of illicit drugs, it has become important to screen these sites for the protection of children on the Internet. Conventional keyword-based approaches are not sufficient because these Web sites often have lots of images and little meaningful words than prices. We propose an AdaBoost-based algorithm for cannabis image recognition. This is the first known attempt at computerized detection of illicit drug Web contents using images. The main technical contributions of our work are two-fold. First, we introduce a novel weak classifier which considers the inherently structural property or ldquoself-similarityrdquo of the cannabis plants. The self-correlation structural characteristics of cannabis can be used as a discriminative property for the purpose of cannabis image recognition. Second, we propose a rapid weak classifier finder, which can efficiently select discriminative weak classifiers from the weak classifier space with little degradation to the classification accuracy. Experiments on real world images have demonstrated improved performance of our method over other methods. Nianhua Xie, Xi Li 0001, Xiaoqin Zhang 0002, Weiming Hu 0004, James Z. Wang 0001 |
ICPR | 4 |
| 2008 | SVD based Kalman particle filter for robust visual trackingabstractObject tracking is one of the most important tasks in computer vision. The unscented particle filter algorithm has been extensively used to tackle this problem and achieved a great success, because it uses the UKF (unscented Kalman filter) to generate a sophisticated proposal distributions which incorporates the newest observations into the state transition distribution and thus overcomes the sample impoverishment problem suffered by the particle filter. However, UKF often encounters the ill-conditioned problem when solving the square root of the covariance matrix in practice. In this paper, we propose a novel Kalman particle filter based on SVD (singular value decomposition), and apply it for visual tracking. Experimental results demonstrate that, compared with the particle filter and the unscented particle filter, the proposed algorithm is more robust in tracking performance. Xiaoqin Zhang 0002, Weiming Hu 0004, Zixiang Zhao, Yanguo Wang, Xi Li 0001, Qingdi Wei |
ICPR | 2 |
| 2008 | Distributed detection of network intrusions based on a parametric modelabstractWith the increasing requirements of fast response and privacy protection, how to detect network intrusions in a distributed architecture becomes a hot research area in the development of modern information security systems. However, it is a challenge to build such a system, given the difficulties brought by the mixed-attribute property of network connection data and the constraints on network communication. In this paper, we present a framework for distributed detection of network intrusions based on a parametric model. The parametric model can explicitly reflect the distributions of different intrusion types and handle the mixed-attribute data naturally. Based on the model, we can generate an accurate global intrusion detector with a very low cost of communication among the distributed detection sites, and no sharing of original network data is needed. Experimental results demonstrate the advantages of the proposed framework in the distributed intrusion detection application. Yanguo Wang, Xi Li 0001, Weiming Hu 0004 |
SMC | 3 |
| 2008 | Recognizing and Filtering Web Images Based on People's ExistenceabstractJudging whether a Web image contains people is useful in both pornographic image recognition and image filtering when searching for images of people. We proposed an approximate but rapid method to solve this problem. For a Web image, three types of probabilities are calculated from the image itself, the imagepsilas associated texts and the title of the Web page is located, respectively. Then a final probability representing the peoplepsilas existence is achieved by fusion of the three probabilistic values. Based on the probability of peoplepsilas existence, we proposed a two-layer framework for pornographic image recognition and a solution of image retrieval respectively. In the experiments conducted, our proposed framework and solution demonstrate good performances in the image recognition and filtering respectively. Ou Wu 0001, Haiqiang Zuo, Weiming Hu 0004, Mingliang Zhu, Shuxiao Li |
Web Intelligence | 3 |
| 2008 | Topic Detection and Tracking for Threaded Discussion CommunitiesabstractThe threaded discussion communities are one of the most common forms of online communities, which are becoming more and more popular among web users. Everyday a huge amount of new discussions are added to these communities, which are difficult to summarize and search. In this paper, we propose a topic detection and tracking (TDT) method for the discussion threads. Most existing TDT methods deal with the news stories, but the language used in discussion data are much more casual, oral and informal compared with news data. To solve this problem, we design several extensions to the basic TDT framework, focusing on the very nature of discussion data, including a thread/post activity validation step, a term pos-weighting strategy, and a two-level decision framework considering not only the content similarity but also the user activity information. Experiment results show that our pro-posed method greatly improves current TDT methods in real discussion community environment. The discussion data can be better organized for searching and visualization with the help of TDT. Mingliang Zhu, Weiming Hu 0004, Ou Wu 0001 |
Web Intelligence | 2 |
| 2008 | User oriented link function classificationabstractCurrently most link-related applications treat all links in the same web page to be identical. One link-related application usually requires one certain property of hyperlinks but actually not all links have this property or they have this property on different levels. Based on a study of how human users judge the links, the idea of the link function classification (LFC) is introduced in this paper. The link functions reflect the purpose that links are created by web page designers and the way they are used by viewers. Links in a certain function class imply one certain relationship between the adjacent pages, and thus they can be assumed to have similar properties. An algorithm is proposed to analyze the link functions based on both vision and structure features which simulates the reaction on the links of human users. Current applications can be enhanced by LFC with a more accurate modeling of the web graph. New mining methods can be also developed by making more and stronger assumptions on links within each function class due to the purer property set they share. Mingliang Zhu, Weiming Hu 0004, Ou Wu 0001, Xi Li 0001, Xiaoqin Zhang 0002 |
WWW | 2 |
| 2008 | AdaBoost-Based Algorithm for Network Intrusion DetectionabstractNetwork intrusion detection aims at distinguishing the attacks on the Internet from normal use of the Internet. It is an indispensable part of the information security system. Due to the variety of network behaviors and the rapid development of attack fashions, it is necessary to develop fast machine-learning-based intrusion detection algorithms with high detection rates and low false-alarm rates. In this correspondence, we propose an intrusion detection algorithm based on the AdaBoost algorithm. In the algorithm, decision stumps are used as weak classifiers. The decision rules are provided for both categorical and continuous features. By combining the weak classifiers for continuous features and the weak classifiers for categorical features into a strong classifier, the relations between these two different types of features are handled naturally, without any forced conversions between continuous and categorical features. Adaptable initial weights and a simple strategy for avoiding overfitting are adopted to improve the performance of the algorithm. Experimental results show that our algorithm has low computational complexity and error rates, as compared with algorithms of higher computational complexity, as tested on the benchmark sample data. Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2007 | Kernel-Bayesian Framework for Object Tracking
Xiaoqin Zhang 0002, Weiming Hu 0004, Guan Luo, Stephen J. Maybank |
ACCV (1) | 2 |
| 2007 | Markov Random Field Modeled Level Sets Method for Object Tracking with Moving Cameras
Weiming Hu 0004, Ying Chen 0018 |
ACCV (1) | 2 |
| 2007 | Robust Visual Tracking Based on Incremental Tensor Subspace LearningabstractMost existing subspace analysis-based tracking algorithms utilize a flattened vector to represent a target, resulting in a high dimensional data learning problem. Recently, subspace analysis is incorporated into the multilinear framework which offline constructs a representation of image ensembles using high-order tensors. This reduces spatio-temporal redundancies substantially, whereas the computational and memory cost is high. In this paper, we present an effective online tensor subspace learning algorithm which models the appearance changes of a target by incrementally learning a low-order tensor eigenspace representation through adaptively updating the sample mean and eigenbasis. Tracking then is led by the state inference within the framework in which a particle filter is used for propagating sample distributions over the time. A novel likelihood function, based on the tensor reconstruction error norm, is developed to measure the similarity between the test image and the learned tensor subspace model during the tracking. Theoretic analysis and experimental evaluations against a state-of-the-art method demonstrate the promise and effectiveness of this algorithm. Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang, Xiaoqin Zhang 0002, Guan Luo |
ICCV | 2 |
| 2007 | Graph Based Discriminative Learning for Robust and Efficient Object TrackingabstractObject tracking is viewed as a two-class 'one-versus-rest' classification problem, in which the sample distribution of the target is approximately Gausian while the background samples are often multimodal. Based on these special properties, we propose a graph embedding based discriminative learning method, in which the topology structures of graphs are carefully designed to reflect the properties of the sample distributions. This method can simultaneously learn the subspace of the target and its local discriminative structure against the background. Moreover, a heuristic negative sample selection scheme is adopted to make the classification more effective. In tracking procedure, the graph based learning is embedded into a Bayesian inference framework cascaded with hierarchical motion estimation, which significantly improves the accuracy and efficiency of the localization. Furthermore, an incremental updating technique for the graphs is developed to capture the changes in both appearance and illumination. Experimental results demonstrate that, compared with two state-of-the-art methods, the proposed tracking algorithm is more efficient and effective, especially in dynamically changing and clutter scenes. Xiaoqin Zhang 0002, Weiming Hu 0004, Stephen J. Maybank, Xi Li 0001 |
ICCV | 2 |
| 2007 | Corner Detection of Contour Images using Spectral ClusteringabstractCorner detection plays an important role in object recognition and motion analysis. In this paper, we propose a hierarchical corner detection framework based on spectral clustering (SC). The framework consists of three stages: contour smoothing, corner cell extraction and corner localization. In the contour smoothing stage, wavelet decomposition is imposed on the raw contour to reduce noise. In the corner cell extraction stage, several atomic corner cells are obtained by SC. In the corner localization stage, the corner points of each corner cell are located by the corner locator based on the kernel-weighted cosine curvature measure. Experimental results demonstrate the superiority of our framework. Xi Li 0001, Weiming Hu 0004, Zhongfei Zhang |
ICIP (3) | 2 |
| 2007 | Dominant Sets-Based Action Recognition using Image Sequence MatchingabstractAction recognition is one of the most active research fields in computer vision. In this paper, we propose a novel method for classifying human actions in a series of image sequences containing certain actions. Human action in image sequences can be recognized by a time-varying contour of human body. We first extract shape context of each contour to form the feature space. Then the dominant sets approach is used for feature clustering and classification to obtain the labeled sequences. Finally, we use a smoothing algorithm upon the labeled sequences to recognize human actions. The proposed dominant sets-based approach has been tested in comparison to three classical methods: K-means, mean shift, and fuzzy-C-mean. Experimental results demonstrate that the dominant sets-based approach achieves the best recognition performance. Moreover, our method is robust to non-rigid deformations, significant scale changes, high action irregularities, and low quality video. Qingdi Wei, Weiming Hu 0004, Xiaoqin Zhang 0002, Guan Luo |
ICIP (6) | 2 |
| 2007 | Customizable Instance-Driven Webpage Filtering Based on Semi-Supervised LearningabstractThe World Wide Web has been growing rapidly in recent years, along with increasing needs for content-based Webpage filtering. But most existing filtering systems cannot easily satisfy the personalized filtering demands from different users at the same time. In this paper, a customizable instance-driven Webpage filtering strategy is proposed. For different users, different Webpage filters are produced by our system through mining the certain Webpage classes they focus on. A semi-supervised learning (SSL) approach is applied for obtaining a precise description of the Webpage class which a user wants to filter based on the small sized user instance set he or she provided. Subsequently, a feature selection step is performed and a Bayes classifier is created over the enlarged training set. Experimental results show the great stability and high performance of our proposed method, and it outperforms existing methods. Mingliang Zhu, Weiming Hu 0004, Xi Li 0001, Ou Wu 0001 |
Web Intelligence | 2 |
| 2007 | Supervised tensor learning
Dacheng Tao, Xuelong Li 0001, Xindong Wu 0001, Weiming Hu 0004, Stephen J. Maybank |
Knowl. Inf. Syst. | 4 |
| 2007 | Recognition of Pornographic Web Pages by Classifying Texts and ImagesabstractWith the rapid development of the World Wide Web, people benefit more and more from the sharing of information. However, Web pages with obscene, harmful, or illegal content can be easily accessed. It is important to recognize such unsuitable, offensive, or pornographic Web pages. In this paper, a novel framework for recognizing pornographic Web pages is described. A C4.5 decision tree is used to divide Web pages, according to content representations, into continuous text pages, discrete text pages, and image pages. These three categories of Web pages are handled, respectively, by a continuous text classifier, a discrete text classifier, and an algorithm that fuses the results from the image classifier and the discrete text classifier. In the continuous text classifier, statistical and semantic features are used to recognize pornographic texts. In the discrete text classifier, the naive Bayes rule is used to calculate the probability that a discrete text is pornographic. In the image classifier, the object's contour-based features are extracted to recognize pornographic images. In the text and image fusion algorithm, the Bayes theory is used to combine the recognition results from images and texts. Experimental results demonstrate that the continuous text classifier outperforms the traditional keyword-statistics-based classifier, the contour-based image classifier outperforms the traditional skin-region-based image classifier, the results obtained by our fusion algorithm outperform those by either of the individual classifiers, and our framework can be adapted to different categories of Web pages. Weiming Hu 0004, Ou Wu 0001, Zhouyao Chen, Zhouyu Fu, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Semantic-Based Surveillance Video RetrievalabstractVisual surveillance produces large amounts of video data. Effective indexing and retrieval from surveillance video databases are very important. Although there are many ways to represent the content of video clips in current video retrieval algorithms, there still exists a semantic gap between users and retrieval systems. Visual surveillance systems supply a platform for investigating semantic-based video retrieval. In this paper, a semantic-based video retrieval framework for visual surveillance is proposed. A cluster-based tracking algorithm is developed to acquire motion trajectories. The trajectories are then clustered hierarchically using the spatial and temporal information, to learn activity models. A hierarchical structure of semantic indexing and retrieval of object activities, where each individual activity automatically inherits all the semantic descriptions of the activity model to which it belongs, is proposed for accessing video clips and individual objects at the semantic level. The proposed retrieval framework supports various queries including queries by keywords, multiple object queries, and queries by sketch. For multiple object queries, succession and simultaneity restrictions, together with depth and breadth first orders, are considered. For sketch-based queries, a method for matching trajectories drawn by users to spatial trajectories is proposed. The effectiveness and efficiency of our framework are tested in a crowded traffic scene. Weiming Hu 0004, Zhouyu Fu, Wenrong Zeng, Stephen J. Maybank |
IEEE Trans. Image Process. | 1 |
| 2006 | Indexing and Matching of Video Shots Based on Motion and Color AnalysisabstractThis paper concerns two fundamental issues in video shots retrieval: key frame identification and similarity measurement between the key frames. We propose a simple key frame extraction algorithm based on optical flow. The algorithm emphasizes the motion extremum in the shot. Color histograms in HSV color space are adopted to describe the content of the extracted key frames and a new model is proposed to measure the similarity between the key frames from different shots. Preliminary experiments have shown that the proposed method outperforms existing ones in retrieving sport video shots Ying Chen 0018, Weiming Hu 0004, Xianglin Zeng, Wanqing Li 0001 |
ICARCV | 2 |
| 2006 | A Novel Web Page Filtering System by Combining Texts and ImagesabstractWith the rapid development of the Internet, people benefit much from the sharing of information. Meanwhile, the WWW era is a double-edged sword which spreads harmful and erotic content widely. In this paper, a new statistical approach has been exploited by combining the results of two or more different classification methods using our filtering system. We first briefly introduce the classification of discrete texts, continuous texts and images separately, and then describe the specific way we have been exploring to merge the text and image classification result. Also there is a section illustrating our system framework. Finally we assess our method by demonstrating the experimental results and comparing it to some common-used filtering methods Zhouyao Chen, Ou Wu 0001, Mingliang Zhu, Weiming Hu 0004 |
Web Intelligence | 4 |
| 2006 | Principal Axis-Based Correspondence between Multiple Cameras for People TrackingabstractVisual surveillance using multiple cameras has attracted increasing interest in recent years. Correspondence between multiple cameras is one of the most important and basic problems which visual surveillance using multiple cameras brings. In this paper, we propose a simple and robust method, based on principal axes of people, to match people across multiple cameras. The correspondence likelihood reflecting the similarity of pairs of principal axes of people is constructed according to the relationship between "ground-points" of people detected in each camera view and the intersections of principal axes detected in different camera views and transformed to the same view. Our method has the following desirable properties: 1) Camera calibration is not needed. 2) Accurate motion detection and segmentation are less critical due to the robustness of the principal axis-based feature to noise. 3) Based on the fused data derived from correspondence results, positions of people in each camera view can be accurately located even when the people are partially occluded in all views. The experimental results on several real video sequences from outdoor environments have demonstrated the effectiveness, efficiency, and robustness of our method. Weiming Hu 0004, Tieniu Tan, Jianguang Lou, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | A System for Learning Statistical Motion PatternsabstractAnalysis of motion patterns is an effective approach for anomaly detection and behavior prediction. Current approaches for the analysis of motion patterns depend on known scenes, where objects move in predefined ways. It is highly desirable to automatically construct object motion patterns which reflect the knowledge of the scene. In this paper, we present a system for automatically learning motion patterns for anomaly detection and behavior prediction based on a proposed algorithm for robustly tracking multiple objects. In the tracking algorithm, foreground pixels are clustered using a fast accurate fuzzy K-means algorithm. Growing and prediction of the cluster centroids of foreground pixels ensure that each cluster centroid is associated with a moving object in the scene. In the algorithm for learning motion patterns, trajectories are clustered hierarchically using spatial and temporal information and then each motion pattern is represented with a chain of Gaussian distributions. Based on the learned statistical motion patterns, statistical methods are used to detect anomalies and predict behaviors. Our system is tested using image sequences acquired, respectively, from a crowded real traffic scene and a model traffic scene. Experimental results show the robustness of the tracking algorithm, the efficiency of the algorithm for learning motion patterns, and the encouraging performance of algorithms for anomaly detection and behavior prediction. Weiming Hu 0004, Xuejuan Xiao, Zhouyu Fu, Tieniu Tan, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Supervised Tensor LearningabstractThis paper aims to take general tensors as inputs for supervised learning. A supervised tensor learning (STL) framework is established for convex optimization based learning techniques such as support vector machines (SVM) and minimax probability machines (MPM). Within the STL framework, many conventional learning machines can be generalized to take n/sup th/-order tensors as inputs. We also study the applications of tensors to learning machine design and feature extraction by linear discriminant analysis (LDA). Our method for tensor based feature extraction is named the tenor rank-one discriminant analysis (TR1DA). These generalized algorithms have several advantages: 1) reduce the curse of dimension problem in machine learning and data mining; 2) avoid the failure to converge; and 3) achieve better separation between the different categories of samples. As an example, we generalize MPM to its STL version, which is named the tensor MPM (TMPM). TMPM learns a series of tensor projections iteratively. It is then evaluated against the original MPM. Our experiments on a binary classification problem show that TMPM significantly outperforms the original MPM. Dacheng Tao, Xuelong Li 0001, Weiming Hu 0004, Stephen J. Maybank, Xindong Wu 0001 |
ICDM | 3 |
| 2005 | Similarity based vehicle trajectory clustering and anomaly detectionabstractIn this paper, we proposed a hierarchical clustering framework to classify vehicle motion trajectories in real traffic video based on their pairwise similarities. First raw trajectories are pre-processed and resampled at equal space intervals. Then spectral clustering is used to group trajectories with similar spatial patterns. Dominant paths and lanes can be distinguished as a result of two-layer hierarchical clustering. Detection of novel trajectories is also possible based on the clustering results. Experimental results demonstrate the superior performance of spectral clustering compared with conventional fuzzy K-means clustering and some results of anomaly detection are presented. Zhouyu Fu, Weiming Hu 0004, Tieniu Tan |
ICIP (2) | 2 |
| 2005 | Kernel Principle Component Analysis in Pixels ClusteringabstractWe propose two new methods in the nonlinear kernel feature space for pixel clustering based on the traditional KMeans and Gaussian mixture model (GMM). Unlike the previous work on the kernel machines, we give out a new perspective on the new developed kernel machines. That is, kernel principle component analysis (KPCA) combined with the KMeans and the GMM are kernel KMeans (KKMeans) and kernel GMM (KGMM), respectively. In this paper, we prove the new perspective on KKMeans and give out a clear statement on the KGMM as well. Based on this new perspectives, we can implement the KKMeans and the KGMM conveniently. At the end of the paper, we utilize these new algorithms on the problem of the colour image segmentation. Based on a series of experimental results on Corel colour images, we find that the KKMeans and KGMM can outperform the traditional KMeans and GMM consistently, respectively. Jing Li 0027, Dacheng Tao, Weiming Hu 0004, Xuelong Li 0001 |
Web Intelligence | 3 |
| 2005 | Stable Third-Order Tensor Representation for Color Image ClassificationabstractGeneral tensors can represent colour images more naturally than conventional features; however, the general tensors' stability properties are not reported and remain to be a key problem. In this paper, we use the tensor minimax probability (TMPM) to prove that the tensor representation is stable. The proof is based on the random subspace method through a large number of experiments. Dacheng Tao, Stephen J. Maybank, Weiming Hu 0004, Xuelong Li 0001 |
Web Intelligence | 3 |
| 2005 | 3-D Model-Based Vehicle TrackingabstractThis paper aims at tracking vehicles from monocular intensity image sequences and presents an efficient and robust approach to three-dimensional (3-D) model-based vehicle tracking. Under the weak perspective assumption and the ground-plane constraint, the movements of model projection in the two-dimensional image plane can be decomposed into two motions: translation and rotation. They are the results of the corresponding movements of 3-D translation on the ground plane (GP) and rotation around the normal of the GP, which can be determined separately. A new metric based on point-to-line segment distance is proposed to evaluate the similarity between an image region and an instantiation of a 3-D vehicle model under a given pose. Based on this, we provide an efficient pose refinement method to refine the vehicle's pose parameters. An improved EKF is also proposed to track and to predict vehicle motion with a precise kinematics model. Experimental results with both indoor and outdoor data show that the algorithm obtains desirable performance even under severe occlusion and clutter. Jianguang Lou, Tieniu Tan, Weiming Hu 0004, Hao Yang 0010, Stephen J. Maybank |
IEEE Trans. Image Process. | 3 |
| 2004 | Gait analysis for human identification in frequency domainabstractIn this paper, we analyze the spatio-temporal human characteristic of moving silhouettes in frequency domain, and find key Fourier descriptors that have better discriminatory capability for recognition than the other Fourier descriptors. A large number of experimental results and analysis show that the proposed algorithm based on the key Fourier descriptors can not only greatly reduce the gait data dimensionality, but also lighten the computation cost, with a satisfactory CCR. Besides that, classification performance can be further improved using feature fusion. Shiqi Yu 0001, Liang Wang 0001, Weiming Hu 0004, Tieniu Tan |
ICIG | 3 |
| 2004 | Multi-camera correspondence based on principal axis of human bodyabstractMulticamera correspondence of moving people is a relatively new issue in computer vision. To cope with it, we propose a simple but effective method based on the principal axis of human body. We apply the method to real video sequences in outdoor environments. The experimental results have demonstrated the efficiency of the proposed method. Jianguang Lou, Weiming Hu 0004, Tieniu Tan |
ICIP | 3 |
| 2004 | Semantic-based traffic video retrieval using activity pattern analysisabstractA semantic based retrieval framework for traffic video sequences is proposed. In order to estimate the low-level motion data, a cluster tracking algorithm is developed. A novel hierarchical self-organizing map is applied to learn the activity patterns. By using activity pattern analysis and semantic concepts assignment, a set of activity models is generated, which is used as the indexing key for accessing video clips and individual vehicles in the semantic level. The proposed retrieval framework supports various queries including query by keywords, query by sketch and multiple object queries. Weiming Hu 0004, Tieniu Tan, Junyi Peng |
ICIP | 2 |
| 2004 | Adaptive skin detection using multiple cuesabstractThis paper presents an adaptive approach to skin detection. First, we propose a nonlinear relationship among R, G and B components and use a closed curve to identify the skin cluster region. Then, a split machine is designed that aids the extraction of the pixels with similar low-level features from images. Finally, a nonlinear skin color classifier with an adaptive threshold is developed by analyzing the properties of the extracted pixels in the HSL, YCbCr, YUV and YIQ color spaces. Experimental results show that our proposed method works very well in skin detection. Jinfeng Yang, Zhouyu Fu, Tieniu Tan, Weiming Hu 0004 |
ICIP | 4 |
| 2004 | Kinematics-based tracking of human walking in monocular video sequences
Huazhong Ning, Tieniu Tan, Liang Wang 0001, Weiming Hu 0004 |
Image Vis. Comput. | 4 |
| 2004 | People tracking based on motion model and motion constraints with automatic initialization
Huazhong Ning, Tieniu Tan, Liang Wang 0001, Weiming Hu 0004 |
Pattern Recognit. | 4 |
| 2004 | Fusion of static and dynamic body biometrics for gait recognitionabstractVision-based human identification at a distance has recently gained growing interest from computer vision researchers. This paper describes a human recognition algorithm by combining static and dynamic body biometrics. For each sequence involving a walker, temporal pose changes of the segmented moving silhouettes are represented as an associated sequence of complex vector configurations and are then analyzed using the Procrustes shape analysis method to obtain a compact appearance representation, called static information of body. In addition, a model-based approach is presented under a Condensation framework to track the walker and to further recover joint-angle trajectories of lower limbs, called dynamic information of gait. Both static and dynamic cues obtained from walking video may be independently used for recognition using the nearest exemplar classifier. They are fused on the decision level using different combinations of rules to improve the performance of both identification and verification. Experimental results of a dataset including 20 subjects demonstrate the feasibility of the proposed algorithm. Liang Wang 0001, Huazhong Ning, Tieniu Tan, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2004 | A hierarchical self-organizing approach for learning the patterns of motion trajectoriesabstractThe understanding and description of object behaviors is a hot topic in computer vision. Trajectory analysis is one of the basic problems in behavior understanding, and the learning of trajectory patterns that can be used to detect anomalies and predict object trajectories is an interesting and important problem in trajectory analysis. In this paper, we present a hierarchical self-organizing neural network model and its application to the learning of trajectory distribution patterns for event recognition. The distribution patterns of trajectories are learnt using a hierarchical self-organizing neural network. Using the learned patterns, we consider anomaly detection as well as object behavior prediction. Compared with the existing neural network structures that are used to learn patterns of trajectories, our network structure has smaller scale and faster learning speed, and is thus more effective. Experimental results using two different sets of data demonstrate the accuracy and speed of our hierarchical self-organizing neural network in learning the distribution patterns of object trajectories. Weiming Hu 0004, Tieniu Tan |
IEEE Trans. Neural Networks | 1 |
| 2004 | A survey on visual surveillance of object motion and behaviorsabstractVisual surveillance in dynamic scenes, especially for humans and vehicles, is currently one of the most active research topics in computer vision. It has a wide spectrum of promising applications, including access control in special areas, human identification at a distance, crowd flux statistics and congestion analysis, detection of anomalous behaviors, and interactive surveillance using multiple cameras, etc. In general, the processing framework of visual surveillance in dynamic scenes includes the following stages: modeling of environments, detection of motion, classification of moving objects, tracking, understanding and description of behaviors, human identification, and fusion of data from multiple cameras. We review recent developments and general strategies of all these stages. Finally, we analyze possible research directions, e.g., occlusion handling, a combination of twoand three-dimensional tracking, a combination of motion analysis and biometrics, anomaly detection and behavior prediction, content-based retrieval of surveillance videos, behavior understanding and natural language description, fusion of information from multiple sensors, and remote surveillance. Weiming Hu 0004, Tieniu Tan, Liang Wang 0001, Stephen J. Maybank |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2004 | Learning activity patterns using fuzzy self-organizing neural networkabstractActivity understanding in visual surveillance has attracted much attention in recent years. In this paper, we present a new method for learning patterns of object activities in image sequences for anomaly detection and activity prediction. The activity patterns are constructed using unsupervised learning of motion trajectories and object features. Based on the learned activity patterns, anomaly detection and activity prediction can be achieved. Unlike existing neural network based methods, our method uses a whole trajectory as an input to the network. This makes the network structure much simpler. Furthermore, the fuzzy set theory based method and the batch learning method are introduced into the network learning process, and make the learning process much more efficient. Two sets of data acquired, respectively, from a model scene and a campus scene are both used to test the proposed algorithms. Experimental results show that the fuzzy self-organizing neural network (fuzzy SOM) is much more efficient than the Kohonen self-organizing feature map (SOFM) and vector quantization in both speed and accuracy, and the anomaly detection and activity prediction algorithms have encouraging performances. Weiming Hu 0004, Tieniu Tan, Stephen J. Maybank |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2003 | Pose Evaluation Based on Bayesian classification ErrorabstractPose evaluation is a fundamental issue in image processing and computer vision. In this paper, we propose a new method called BCE for pose evaluation based on Bayesian classification error. Various image cues are incorporated to depict an object including object shape, side region statistics and temporal information. Then a PEF (Pose Evaluation Function) is constructed based on Bayesian classification error, and an efficient algorithm to calculate it is developed. We test our new method with real outdoor image sequences, and use two criteria to compare it with two other representative ones. It is shown that our new method leads to better performance with respect to localization accuracy and robustness against general clutter and occlusion. 1 Jianguang Lou, Weiming Hu 0004, Tieniu Tan |
BMVC | 3 |
| 2003 | Fusion of Static and Dynamic Body Biometrics for Gait RecognitionabstractHuman identification at a distance has recently gained growing interest from computer vision researchers. This paper aims to propose a visual recognition algorithm based upon fusion of static and dynamic body biometrics. For each sequence involving a walking figure, pose changes of the segmented moving silhouettes are represented as an associated sequence of complex vector configurations, and are then analyzed using the Procrustes shape analysis method to obtain a compact appearance representation, called static information of body. Also, a model-based approach is presented under a condensation framework to track the walker and to recover joint-angle trajectories of lower limbs, called dynamic information of gait. Both static and dynamic cues are respectively used for recognition using the nearest exemplar classifier. They are also effectively fused on decision level using different combination rules to improve the performance of both identification and verification. Experimental results on a dataset including 20 subjects demonstrate the validity of the proposed algorithm. Liang Wang 0001, Huazhong Ning, Tieniu Tan, Weiming Hu 0004 |
ICCV | 4 |
| 2003 | Silhouette Analysis-Based Gait Recognition for Human IdentificationabstractHuman identification at a distance has recently gained growing interest from computer vision researchers. Gait recognition aims essentially to address this problem by identifying people based on the way they walk. In this paper, a simple but efficient gait recognition algorithm using spatial-temporal silhouette analysis is proposed. For each image sequence, a background subtraction algorithm and a simple correspondence procedure are first used to segment and track the moving silhouettes of a walking figure. Then, eigenspace transformation based on principal component analysis (PCA) is applied to time-varying distance signals derived from a sequence of silhouette images to reduce the dimensionality of the input feature space. Supervised pattern classification techniques are finally performed in the lower-dimensional eigenspace for recognition. This method implicitly captures the structural and transitional characteristics of gait. Extensive experimental results on outdoor image sequences demonstrate that the proposed algorithm has an encouraging recognition performance with relatively low computational cost. Liang Wang 0001, Tieniu Tan, Huazhong Ning, Weiming Hu 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2003 | Recent developments in human motion analysis
Liang Wang 0001, Weiming Hu 0004, Tieniu Tan |
Pattern Recognit. | 2 |
| 2003 | Automatic gait recognition based on statistical shape analysisabstractGait recognition has recently gained significant attention from computer vision researchers. This interest is strongly motivated by the need for automated person identification systems at a distance in visual surveillance and monitoring applications. The paper proposes a simple and efficient automatic gait recognition algorithm using statistical shape analysis. For each image sequence, an improved background subtraction procedure is used to extract moving silhouettes of a walking figure from the background. Temporal changes of the detected silhouettes are then represented as an associated sequence of complex vector configurations in a common coordinate frame, and are further analyzed using the Procrustes shape analysis method to obtain mean shape as gait signature. Supervised pattern classification techniques, based on the full Procrustes distance measure, are adopted for recognition. This method does not directly analyze the dynamics of gait, but implicitly uses the action of walking to capture the structural characteristics of gait, especially the shape cues of body biometrics. The algorithm is tested on a database consisting of 240 sequences from 20 different subjects walking at 3 viewing angles in an outdoor environment. Experimental results are included to demonstrate the encouraging performance of the proposed algorithm. Liang Wang 0001, Tieniu Tan, Weiming Hu 0004, Huazhong Ning |
IEEE Trans. Image Process. | 3 |
| 2002 | Gait recognition based on Procrustes shape analysisabstractGait recognition has recently attracted increasing attention, especially in vision-based human identification-at-a-distance in visual surveillance. The paper proposes a simple but efficient gait recognition algorithm, based on statistical shape analysis. For each gait sequence, a background subtraction procedure is used to segment spatial silhouettes of the walking figures from the background. Static pose changes of these silhouettes over time are represented as a sequence of associated complex configurations in a common coordinate, and are then analyzed using the Procrustes shape analysis method to obtain a gait signature. The k-nearest neighbor classifier and the nearest exemplar classifier based on the full Procrustes distance measure are adopted for recognition. Experimental results demonstrate that the proposed algorithm has an encouraging recognition performance. Liang Wang 0001, Huazhong Ning, Weiming Hu 0004, Tieniu Tan |
ICIP (3) | 3 |
| 2002 | Articulated Model Based People Tracking Using Motion ModelsabstractThis paper focuses on acquisition of human motion data such as joint angles and velocity for applications of virtual reality, using both an articulated body model and a motion model in the CONDENSATION framework. Firstly, we learn a motion model represented by Gaussian distributions, and explore motion constraints by considering the dependency of motion parameters and represent them as conditional distributions. Both are integrated into the dynamic model to concentrate factored sampling in the areas of state-space with most posterior information. To measure the observing density with accuracy and robustness, a PEF (pose evaluation function) modeled with a radial term is proposed. We also address the issue of automatic acquisition of initial model posture and recovery from severe failures. A large number of experiments on several persons demonstrate that our approach works well. Huazhong Ning, Liang Wang 0001, Weiming Hu 0004, Tieniu Tan |
ICMI | 3 |
| 2001 | Efficient and robust vehicle localizationabstractWe propose a novel algorithm for efficient and robust pose determination of vehicles in traffic scenes from single monocular intensity images using calibrated cameras. We consider the pose determination process as a series of evolutions from initial pose to correct pose in 3D space, which can be decomposed into two independent 3D motions: translation and rotation. The translation parameters are obtained based on point-to-line-segment distance (PLS distance), while the rotation parameters are determined by geometric relationships among a set of specially constructed but imaginary planes. Closed-form solutions to both sub-problems are obtained, thus avoiding the usual shortcoming of relatively high computational cost of traditional 3D-model based approaches. In addition, vertex neighborhood constraint (VNC) is introduced to improve the robustness of the method. Experimental results show that the algorithm works well even under severe occlusion and clutter. Hao Yang 0010, Jianguang Lou, Hongzan Sun, Weiming Hu 0004, Tieniu Tan |
ICIP (2) | 4 |