Guorong Li

dblp:28/4782 · DBLP profile ↗
← Back
111ranked-venue papers
13as first author
59since 2021 · last 2026
0000-0003-3954-2387ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 75 · 7 first-author · 34 since 2021Artificial intelligence and machine learning · 47 · 5 first-author · 29 since 2021Databases, data management, data science and information retrieval · 8 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Computer networks · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Mamba-Based Temporally Guided Spatial Alignment for Unaligned RGB-T Tracking
Guorong Li, Chen Zhang 0013, Wentao Cao, Laiyun Qing
ICIC (1)2
2026 Synergistic Dual-Graph Co-Evolutionary Network for point-supervised temporal action localization
Laiyun Qing, Guorong Li, Qingming Huang
Comput. Vis. Image Underst.4
2026 A Multi-Modal Knowledge-Driven Approach for Generalized Zero-shot Video Classification
Mingyao Hong, Xinfeng Zhang 0001, Guorong Li, Qingming Huang
Int. J. Comput. Vis.3
2026 SeqCount: A sequence modeling framework for class-agnostic counting
abstract
Class-agnostic counting aims to count the number of objects in any category with only a few exemplars. It is crucial for solving the challenge of counting any visual class without re-finetuning, which in turn lowers deployment costs across diverse scenarios. Existing methods count the number of exemplar objects by integrating them over the density map smoothed by Gaussian kernels. However, designing generic kernels to generate density maps is challenging due to the different sizes and shapes of objects. To solve this problem, we propose SeqCount, which eliminates the need for density maps and treats object counting as a sequence generation problem. Specifically, we consider an input image as $N\times N$ patches and propose a serialization scheme. Then, we use an encoding-decoding structure to exploit the correlation among the patches. Experimental results on five challenging datasets demonstrate that our method performs favorably against the state-of-the-art models.
Guorong Li, Xinyan Liu 0008, Zhenjun Han, Yuankai Qi
J. Vis. Commun. Image Represent.2
2026 Consistency-Aware Anchor Pyramid Network for Crowd Localization
abstract
Crowd localization aims to predict the positions of humans in images of crowded scenes. While existing methods have made significant progress, two primary challenges remain: (i) a fixed number of evenly distributed anchors can cause excessive or insufficient predictions across regions in an image with varying crowd densities, and (ii) ranking inconsistency of predictions between the testing and training phases leads to the model being sub-optimal in inference. To address these issues, we propose a Consistency-Aware Anchor Pyramid Network (CAAPN) comprising two key components: an Adaptive Anchor Generator (AAG) and a Localizer with Augmented Matching (LAM). The AAG module adaptively generates anchors based on estimated crowd density in local regions to alleviate the anchor deficiency or excess problem. It also considers the spatial distribution prior to heads for better performance. The LAM module is designed to augment the predictions which are used to optimize the neural network during training by introducing an extra set of target candidates and correctly matching them to the ground truth. The proposed method achieves favorable performance against state-of-the-art approaches on five challenging datasets: ShanghaiTech A and B, UCF-QNRF, JHU-CROWD++, and NWPU-Crowd.
Xinyan Liu 0008, Guorong Li, Yuankai Qi, Zhenjun Han, Anton van den Hengel, Nicu Sebe, Ming-Hsuan Yang 0001, Qingming Huang
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 SAPNet++: Evolving Point-Prompted Instance Segmentation With Semantic and Spatial Awareness
abstract
Single-point annotation is increasingly prominent in visual tasks for labeling cost reduction. However, it challenges tasks requiring high precision, such as the point-prompted instance segmentation (PPIS) task, which aims to estimate precise masks using single-point prompts to train a segmentation network. Due to the constraints of point annotations, granularity ambiguity and boundary uncertainty arise i.e., the difficulty distinguishing between different levels of detail (e.g., whole object vs. parts) and the challenge of precisely delineating object boundaries. Previous works have usually inherited the paradigm of mask generation along with proposal selection to achieve PPIS. However, proposal selection relies solely on category information, failing to resolve the ambiguity of different granularity. Furthermore, mask generators offer only finite discrete solutions that often deviate from actual masks, particularly at boundaries. To address these issues, we propose the Semantic-Aware Point-Prompted Instance Segmentation Network (SAPNet). It integrates Point Distance Guidance and Box Mining Strategy to tackle group and local issues caused by the point's granularity ambiguity. Additionally, we incorporate completeness scores within proposals to add spatial granularity awareness, enhancing multiple instance learning (MIL) in proposal selection termed S-MIL. The Multi-level Affinity Refinement conveys pixel and semantic clues, narrowing boundary uncertainty during mask refinement. These modules culminate in SAPNet++, mitigating point prompt's granularity ambiguity and boundary uncertainty and significantly improving segmentation performance. Extensive experiments on four challenging datasets validate the effectiveness of our methods, highlighting the potential to advance PPIS.
Zhaoyang Wei, Xumeng Han, Xuehui Yu, Xue Yang 0005, Guorong Li, Zhenjun Han, Jianbin Jiao
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Dynamic example network for class-agnostic object counting
Xinyan Liu 0008, Guorong Li, Yuankai Qi, Ziheng Yan, Weigang Zhang, Laiyun Qing, Qingming Huang
Pattern Recognit.2
2026 RETTA: Retrieval-enhanced test-time adaptation for zero-shot video captioning
abstract
Despite the significant progress of fully-supervised video captioning, zero-shot methods remain much less explored. In this paper, we propose a novel zero-shot video captioning framework named R etrieval- E nhanced T est- T ime A daptation (RETTA), which takes advantage of existing pre-trained large-scale vision and language models to directly generate captions with test-time adaptation. Specifically, we bridge video and text using four key models: a general video-text retrieval model XCLIP, a general image-text matching model CLIP, a text alignment model AnglE, and a text generation model GPT-2, due to their source-code availability. The main challenge is how to enable the text generation model to be sufficiently aware of the content in a given video so as to generate corresponding captions. To address this problem, we propose using learnable tokens as a communication medium among these four frozen models GPT-2, XCLIP, CLIP, and AnglE. Different from the conventional way that trains these tokens with training data, we propose to learn these tokens with soft targets of the inference data under several carefully crafted loss functions, which enable the tokens to absorb video information catered for GPT-2. This adaptation requires only a few iterations ( e.g. , 16) and does not require ground truth data. Extensive experimental on MSR-VTT, MSVD, and VATEX, show absolute 5.1 % ∼ 32.4 % improvements in CIDEr scores compared to several state-of-the-art zero-shot video captioning methods.
Yunchuan Ma, Laiyun Qing, Guorong Li, Yuankai Qi, Amin Beheshti, Quan Z. Sheng, Qingming Huang
Pattern Recognit.3
2026 Compactness driven Co-learning for crowd counting and localization
Ziheng Yan, Xinyan Liu 0008, Guorong Li, Weigang Zhang, Fang Wan 0001, Qingming Huang
Pattern Recognit.3
2025 MambaLCT: Boosting Tracking via Long-term Context State Space Model
abstract
Effectively constructing context information with long-term dependencies from video sequences is crucial for object tracking. However, the context length constructed by existing work is limited, only considering object information from adjacent frames or video clips, leading to insufficient utilization of contextual information. To address this issue, we propose MambaLCT, which constructs and utilizes target variation cues from the first frame to the current frame for robust tracking. First, a novel unidirectional Context Mamba module is designed to scan frame features along the temporal dimension, gathering target change cues throughout the entire sequence. Specifically, target-related information in frame features is compressed into a hidden state space through a selective scanning mechanism. The target information across the entire video is continuously aggregated into target variation cues. Next, we inject the target change cues into the attention mechanism, providing temporal information for modeling the relationship between the template and search frames. The advantage of MambaLCT is its ability to continuously extend the length of the context, capturing complete target change cues, which enhances the stability and robustness of the tracker. Extensive experiments show that long-term context information enhances the model's ability to perceive targets in complex scenarios. MambaLCT achieves new SOTA performance on six benchmarks while maintaining real-time runing speeds.
Xiaohai Li, Bineng Zhong 0001, Qihua Liang, Guorong Li, Zhiyi Mo, Shuxiang Song 0001
AAAI4
2025 Less Is More: Token Context-Aware Learning for Object Tracking
abstract
Recently, several studies have shown that utilizing contextual information to perceive target states is crucial for object tracking. They typically capture context by incorporating multiple video frames. However, these naive frame-context methods fail to consider the importance of each patch within a reference frame, making them susceptible to noise and redundant tokens, which deteriorates tracking performance. To address this challenge, we propose a new token context-aware tracking pipeline named LMTrack, designed to automatically learn high-quality reference tokens for efficient visual tracking. Embracing the principle of Less is More, the core idea of LMTrack is to analyze the importance distribution of all reference tokens, where important tokens are collected, continually attended to, and updated. Specifically, a novel Token Context Memory module is designed to dynamically collect high-quality spatio-temporal information of a target in an autoregressive manner, eliminating redundant background tokens from the reference frames. Furthermore, an effective Unidirectional Token Attention mechanism is designed to establish dependencies between reference tokens and search frame, enabling robust cross-frame association and target localization. Extensive experiments demonstrate the superiority of our tracker, achieving state-of-the-art results on tracking benchmarks such as GOT-10K, TrackingNet, and LaSOT.
Chenlong Xu, Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Guorong Li, Shuxiang Song 0001
AAAI5
2025 Subpart Suppression Network for Few-Shot Object Counting
abstract
Few-shot object counting and detection aim to count objects along with their bounding boxes specified by exemplar bounding boxes. Current mainstream methods predict density maps by applying similarity between exemplar and image features to get counting results and detecting peak points from the density map as the positions of objects. However, the sub-parts of objects can also have high similarity to the exemplars, harming the counting and detection performance. To address these issues, we propose a two-stage Subpart Suppression Network (SSN), consisting of a Subpart Suppression Density map Predictor (SSDP), which enforces the model focus on whole objects rather than subparts, and a SAM-based Detector and Verifier, which refines the final prediction by a clustering method. Extensive experiments on two popular datasets show an advantage in performance over the state-of-the-art methods and prove the components’ effectiveness.
Lanxin Liu, Xinyan Liu 0008, Guorong Li
ICASSP3
2025 SDVPT: Semantic-Driven Visual Prompt Tuning for Open-world Object Counting
abstract
Open-world object counting leverages the robust text-image alignment of pre-trained vision-language models (VLMs) to enable counting of arbitrary categories in images specified by textual queries. However, widely adopted naive fine-tuning strategies concentrate exclusively on text-image consistency for categories contained in, which leads to limited generalizability for unseen categories. In this work, we propose a plug-and-play Semantic-Driven Visual Prompt Tuning framework (SDVPT) that transfers knowledge from the training set to unseen categories with minimal overhead in parameters and inference time. First, we introduce a two-stage visual prompt learning strategy composed of Category-Specific Prompt Initialization (CSPI) and Topology-Guided Prompt Refinement (TGPR). The CSPI generates category-specific visual prompts, and then TGPR distills latent structural patterns from the VLM's text encoder to refine these prompts. During inference, we dynamically synthesize the visual prompts for unseen categories based on the semantic correlation between unseen and training categories, facilitating robust text-image alignment for unseen categories. Extensive experiments integrating SDVPT with all available open-world object counting models demonstrate its effectiveness and adaptability across three widely used datasets: FSC-147, CARPK, and PUCPR+. Code is available https://github.com/Eamon-0v0/SDVPT
Guorong Li, Laiyun Qing, Amin Beheshti, Jian Yang 0001, Quan Z. Sheng, Yuankai Qi, Qingming Huang
ACM Multimedia2
2025 P2Object: Single Point Supervised Object Detection and Instance Segmentation
Pengfei Chen 0004, Xuehui Yu, Xumeng Han, Kuiran Wang, Guorong Li, Lingxi Xie, Zhenjun Han, Jianbin Jiao
Int. J. Comput. Vis.5
2025 Towards Universal Modal Tracking With Online Dense Temporal Token Learning
abstract
We propose a universal video-level modality-awareness tracking model with online dense temporal token learning (called UM-ODTrack). It is designed to support various tracking tasks, including RGB, RGB+Thermal, RGB+Depth, and RGB+Event, utilizing the same model architecture and parameters. Specifically, our model is designed with three core goals: Video-level Sampling. We expand the model's inputs to a video sequence level, aiming to see a richer video context from an near-global perspective. Video-level Association. Furthermore, we introduce two simple yet effective online dense temporal token association mechanisms to propagate the appearance and motion trajectory information of target via a video stream manner. Modality Scalable. We propose two novel gated perceivers that adaptively learn cross-modal representations via a gated attention mechanism, and subsequently compress them into the same set of model parameters via a one-shot training manner for multi-task inference. This new solution brings the following benefits: (i) The purified token sequences can serve as temporal prompts for the inference in the next video frames, whereby previous information is leveraged to guide future inference. (ii) Unlike multi-modal trackers that require independent training, our one-shot training scheme not only alleviates the training burden, but also improves model representation. Extensive experiments on visible and multi-modal benchmarks show that our UM-ODTrack achieves a new SOTA performance.
Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Guorong Li, Xianxian Li, Rongrong Ji
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 ClickTrack: Towards real-time interactive single object tracking
Kuiran Wang, Xuehui Yu, Wenwen Yu, Guorong Li, Xiangyuan Lan, Qixiang Ye, Jianbin Jiao, Zhenjun Han
Pattern Recognit.4
2025 MMA: Video Reconstruction for Spike Camera Based on Multiscale Temporal Modeling and Fine-Grained Attention
abstract
This paper presents a Multiscale Temporal Correlation Learning with the Mamba-Fused Attention Model (MMA), an efficient and effective method for reconstructing a video clip from a spike stream. Spike cameras offer unique advantages for capturing rapid scene changes with high temporal resolution. A spike stream contains sufficient information for multiple image reconstructions. However, existing methods generate only a single image at a time for a given spike stream, which results in excessive redundant computations between consecutive frames when aiming at restoring a video clip, thereby increasing computational costs significantly. The proposed MMA addresses such challenges by constructing a spike-to-video model, directly producing an image sequence at a time. Specifically, we propose a U-shaped Multiscale Temporal Correlation Learning (MTCL) to fuse the features at different temporal resolutions for clear video reconstruction. At each scale, we introduce a Fine-Grained Attention (FGA) module for fine-spatial context modeling within a patch and a Mamba module for integrating features across patches. Adopting a lightweight U-shaped structure and fine-grained feature extraction at each level, our method reconstructs high-quality image sequences quickly. The experimental results show that the proposed MMA surpasses current state-of-the-art methods in image quality, computation cost, and model size.
Dilmurat Alim, Chen Yang 0034, Laiyun Qing, Guorong Li, Qingming Huang
IEEE Signal Process. Lett.4
2025 Boosting UAV Detection via Memory-Enhanced Attention and Contrastive Learning
abstract
With unmanned aerial vehicles (UAVs) having emerged in diverse application domains, visual detection of UAVs has become a critical research focus in recent years. However, most existing methods are limited in capturing small UAVs and may not perform well in complex backgrounds. To address these challenges, we propose a novel detection framework that integrates newly designed memory mechanism and contrastive loss to improve UAV detection. Specifically, we first utilize a clustering algorithm to gather representative UAV prototypes, which are then utilized to construct a reliable memory bank. Then, we design a UAV Memory-Enhanced Attention (UMEA) module to propagate high-confidence prototypes from the memory bank, thereby enhancing the appearance features of UAVs in the input frame. Furthermore, we introduce a Memory-Driven Contrastive Learning (MDCL) loss function to pull UAVs closer in the feature space while pushing them further away from the background. Extensive experiments conducted on three challenging datasets, NPS-Drones, ARD-MAV and Drone-vs-Bird demonstrate that the proposed method outperforms several state-of-the-art models in terms of the main metric AP with a large absolute margin, 2.1%, 3.6%, and 4.4%, respectively.
Yunchuan Ma, Yuankai Qi, Laiyun Qing, Guorong Li
IEEE Signal Process. Lett.5
2025 Iterative Bounded Distance Decoding With Random Flipping for Product-Like Codes
abstract
Product-like codes are widely used in high-speed communication systems since they can be decoded with low-complexity hard decision decoders (HDDs). To meet the growing demand of data rates, enhanced HDDs are required. In this paper, we propose a novel soft-aided HDD (SA-HDD), termed iterative bounded distance decoding with random flipping (iBDD-RF), for product-like codes. In iBDD-RF, the soft reliability of a bit is a weighted sum of the output of bounded distance decoder (BDD) and the channel log-likelihood ratio (LLR). When the amplitude of the soft reliability of a bit is less than a given threshold, it is flipped with a given probability. This random flipping may make the decoder escape from the local optimum. To optimize the threshold and the flipping probability, we derive the density evolution (DE) equations of iBDD-RF for product codes (PCs) and staircase codes (SCs). Our extensive numerical results show that iBDD-RF outperforms iBDD with scaled reliability (iBDD-SR) over the binary-input additive white Gaussian noise (Bi-AWGN) channel. Particularly, for a PC with (255,239,2) Bose-Chaudhuri-Hocquenghem (BCH) code and an SC with (254,230,3) BCH code, iBDD-RF performs about 0.25 dB and 0.28 dB better than iBDD-SR, respectively.
Guorong Li, Shiguo Wang, Shancheng Zhao
IEEE Trans. Commun.1
2025 QuadrantSearch: A Novel Method for Registering UAV and Backpack LiDAR Point Clouds in Forested Areas
abstract
Unmanned aerial vehicle (UAV) laser scanning (ULS) and backpack laser scanning (BLS) are two commonly employed technologies in precision forestry. However, data acquired by these two types of light detection and ranging (LiDAR) are distinct, with one capturing point clouds beneath the canopy and the other above. Consequently, there is minimal overlap in the point clouds collected by both methods, especially in dense forests, presenting significant challenges for data registration. Furthermore, many trees in forests (particularly broadleaf trees) have the tree tops and trunk centers not aligned vertically, which greatly increases the difficulty of the data registration methods based on tree position. To solve the above-mentioned problems, we here propose a novel and robust method to register ULS and BLS point clouds in forested areas. Our method consists of three key steps, that is, tree location extraction, quadrant search-based minimum spanning tree (MST) matching, and registration. The quadrant searching strategy dynamically searches for potential candidates in four quadrants centered on the initial tree locations. By constructing MSTs for the potential tree locations, triangle constraints require only four topologically similar tree locations to find one-to-one correspondences during the stepwise MST matching process. The proposed method was evaluated in five urban forest sample plots and one natural forest sample plot located in China, covering both coniferous and broadleaf forests. The results show that our method obtained good registration results on all six sample plots, with an averaged rotation error, translation error, pointwise error, and root-mean-square error (RMSE) of 0.012 rad, 0.354, 0.378, and 0.379 m, respectively. Comparative studies indicate that our method outperformed existing registration methods, demonstrating its effectiveness and robustness. Our method allows for the creation of a more complete picture of forest vertical structure and holds great potential for informing sustainable forest management practices and supporting critical ecological assessments.
Guorong Li, Bin Wu 0010, Zhan Pan, Linxin Dong, Guochun Shen, Tian Xiao, Lefeng Zhang, Bailang Yu
IEEE Trans. Geosci. Remote. Sens.1
2025 Boost Tracking by Natural Language With Prompt-Guided Grounding
abstract
TNL (Tracking by Natural Language) aims to locate the target described by a natural language sentence in a video. Most existing TNL methods are typically composed of three modules: object grounding, object tracking, and switching module, and their performance is limited by the poor performance of the grounding and switching modules due to the complex backgrounds and inaccurate information stored in the memory. This paper presents a global-local framework to address these issues, which includes a prompt-guided grounding module, a trained local tracking module, and a memory-based switcher module. The prompt-guided grounding module uses noun prompts to guide the CLIP model in focusing more on target regions and aligning visual features semantically with linguistic features, avoiding being misled by distractors and background. The memory-based switch module stores historical information with higher-quality memory, allowing the model to make more accurate decisions based on reliable data, thus improving the overall performance. Experiments on TNL2K, LaSOT, and OTB-Lang demonstrate the effectiveness and generalizability of the proposed framework.
Hengyou Li, Xinyan Liu 0008, Guorong Li, Shuhui Wang, Laiyun Qing, Qingming Huang
IEEE Trans. Intell. Transp. Syst.3
2025 Dynamic Erasing Network With Adaptive Temporal Modeling for Weakly Supervised Video Anomaly Detection
abstract
The weakly supervised video anomaly detection aims to learn a detection model using only video-level labeled data. Prior studies ignore the complexity or duration of anomalies present in abnormal videos during temporal modeling. Moreover, existing works usually detect the most abnormal segments, potentially overlooking the completeness of anomalies. We propose a dynamic erasing network (DE-Net) for weakly supervised video anomaly detection, which learns video-specific temporal features via adaptive temporal modeling (ATM) to address these limitations. Specifically, to handle duration variations of abnormal events, we propose an ATM module capable of adaptively selecting and aggregating the most appropriate K temporal scale features for each video. Then, we design a dynamic erasing (DE) strategy that dynamically assesses the completeness of the detected anomalies and erases prominent abnormal segments to encourage the model to discover gentle abnormal segments. The proposed method achieves favorable performance compared to several state-of-the-art approaches on the widely used XD-Violence, TAD, and UCF-Crime datasets.
Chen Zhang 0013, Guorong Li, Yuankai Qi, Hanhua Ye, Laiyun Qing, Ming-Hsuan Yang 0001, Qingming Huang
IEEE Trans. Neural Networks Learn. Syst.2
2024 Weak-Evidence Aggregation for the Choice of Plausible Alternatives Task
Zhipeng Xie, Guorong Li
ADMA (1)2
2024 Weakly Supervised Video Individual Counting
abstract
Video Individual Counting (VIC) aims to predict the number of unique individuals in a single video. Existing methods learn representations based on trajectory labels for individuals, which are annotation-expensive. To provide a more realistic reflection of the underlying practical challenge, we introduce a weakly supervised VIC task, wherein trajectory labels are not provided. Instead, two types of labels are provided to indicate traffic entering the field of view (inflow) and leaving the field view (outflow). We also propose the first solution as a baseline that formulates the task as a weakly supervised contrastive learning problem under group-level matching. In doing so, we devise an end-to-end trainable soft contrastive loss to drive the network to distin-guish inflow, outflow, and the remaining. To facilitate future study in this direction, we generate annotations from the existing VIC datasets Sense Crowd and CroHD and also build a new dataset, UAVVIC. Extensive results show that our baseline weakly supervised method outperforms supervised methods, and thus, little information is lost in the transition to the more practically relevant weakly supervised task. The code and trained model can be found at CGNet.
Xinyan Liu 0008, Guorong Li, Yuankai Qi, Ziheng Yan, Zhenjun Han, Anton van den Hengel, Ming-Hsuan Yang 0001, Qingming Huang
CVPR2
2024 Semantic-aware SAM for Point-Prompted Instance Segmentation
abstract
Single-point annotation in visual tasks, with the goal of minimizing labelling costs, is becoming increasingly prominent in research. Recently, visual foundation models, such as Segment Anything (SAM), have gained widespread usage due to their robust zero-shot capabilities and exceptional annotation performance. However, SAM's class-agnostic output and high confidence in local segmentation introduce semantic ambiguity, posing a challenge for precise category-specific segmentation. In this paper, we introduce a cost-effective category-specific segmenter using SAM. To tackle this challenge, we have devised a Semantic-Aware Instance Segmentation Network (SAPNet) that integrates Multiple Instance Learning (MIL) with matching capability and SAM with point prompts. SAPNet strategically selects the most representative mask proposals generated by SAM to supervise segmentation, with a specific focus on object category information. Moreover, we introduce the Point Distance Guidance and Box Mining Strategy to mitigate inherent challenges: group and local issues in weakly supervised segmentation. These strategies serve to further enhance the overall segmentation performance. The experimental results on Pascal VOC and COCO demonstrate the promising performance of our proposed SAPNet, emphasizing its semantic matching capabilities and its potential to advance point-prompted instance segmentation. The code is available at https://github.com/zhaoyangwei123/SAPNet.
Zhaoyang Wei, Pengfei Chen 0004, Xuehui Yu, Guorong Li, Jianbin Jiao, Zhenjun Han
CVPR4
2024 Directly Locating Actions in Video with Single Frame Annotation
abstract
We propose a novel method for point-supervised action localization.Differs from the common practice of locating actions by first categorizing each video frame, our method directly predicts actions' positions and length. Specifically, point-supervised action localization is achieved by a series of fully supervised action location iteratively. In each iteration, the input video are used as input tokens and fed into a transformer, where the encoder extracts global context of the clips, and the decoder generates queries containing information for action localization. Three MLP heads are built on each query to obtain the probability, the center, and the length of each action instance respectively. Experiments on three popular datasets prove the potential of our method.
Haoran Tong, Xinyan Liu 0008, Guorong Li, Laiyun Qing
ICMR3
2024 Style-aware two-stage learning framework for video captioning
abstract
Significant progress has been made in video captioning in recent years. However, most existing methods directly learn from all given captions without distinguishing the styles of captions. The large diversity in these captions might bring ambiguity to the model learning. To address this issue, we propose a style-aware two-stage learning framework. In the first stage, the model is trained with captions of separate styles, including length style (short, medium, long), action style (single action or multiple actions), and object style (one object or more). For efficiency, a shared model with multiple individual style vectors is learned. In the second stage, a video style encoder is devised to capture style information from the input video, and it outputs a guidance signal of how to utilize the style vectors for the final caption generation. Without whistles and bells, our method achieves state-of-the-art performance on three widely-used public datasets, MSVD, MSR-VTT and VATEX. The source code and trained models will be made available to the public.
Yunchuan Ma, Yuankai Qi, Amin Beheshti, Laiyun Qing, Guorong Li
Knowl. Based Syst.7
2024 Learning Hierarchical Modular Networks for Video Captioning
abstract
Video captioning aims to generate natural language descriptions for a given video clip. Existing methods mainly focus on end-to-end representation learning via word-by-word comparison between predicted captions and ground-truth texts. Although significant progress has been made, such supervised approaches neglect semantic alignment between visual and linguistic entities, which may negatively affect the generated captions. In this work, we propose a hierarchical modular network to bridge video representations and linguistic semantics at four granularities before generating captions: entity, verb, predicate, and sentence. Each level is implemented by one module to embed corresponding semantics into video representations. Additionally, we present a reinforcement learning module based on the scene graph of captions to better measure sentence similarity. Extensive experimental results show that the proposed method performs favorably against the state-of-the-art models on three widely-used benchmark datasets, including microsoft research video description corpus (MSVD), MSR-video to text (MSR-VTT), and video-and-TEXt (VATEX).
Guorong Li, Hanhua Ye, Yuankai Qi, Shuhui Wang, Laiyun Qing, Qingming Huang, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 CPR++: Object Localization via Single Coarse Point Supervision
abstract
Point-based object localization (POL), which pursues high-performance object sensing under low-cost data annotation, has attracted increased attention. However, the point annotation mode inevitably introduces semantic variance due to the inconsistency of annotated points. Existing POL heavily rely on strict annotation rules, which are difficult to define and apply, to handle the problem. In this study, we propose coarse point refinement (CPR), which to our best knowledge is the first attempt to alleviate semantic variance from an algorithmic perspective. CPR reduces the semantic variance by selecting a semantic centre point in a neighbourhood region to replace the initial annotated point. Furthermore, We design a sampling region estimation module to dynamically compute a sampling region for each object and use a cascaded structure to achieve end-to-end optimization. We further integrate a variance regularization into the structure to concentrate the predicted scores, yielding CPR++. We observe that CPR++ can obtain scale information and further reduce the semantic variance in a global region, thus guaranteeing high-performance object localization. Extensive experiments on four challenging datasets validate the effectiveness of both CPR and CPR++. We hope our work can inspire more research on designing algorithms rather than annotation rules to address the semantic variance problem in POL.
Xuehui Yu, Pengfei Chen 0004, Kuiran Wang, Xumeng Han, Guorong Li, Zhenjun Han, Qixiang Ye, Jianbin Jiao
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Rethink video retrieval representation for video captioning
Mingkai Tian, Guorong Li, Yuankai Qi, Shuhui Wang, Quan Z. Sheng, Qingming Huang
Pattern Recognit.2
2024 Progressive Multi-Resolution Loss for Crowd Counting
abstract
Crowd counting is usually handled in a density map regression fashion, which is supervised via an L2 loss between the predicted density map and ground truth. To effectively regulate models, various improved L2 loss functions have been developed to find a better correspondence between predicted density and annotation positions. In this paper, we propose to predict the density map at one resolution but measure its quality via a derived log-formed loss at multiple resolutions. Unlike existing methods that assume density maps at different resolutions are independent, our loss is obtained by modeling the likelihood function inspired by the relationship of density maps across multi-resolutions. We find that the traditional single-resolution L2 loss is a particular case of our derived log-likelihood. We mathematically prove it is superior to a single-resolution L2 loss. Without bells and whistles, the proposed loss substantially improves several baselines and performs favorably compared to state-of-the-art methods on five crowd counting datasets: NWPU-Crowd, ShanghaiTech A & B, UCF-QNRF, and JHU-Crowd++. The source code and trained models are released athttps://github.com/streamer-AP/PML_Loss.git.
Ziheng Yan, Yuankai Qi, Guorong Li, Xinyan Liu 0008, Weigang Zhang, Ming-Hsuan Yang 0001, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.3
2024 SpikeODE: Image Reconstruction for Spike Camera With Neural Ordinary Differential Equation
abstract
The recently invented retina-inspired spike camera has shown great potential for capturing dynamic scenes. However, reconstructing high-quality images from the binary spike data remains a challenge due to the existence of noises in the camera. This paper proposes SpikeODE, a novel approach to reconstructing clear images by exploring temporal-spatial correlation to depress noises. The main idea of our method is to restore the continuous dynamic process of real scenes in a latent space and learn the temporal correlations in a fine-grained manner. Furthermore, to model the dynamic process more effectively, we design a conditional ODE where the latent state of each timestamp is conditioned on the observed spike data. Subsequently, forward and backward inferences are conducted through the ODE to investigate the correlations between the representation of the target timestamp and the information from both past and future contexts. Additionally, we incorporate a Unet structure with a pixel-wise attention mechanism at each level to learn spatial correlations. Experimental results demonstrate that our method outperforms state-of-the-art methods across several metrics.
Chen Yang 0034, Guorong Li, Shuhui Wang, Li Su 0003, Laiyun Qing, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.2
2024 Toward Unified Token Learning for Vision-Language Tracking
abstract
In this paper, we present a simple, flexible and effective vision-language (VL) tracking pipeline, termed MMTrack, which casts VL tracking as a token generation task. Traditional paradigms address VL tracking task indirectly with sophisticated prior designs, making them over-specialize on the features of specific architectures or mechanisms. In contrast, our proposed framework serializes language description and bounding box into a sequence of discrete tokens. In this new design paradigm, all token queries are required to perceive the desired target and directly predict spatial coordinates of the target in an auto-regressive manner. The design without other prior modules avoids multiple sub-tasks learning and hand-designed loss functions, significantly reducing the complexity of VL tracking modeling and allowing our tracker to use a simple cross-entropy loss as unified optimization objective for VL tracking task. Extensive experiments on TNL2K, LaSOT, LaSOT$_{\mathrm{ext}}$and OTB99-Lang benchmarks show that our approach achieves promising results, compared to other state-of-the-arts.
Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Guorong Li, Rongrong Ji, Xianxian Li
IEEE Trans. Circuits Syst. Video Technol.4
2024 A Novel Intercalibration Method for Fengyun(FY)-3 VIRR Using MERSI Onboard the Same Satellite Based on Pseudo-Invariant Pixels
abstract
This study presents a novel approach to the radiometric inter-calibration between two sensors onboard the same satellite based on pseudo-invariant pixels (PIPs) using iteratively re-weighted multivariate alteration detection (IR-MAD) method. The IR-MAD algorithm can statistically select pseudo-invariant pixels from the multispectral image pair to assess the radiometric differences between them. Analysis of multiple image pairs from different acquisition times can provide long-term inter-calibration results of the two sensors. The procedure is applied to Fengyun(FY)-3A&3B Visible Infrared Radiometer (VIRR), with the Medium Resolution Spectral Imager (MERSI) onboard the same platform as the reference. Consistency of the spatial distribution of the PIPs selected by IR-MAD with pseudo-invariant calibration sites (PICS) given by other scientists demonstrates the effectiveness of our method. The long-term time series trending of top-of-atmosphere VIRR reflectance over LIBYA1 and LIBYA4 after inter-calibration correction shows that the inter-calibrated VIRR has good agreement with MERSI, with a mean bias of less than 1% and an uncertainty of less than 2% for most channels. The approach requires no prior knowledge of the inter-calibration targets and extends PICS to the pixel-level targets, which results in more diverse samples, broader dynamic ranges and lower uncertainty, yielding consistent and reliable long-term inter-calibration results.
Xiuqing Hu, Kun Gao 0001, Guorong Li, Na Xu 0001, Peng Zhang 0024
IEEE Trans. Geosci. Remote. Sens.6
2024 Self Supervised Progressive Network for High Performance Video Object Segmentation
abstract
Recently, self-supervised video object segmentation (VOS) has attracted much interest. However, most proxy tasks are proposed to train only a single backbone, which relies on a point-to-point correspondence strategy to propagate masks through a video sequence. Due to its simple pipeline, the performance of the single backbone paradigm is still unsatisfactory. Instead of following the previous literature, we propose our self-supervised progressive network (SSPNet) which consists of a memory retrieval module (MRM) and collaborative refinement module (CRM). The MRM can perform point-to-point correspondence and produce a propagated coarse mask for a query frame through self-supervised pixel-level and frame-level similarity learning. The CRM, which is trained via cycle consistency region tracking, aggregates the reference & query information and learns the collaborative relationship among them implicitly to refine the coarse mask. Furthermore, to learn semantic knowledge from unlabeled data, we also design two novel mask-generation strategies to provide the training data with meaningful semantic information for the CRM. Extensive experiments conducted on DAVIS-17, YouTube- VOS and SegTrack v2 demonstrate that our method surpasses the state-of-the-art self-supervised methods and narrows the gap with the fully supervised methods.
Guorong Li, Dexiang Hong, Kai Xu 0013, Bineng Zhong 0001, Li Su 0003, Zhenjun Han, Qingming Huang
IEEE Trans. Neural Networks Learn. Syst.1
2023 Exploiting Completeness and Uncertainty of Pseudo Labels for Weakly Supervised Video Anomaly Detection
abstract
Weakly supervised video anomaly detection aims to identify abnormal events in videos using only video-level labels. Recently, two-stage self-training methods have achieved significant improvements by self-generating pseudo labels and self-refining anomaly scores with these labels. As the pseudo labels play a crucial role, we propose an enhancement framework by exploiting completeness and uncertainty properties for effective self-training. Specifically, we first design a multi-head classification module (each head serves as a classifier) with a diversity loss to maximize the distribution differences of predicted pseudo labels across heads. This encourages the generated pseudo labels to cover as many abnormal events as possible. We then devise an iterative uncertainty pseudo label refinement strategy, which improves not only the initial pseudo labels but also the updated ones obtained by the desired classifier in the second stage. Extensive experimental results demonstrate the proposed method performs favorably against state-of-the-art approaches on the UCF-Crime, TAD, and XD-Violence benchmark datasets.
Chen Zhang 0013, Guorong Li, Yuankai Qi, Shuhui Wang, Laiyun Qing, Qingming Huang, Ming-Hsuan Yang 0001
CVPR2
2023 Spatial Self-Distillation for Object Detection with Inaccurate Bounding Boxes
abstract
Object detection via inaccurate bounding boxes supervision has boosted a broad interest due to the expensive high-quality annotation data or the occasional inevitability of low annotation quality (e.g. tiny objects). The previous works usually utilize multiple instance learning (MIL), which highly depends on category information, to select and refine a low-quality box. Those methods suffer from object drift, group prediction and part domination problems without exploring spatial information. In this paper, we heuristically propose a Spatial Self-Distillation based Object Detector (SSD-Det) to mine spatial information to refine the inaccurate box in a self-distillation fashion. SSD-Det utilizes a Spatial Position Self-Distillation (SPSD) module to exploit spatial information and an interactive structure to combine spatial information and category information, thus constructing a high-quality proposal bag. To further improve the selection procedure, a Spatial Identity Self-Distillation (SISD) module is introduced in SSD-Det to obtain spatial confidence to help select the best proposals. Experiments on MS-COCO and VOC datasets with noisy box annotation verify our method’s effectiveness and achieve state-of-the-art performance. The code is available at https://github.com/ucas-vg/PointTinyBenchmark/tree/SSD-Det.
Pengfei Chen 0004, Xuehui Yu, Guorong Li, Zhenjun Han, Jianbin Jiao
ICCV4
2023 SiamBAN: Target-Aware Tracking With Siamese Box Adaptive Network
abstract
Variation of scales or aspect ratios has been one of the main challenges for tracking. To overcome this challenge, most existing methods adopt either multi-scale search or anchor-based schemes, which use a predefined search space in a handcrafted way and therefore limit their performance in complicated scenes. To address this problem, recent anchor-free based trackers have been proposed without using prior scale or anchor information. However, an inconsistency problem between classification and regression degrades the tracking performance. To address the above issues, we propose a simple yet effective tracker (named Siamese Box Adaptive Network, SiamBAN) to learn a target-aware scale handling schema in a data-driven manner. Our basic idea is to predict the target boxes in a per-pixel fashion through a fully convolutional network, which is anchor-free. Specifically, SiamBAN divides the tracking problem into classification and regression tasks, which directly predict objectiveness and regress bounding boxes, respectively. A no-prior box design is proposed to avoid tuning hyper-parameters related to candidate boxes, which makes SiamBAN more flexible. SiamBAN further uses a target-aware branch to address the inconsistency problem. Experiments on benchmarks including VOT2018, VOT2019, OTB100, UAV123, LaSOT and TrackingNet show that SiamBAN achieves promising performance and runs at 35 FPS.
Zedu Chen, Bineng Zhong 0001, Guorong Li, Shengping Zhang, Rongrong Ji, Zhenjun Tang, Xianxian Li
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 A Tale of HodgeRank and Spectral Method: Target Attack Against Rank Aggregation is the Fixed Point of Adversarial Game
abstract
Rank aggregation with pairwise comparisons has shown promising results in elections, sports competitions, recommendations, and information retrieval. However, little attention has been paid to the security issue of such algorithms, in contrast to numerous research work on the computational and statistical characteristics. Driven by huge profit, the potential adversary has strong motivation and incentives to manipulate the ranking list. Meanwhile, the intrinsic vulnerability of the rank aggregation methods is not well studied in the literature. To fully understand the possible risks, we focus on the purposeful adversary who desires to designate the aggregated results by modifying the pairwise data in this paper. From the perspective of the dynamical system, the attack behavior with a target ranking list is a fixed point belonging to the composition of the adversary and the victim. To perform the targeted attack, we formulate the interaction between the adversary and the victim as a game-theoretic framework consisting of two continuous operators while Nash equilibrium is established. Then two procedures against HodgeRank and RankCentrality are constructed to produce the modification of the original data. Furthermore, we prove that the victims will produce the target ranking list once the adversary masters the complete information. It is noteworthy that the proposed methods allow the adversary only to hold incomplete information or imperfect feedback and perform the purposeful attack. The effectiveness of the suggested target attack strategies is demonstrated by a series of toy simulations and several real-world data experiments. These experimental results show that the proposed methods could achieve the attacker's goal in the sense that the leading candidate of the perturbed ranking list is the designated one by the adversary.
Ke Ma 0001, Qianqian Xu 0001, Jinshan Zeng, Guorong Li, Xiaochun Cao, Qingming Huang
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Multi-Modal Multi-Grained Embedding Learning for Generalized Zero-Shot Video Classification
abstract
Zero-shot learning aims to learn knowledge from existing information to classify new classes with no visual training data. In the current work on zero-shot video classification, only the category name information can be used for unseen classes. While, most of the category names cannot fully describe the entire video information, but are only precise labels assigned by humans to actions, in which the amount of information is very small. In order to make up for the semantic deficiencies of video databases and build relationships between categories, we propose a multi-modal generalized zero-shot video classification framework based on multi-grained semantic information with a proposed video description text database. Our model explores semantic knowledge from accurate but lacking informative category names and exhaustive but redundant description texts, and learns visual knowledge from semantic embeddings of varying granularity. Further, we use the learned semantic and visual knowledge to perform multi-grained classification on test video data with both seen and unseen classes. To describe actions in detail and provide complete semantic information, we propose a description text database. The textual descriptions, including category definitions and explanations, in our proposed textual database effectively help establish relationships between categories, thus providing a more reliable basis for visual feature synthesis. Furthermore, our framework generates synthesized features for unseen classes from both coarse-grained and fine-grained semantic information, which would effectively avoid the bias of generalized zero-shot learning on seen classes. Extensive experimental results on the database prove the validity of our method and the effectiveness of the description texts in generalized zero-shot video classification problems.
Mingyao Hong, Xinfeng Zhang 0001, Guorong Li, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.3
2023 Robust Tracking via Uncertainty-Aware Semantic Consistency
abstract
Robust tracking has a variety of practical applications. Despite many years of progress, it is still a difficult problem due to enormous uncertainties in real-world scenes. To address this issue, we propose a robust anchor-free based tracking model with uncertainty estimation. Within the model, a new data-driven uncertainty estimation strategy is proposed to generate uncertainty-aware features with promising discriminative and descriptive power. Then, a simple yet effective pyramid-wise cross correlation operation is constructed to extract multi-scale semantic features that provide rich correlation information for uncertainty-aware estimation and thus enhances the tracking robustness. Finally, a semantic consistency checking branch is designed to further estimate uncertainty of output results from the classification and regression branches by adaptively generating semantically consistent labels. Experiments on six benchmarks (i.e., OTB100, VOT2018, VOT2020, TrackingNet, GOT-10K and LaSOT) show the competing performance of our tracker with 130 FPS.
Jie Ma 0006, Xiangyuan Lan, Bineng Zhong 0001, Guorong Li, Zhenjun Tang, Xianxian Li, Rongrong Ji
IEEE Trans. Circuits Syst. Video Technol.4
2023 Rethinking Sampling Strategies for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (re-ID) remains a challenging task. While extensive research has focused on the framework design and loss function, this paper shows that sampling strategy plays an equally important role. We analyze the reasons for the performance differences between various sampling strategies under the same framework and loss function. We suggest that deteriorated over-fitting is an important factor causing poor performance, and enhancing statistical stability can rectify this problem. Inspired by that, a simple yet effective approach is proposed, termed group sampling, which gathers samples from the same class into groups. The model is thereby trained using normalized group samples, which helps alleviate the negative impact of individual samples. Group sampling updates the pipeline of pseudo-label generation by guaranteeing that samples are more efficiently classified into the correct classes. It regulates the representation learning process, enhancing statistical stability for feature representation in a progressive fashion. Extensive experiments on Market-1501, DukeMTMC-reID and MSMT17 show that group sampling achieves performance comparable to state-of-the-art methods and outperforms the current techniques under purely camera-agnostic settings. Code has been available at https://github.com/ucas-vg/GroupSampling.
Xumeng Han, Xuehui Yu, Guorong Li, Jian Zhao 0006, Gang Pan 0002, Qixiang Ye, Jianbin Jiao, Zhenjun Han
IEEE Trans. Image Process.3
2023 Fine-Grained Feature Generation for Generalized Zero-Shot Video Classification
abstract
Generalized zero-shot video classification aims to train a classifier to classify videos including both seen and unseen classes. Since the unseen videos have no visual information during training, most existing methods rely on the generative adversarial networks to synthesize visual features for unseen classes through the class embedding of category names. However, most category names only describe the content of the video, ignoring other relational information. As a rich information carrier, videos include actions, performers, environments, etc., and the semantic description of the videos also express the events from different levels of actions. In order to use fully explore the video information, we propose a fine-grained feature generation model based on video category name and its corresponding description texts for generalized zero-shot video classification. To obtain comprehensive information, we first extract content information from coarse-grained semantic information (category names) and motion information from fine-grained semantic information (description texts) as the base for feature synthesis. Then, we subdivide motion into hierarchical constraints on the fine-grained correlation between event and action from the feature level. In addition, we propose a loss that can avoid the imbalance of positive and negative examples to constrain the consistency of features at each level. In order to prove the validity of our proposed framework, we perform extensive quantitative and qualitative evaluations on two challenging datasets: UCF101 and HMDB51, and obtain a positive gain for the task of generalized zero-shot video classification.
Mingyao Hong, Xinfeng Zhang 0001, Guorong Li, Qingming Huang
IEEE Trans. Image Process.3
2023 Anti-UAV: A Large-Scale Benchmark for Vision-Based UAV Tracking
abstract
Unmanned Aerial Vehicles (UAV) have many applications in both commerce and recreation. However, irresponsibly operated UAVs will pose a threat to public safety. Therefore, developing our understanding of UAVs and their uses is of particular interest. This paper considers tracking UAVs, which provide multifaceted information around location, paths and trajectories. To facilitate research on this topic, we introduce a new benchmark, herein referred to as Anti-UAV, which provides a novel direction for UAV tracking with more than 300 video pairs containing over 580 k manually annotated bounding boxes. Addressing anti-UAV research challenges could help to design anti-UAV systems, which in turn may improve surveillance. Accordingly, we have proposed a simple yet effective approach, called dual-flow semantic consistency (DFSC) is proposed for UAV tracking. Modulated by the semantic flow across video sequences, tracker learns more robust class-level semantic information and obtains more discriminative instance-level features. Experiments highlight significant performance gain with the proposed approach over state-of-the-art trackers and the challenging aspects of Anti-UAV. The Anti-UAV benchmark and the code for the proposed approach have been made publicly available athttps://github.com/ucas-vg/Anti-UAVandhttps://github.com/ZhaoJ9014/Anti-UAV.
Kuiran Wang, Xiaoke Peng, Xuehui Yu, Qiang Wang 0051, Junliang Xing, Guorong Li, Guodong Guo, Qixiang Ye, Jianbin Jiao, Jian Zhao 0006, Zhenjun Han
IEEE Trans. Multim.7
2023 Weakly Supervised Text-based Actor-Action Video Segmentation by Clip-level Multi-instance Learning
abstract
In real-world scenarios, it is common that a video contains multiple actors and their activities. Selectively localizing one specific actor and its action spatially and temporally via a language query becomes a vital and challenging task. Existing fully supervised methods require extensive elaborately annotated data and are sensitive to the class labels, which cannot satisfy real-world applications’ needs. Thus, we introduce the task of weakly supervised actor-action video segmentation from a sentence query (AAVSS) in this work, where only the video-sentence pairs are provided. To the best of our knowledge, our work is the first to perform AAVSS under weakly supervised situations. However, this task is extremely challenging not only because the task aims to learn the complex interactions between two heterogeneous modalities but also because the task needs to learn fine-grained analysis of video content without pixel-level annotations. To overcome the challenges, we propose a two-stage network. The network first follows the sentence guidance to localize the candidate region and then performs segmentation to achieve selective segmentation. Specifically, a novel tracker-based clip-level multiple instance learning paradigm is proposed in this article to learn the matches between regions and sentences, which makes our two-stage network robust to the region proposal network. Furthermore, two intrinsic characteristics of the video, temporal consistency and motion information, are utilized in companion with the weak supervision to facilitate the region-query matching. Through extensive experiments, the proposed method achieves comparable performance to state-of-the-art fully supervised approaches on two large-scale benchmarks, including A2D Sentences and J-HMDB Sentences.
Weidong Chen 0013, Guorong Li, Xinfeng Zhang 0001, Shuhui Wang, Liang Li 0003, Qingming Huang
ACM Trans. Multim. Comput. Commun. Appl.2
2022 Hierarchical Modular Network for Video Captioning
abstract
Video captioning aims to generate natural language descriptions according to the content, where representation learning plays a crucial role. Existing methods are mainly developed within the supervised learning framework via word-by-word comparison of the generated caption against the ground-truth text without fully exploiting linguistic semantics. In this work, we propose a hierarchical modular network to bridge video representations and linguistic semantics from three levels before generating captions. In particular, the hierarchy is composed of: (I) Entity level, which highlights objects that are most likely to be mentioned in captions. (II) Predicate level, which learns the actions conditioned on highlighted objects and is supervised by the predicate in captions. (III) Sentence level, which learns the global semantic representation and is supervised by the whole caption. Each level is implemented by one module. Extensive experimental results show that the proposed method performs favorably against the state-of-the-art models on the two widely-used benchmarks: MSVD 104.0% and MSR-VTT 51.5% in CIDEr score. Code will be made available at https://github.com/MarcusNerva/HMN.
Hanhua Ye, Guorong Li, Yuankai Qi, Shuhui Wang, Qingming Huang, Ming-Hsuan Yang 0001
CVPR2
2022 Object Localization under Single Coarse Point Supervision
abstract
Point-based object localization (POL), which pursues high-performance object sensing under low-cost data annotation, has attracted increased attention. However, the point annotation mode inevitably introduces semantic variance for the inconsistency of annotated points. Existing POL methods heavily reply on accurate keypoint annotations which are difficult to define. In this study, we propose a POL method using coarse point annotations, relaxing the supervision signals from accurate key points to freely spotted points. To this end, we propose a coarse point refinement (CPR) approach, which to our best knowledge is the first attempt to alleviate semantic variance from the perspective of algorithm. CPR constructs point bags, selects semantic-correlated points, and produces semantic center points through multiple instance learning (MIL). In this way, CPR defines a weakly supervised evolution procedure, which ensures training high-performance object localizer under coarse point supervision. Experimental results on COCO, DOTA and our proposed SeaPerson dataset validate the effectiveness of the CPR approach. The dataset and code will be available at https://github.com/ucas-vg/PointTinyBenchmark/
Xuehui Yu, Pengfei Chen 0004, Najmul Hassan, Guorong Li, Junchi Yan, Humphrey Shi, Qixiang Ye, Zhenjun Han
CVPR5
2022 Enhanced Semantic Head for Cascade Instance Segmentation
abstract
Recently, cascade instance segmentation inspired by cascade object detection has achieved notable performance. Due to the lack of global information, many methods suffer from incomplete segmentation such as missing edge regions and discontinuities within instances. To solve this problem, we proposed an effective and flexible semantic head to extract enhanced spatial context information. A vision transformer is utilized to generate global context features, and a convolution network is adopted to generate spatial context features. After combining the two modules, we obtain enhanced semantic segmentation features for segmentation. Extensive experiments show that the enhanced semantic head achieves 40.6% and 42.3% mask AP for cascade predictor HTC and DSC, which surpass about 0.9 and 1.4 percentage points respectively. The enhanced semantic head is universal and effective to improve the performance of different cascade predictors.
Xuerong Huang, Li Su 0003, Guorong Li, Xinfeng Zhang 0001, Laiyun Qing, Qingming Huang
ICME3
2022 Multi-Attention Network for Compressed Video Referring Object Segmentation
abstract
Referring video object segmentation aims to segment the object referred by a given language expression. Existing works typically require compressed video bitstream to be decoded to RGB frames before being segmented, which increases computation and storage requirements and ultimately slows the inference down. This may hamper its application in real-world computing resource limited scenarios, such as autonomous cars and drones. To alleviate this problem, in this paper, we explore the referring object segmenta- tion task on compressed videos, namely on the original video data flow. Besides the inherent difficulty of the video referring object segmentation task itself, obtaining discriminative representation from compressed video is also rather challenging. To address this problem, we propose a multi-attention network which consists of dual-path dual-attention module and a query-based cross-modal Transformer module. Specifically, the dual-path dual-attention module is designed to extract effective representation from compressed data in three modalities, i.e., I-frame, Motion Vector and Residual. The query-based cross-modal Transformer firstly models the corre- lation between linguistic and visual modalities, and then the fused multi-modality features are used to guide object queries to generate a content-aware dynamic kernel and to predict final segmentation masks. Different from previous works, we propose to learn just one kernel, which thus removes the complicated post mask-matching procedure of existing methods. Extensive promising experimental results on three challenging datasets show the effectiveness of our method compared against several state-of-the-art methods which are proposed for processing RGB data. Source code is available at: https://github.com/DexiangHong/MANet.
Weidong Chen 0013, Dexiang Hong, Yuankai Qi, Zhenjun Han, Shuhui Wang, Laiyun Qing, Qingming Huang, Guorong Li
ACM Multimedia8
2022 Weakly Supervised Anomaly Detection in Videos Considering the Openness of Events
abstract
Although various weakly supervised anomaly detection methods have been proposed in recent years, generalization of anomaly detection is still not well-explored. Existing weakly supervised methods usually use normal and abnormal events to pose anomaly detection as a regression problem. However, defining concepts that encompass all possible normal and abnormal event patterns is nearly unrealistic, so the anomaly detection model is likely to face both open normal and abnormal events in practical applications. We find some weakly supervised anomaly detection methods suffer from performance degradation when faced with open events due to their poor generalization. To tackle this issue, we propose a two-branch weakly supervised approach, which can improve the anomaly detection performance of open events without affecting the performance of the seen events. Specifically, considering that the pattern of open events is different from that of seen events, we design a Test Data Analyzer (TDA) that determines whether the test video features belong to seen or open data and argue for separate treatment for them. For the seen data, a classifier trained by multiple instance learning is used to predict anomaly scores. For the open data, we design an anomaly detection model via meta-learning named Meta-Learning Anomaly Detection (MLAD), which can directly determine whether open data is abnormal without updating model parameters. In detail, MLAD synthesizes pseudo-seen data and pseudo-open data so that the model can learn to detect anomalies in open data by transferring the knowledge of seen data. Experimental results validate the effectiveness of our proposed method.
Chen Zhang 0013, Guorong Li, Qianqian Xu 0001, Xinfeng Zhang 0001, Li Su 0003, Qingming Huang
IEEE Trans. Intell. Transp. Syst.2
2022 Introduction to the Special Issue on Fine-Grained Visual Recognition and Re-Identification
abstract
introduction Share on Introduction to the Special Issue on Fine-Grained Visual Recognition and Re-Identification Authors: Shiliang Zhang Peking University Peking UniversityView Profile , Guorong Li University of Chinese Academy of Sciences University of Chinese Academy of SciencesView Profile , Weigang Zhang Harbin Institute of Technology Harbin Institute of TechnologyView Profile , Qingming Huang University of Chinese Academy of Sciences University of Chinese Academy of SciencesView Profile , Tiejun Huang Peking University Peking UniversityView Profile , Mubarak Shah University of Central Florida University of Central FloridaView Profile , Nicu Sebe University of Trento University of TrentoView Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 18Issue 1sFebruary 2022 Article No.: 24pp 1–3https://doi.org/10.1145/3505280Online:25 January 2022Publication History 0citation169DownloadsMetricsTotal Citations0Total Downloads169Last 12 Months169Last 6 weeks22 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Shiliang Zhang, Guorong Li, Weigang Zhang, Qingming Huang, Tiejun Huang 0001, Mubarak Shah, Nicu Sebe
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Learning To Filter: Siamese Relation Network for Robust Tracking
abstract
Despite the great success of Siamese-based trackers, their performance under complicated scenarios is still not satisfying, especially when there are distractors. To this end, we propose a novel Siamese relation network, which introduces two efficient modules, i.e. Relation Detector (RD) and Refinement Module (RM). RD performs in a meta-learning way to obtain a learning ability to filter the distractors from the background while RM aims to effectively integrate the proposed RD into the Siamese framework to generate accurate tracking result. Moreover, to further improve the discriminability and robustness of the tracker, we introduce a contrastive training strategy that attempts not only to learn matching the same target but also to learn how to distinguish the different objects. Therefore, our tracker can achieve accurate tracking results when facing background clutters, fast motion, and occlusion. Experimental results on five popular benchmarks, including VOT2018, VOT2019, OTB100, LaSOT, and UAV123, show that the proposed method is effective and can achieve state-of-the-art results. The code will be available at https://github.com/hqucv/siamrn
Siyuan Cheng 0003, Bineng Zhong 0001, Guorong Li, Xin Liu 0011, Zhenjun Tang, Xianxian Li, Jing Wang 0049
CVPR3
2021 Exploiting sample correlation for crowd counting with multi-expert network
abstract
Crowd counting is a difficult task because of the diversity of scenes. Most of the existing crowd counting methods adopt complex structures with massive backbones to enhance the generalization ability. Unfortunately, the performance of existing methods on large-scale data sets is not satisfactory. In order to handle various scenarios with less complex network, we explored how to efficiently use the multi-expert model for crowd counting tasks. We mainly focus on how to train more efficient expert networks and how to choose the most suitable expert. Specifically, we propose a task-driven similarity metric based on sample’s mutual enhancement, referred as co-fine-tune similarity, which can find a more efficient subset of data for training the expert network. Similar samples are considered as a cluster which is used to obtain parameters of an expert. Besides, to make better use of the proposed method, we design a simple network called FPN with Deconvolution Counting Network, which is a more suitable base model for the multi-expert counting network. Experimental results show that multiple experts FDC (MFDC) achieves the best performance on four public data sets, including the large scale NWPU-Crowd data set. Furthermore, the MFDC trained on an extensive dense crowd data set can generalize well on the other data sets without extra training or fine-tuning.1
Xinyan Liu 0008, Guorong Li, Zhenjun Han, Weigang Zhang, Qingming Huang, Nicu Sebe
ICCV2
2021 DBAM: Dense Boundary and Actionness Map for Action Localization in Videos via Sentence Query
Weigang Zhang, Yushu Liu, Jianping Zhong, Guorong Li, Qingming Huang
ICIG (3)4
2021 Action Category and Phase Consistency Regularization for High-Quality Temporal Action Proposal Generation
abstract
Temporal action detection is a fundamental yet challenging task in video content analysis. The performance of existing methods still remains far from satisfactory as the mAP reduces dramatically at high tIoU threshold. With the goal of predicting the starting and ending points more precisely, this work first introduces the action category label into the temporal proposal generation stage of the training process. Specifically, with the category information, we proposed two extra constrains, i.e, action based constraint and action-class agnostic constraints. The former aims at minimizing the discrepancy inside the same action category while the latter forces the feature of the samples aggregates in the same phase. Comprehensive experiments are conducted on the THUMOS’14 benchmark. A remarkable improvement of average recall is attained especially when the number of proposals is small. And our approach achieves 29.0% mAP at a strict [email protected].
Yushu Liu, Weigang Zhang, Guorong Li, Qingming Huang
ICME3
2021 Cascade Cross-modal Attention Network for Video Actor and Action Segmentation from a Sentence
abstract
In this paper, we address the problem that selectively segments the actor and its action in the video clip given the sentence description. The main challenge is to match the local semantic features of the video with the heterogeneous textual features. A widely used language processing method in previous works is to leverage bi-LSTM and self-attention, which fixed the attention of the sentence and neglected the personality of the video, leading the attention of the sentence mismatch the most discriminative feature of the video. The proposed algorithm in this paper allows the sentence to learn the most discriminative features of the video, remarkably improving the accuracy of matching and segmentation. Specifically, we propose a cascade cross-modal attention to leverage two perspectives visual features to attend language from coarse to fine to generate the discriminative vision-aware language features. Moreover, equipping our framework with a contrastive learning method and a designed hard negative mining strategy benefits our proposed network from identifying the positive sample from numbers of negatives, and further improving the performance. To demonstrate the effectiveness of our approach, we conduct experiments on two datasets: A2D Sentences and J-HMDB Sentences. Experimental results show that our method significantly improves the performance over recent state-of-the-art methods.
Weidong Chen 0013, Guorong Li, Xinfeng Zhang 0001, Hongyang Yu 0001, Shuhui Wang, Qingming Huang
ACM Multimedia2
2021 Learning Self-Supervised Space-Time CNN for Fast Video Style Transfer
abstract
Style transfer on images has achieved significant advances in recent years, with the deep convolutional neural network (CNN). Directly applying image style transfer algorithms to each frame of a video independently often leads to flickering and unstable results. In this work, we present a self-supervised space-time convolutional neural network (CNN) based method for online video style transfer, named as VTNet, which is end-to-end trained from nearly unlimited unlabeled video data to produce temporally coherent stylized videos in real-time. Specifically, our VTNet transfer the style of a reference image to the source video frames, which is formed by the temporal prediction branch and the stylizing branch. The temporal prediction branch is used to capture discriminative spatiotemporal features for temporal consistency, pretrained in an adversarial manner from unlabeled video data. The stylizing branch is used to transfer the style image to a video frame with the guidance from the temporal prediction branch to ensure temporal consistency. To guide the training of VTNet, we introduce the style-coherence loss net (SCNet), which assembles the content loss, the style loss, and the new designed coherence loss. These losses are computed based on high-level features extracted from a pretrained VGG-16 network. The content loss is used to preserve high-level abstract contents of the input frames, and the style loss introduces new colors and patterns from the style image. Instead of using optical flow to explicitly redress the stylized video frames, we design the coherence loss to make the stylized video inherit the dynamics and motion patterns from the source video to remove temporal flickering. Extensive subjective and objective evaluations on various styles demonstrate that the proposed method achieves favorable results against the state-of-the-arts with high efficiency.
Kai Xu 0013, Longyin Wen, Guorong Li, Honggang Qi, Liefeng Bo, Qingming Huang
IEEE Trans. Image Process.3
2021 Embedding Perspective Analysis Into Multi-Column Convolutional Neural Network for Crowd Counting
abstract
The crowd counting is challenging for deep networks due to several factors. For instance, the networks can not efficiently analyze the perspective information of arbitrary scenes, and they are naturally inefficient to handle the scale variations. In this work, we deliver a simple yet efficient multi-column network, which integrates the perspective analysis method with the counting network. The proposed method explicitly excavates the perspective information and drives the counting network to analyze the scenes. More concretely, we explore the perspective information from the estimated density maps and quantify the perspective space into several separate scenes. We then embed the perspective analysis into the multi-column framework with a recurrent connection. Therefore, the proposed network matches various scales with the different receptive fields efficiently. Secondly, we share the parameters of the branches with various receptive fields. This strategy drives the convolutional kernels to be sensitive to the instances with various scales. Furthermore, to improve the evaluation accuracy of the column with a large receptive field, we propose a transform dilated convolution. The transform dilated convolution breaks the fixed sampling structure of the deep network. Moreover, it needs no extra parameters and training, and the offsets are constrained in a local region, which is designed for the congested scenes. The proposed method achieves state-of-the-art performance on five datasets (ShanghaiTech, UCF CC 50, WorldEXPO'10, UCSD, and TRANCOS).
Guorong Li, Dawei Du, Qingming Huang, Nicu Sebe
IEEE Trans. Image Process.2
2021 Self-Supervised Deep TripleNet for Video Object Segmentation
abstract
Most of previous video object segmentation methods require a large amount of pixel-level annotated video data to construct a robust model. It is quite expensive to label segmentation mask in video. In this paper, we propose a self-supervised triplenet for video object segmentation, which only leverages nearly unlimited unlabeled video data in training phase. Our method consists of two modules, i.e., the temporal motion module and appearance matching module. The temporal motion module is trained based on the pixel correspondence between two video frames in a self-supervised manner, which models the motion patterns between two video frames and propagates the labels from one frame to another. Meanwhile, the appearance matching module encodes the reference frame and its corresponding mask, and generates the segmentation mask of the same object in target frame. The appearance matching module can adjust and refine the outout of temporal motion module, and avoid error accumulation by matching the reference appearance. In order to train the appearance matching module in self-supervised manner, we propose two mask generation strategies: foreground region mask generation and random color region mask generation. Extensive experiments conducted on four challenging video object segmentation datasets, i.e., DAVIS-2017, Youtube-VOS, DAVIS- 2016 and SegTrack v2, demonstrate that the proposed method performs favorable against the state-of-the-art self-supervised methods, and performs even competitively with fully-supervised methods. We also show our self-supervised approach has actually superior generalizability to the majority of supervised methods.
Kai Xu 0013, Longyin Wen, Guorong Li, Qingming Huang
IEEE Trans. Multim.3
2020 Release the Power of Online-Training for Robust Visual Tracking
abstract
Convolutional neural networks (CNNs) have been widely adopted in the visual tracking community, significantly improving the state-of-the-art. However, most of them ignore the important cues lying in the distribution of training data and high-level features that are tightly coupled with the target/background classification. In this paper, we propose to improve the tracking accuracy via online training. On the one hand, we squeeze redundant training data by analyzing the dataset distribution in low-level feature space. On the other hand, we design statistic-based losses to increase the inter-class distance while decreasing the intra-class variance of high-level semantic features. We demonstrate the effectiveness on top of two high-performance tracking methods: MDNet and DAT. Experimental results on the challenging large-scale OTB2015 and UAVDT demonstrate the outstanding performance of our tracking method.
Guorong Li, Yuankai Qi, Qingming Huang
AAAI2
2020 Siamese Box Adaptive Network for Visual Tracking
abstract
Most of the existing trackers usually rely on either a multi-scale searching scheme or pre-defined anchor boxes to accurately estimate the scale and aspect ratio of a target. Unfortunately, they typically call for tedious and heuristic configurations. To address this issue, we propose a simple yet effective visual tracking framework (named Siamese Box Adaptive Network, SiamBAN) by exploiting the expressive power of the fully convolutional network (FCN). SiamBAN views the visual tracking problem as a parallel classification and regression problem, and thus directly classifies objects and regresses their bounding boxes in a unified FCN. The no-prior box design avoids hyper-parameters associated with the candidate boxes, making SiamBAN more flexible and general. Extensive experiments on visual tracking benchmarks including VOT2018, VOT2019, OTB100, NFS, UAV123, and LaSOT demonstrate that SiamBAN achieves state-of-the-art performance and runs at 40 FPS, confirming its effectiveness and efficiency. The code will be available at https://github.com/hqucv/siamban.
Zedu Chen, Bineng Zhong 0001, Guorong Li, Shengping Zhang, Rongrong Ji
CVPR3
2020 Reverse Perspective Network for Perspective-Aware Object Counting
abstract
One of the critical challenges of object counting is the dramatic scale variations, which is introduced by arbitrary perspectives. We propose a reverse perspective network to solve the scale variations of input images, instead of generating perspective maps to smooth final outputs. The reverse perspective network explicitly evaluates the perspective distortions, and efficiently corrects the distortions by uniformly warping the input images. Then the proposed network delivers images with similar instance scales to the regressor. Thus the regression network doesn't need multi-scale receptive fields to match the various scales. Besides, to further solve the scale problem of more congested areas, we enhance the corresponding regions of ground-truth with the evaluation errors. Then we force the regressor to learn from the augmented ground-truth via an adversarial process. Furthermore, to verify the proposed model, we collected a vehicle counting dataset based on Unmanned Aerial Vehicles (UAVs). The proposed dataset has fierce scale variations. Extensive experimental results on four benchmark datasets show the improvements of our method against the state-of-the-arts.
Guorong Li, Zhe Wu 0006, Li Su 0003, Qingming Huang, Nicu Sebe
CVPR2
2020 Weakly-Supervised Crowd Counting Learns from Sorting Rather Than Locations
Guorong Li, Zhe Wu 0006, Li Su 0003, Qingming Huang, Nicu Sebe
ECCV (8)2
2020 Siamese Dynamic Mask Estimation Network for Fast Video Object Segmentation
abstract
Video object segmentation(VOS) has been a fundamental topic in recent years, and many deep learning-based methods have achieved state-of-the-art performance on multiple benchmarks. However, most of these methods rely on pixel-level matching between the template and the searched frames on the whole image while the targets only occupy a small region. Calculating on the entire image brings lots of additional computation cost. Besides, the whole image may contain some distracting information resulting in many false-positive matching points. To address this issue, motivated by one-stage instance object segmentation methods, we propose an efficient siamese dynamic mask estimation network for fast video object segmentation. The VOS is decoupled into two tasks, i.e., mask feature learning and dynamic kernel prediction. The former is responsible for learning high-quality features to preserve structural geometric information, and the latter learns a dynamic kernel that is used to convolve with the mask feature to generate a mask output. We use Siamese neural network as a feature extractor and directly predict masks after correlation. In this way, we can avoid using pixel-level matching, making our framework more simple and efficient. Experiment results on DAVIS 2016 /2017 datasets show that our proposed methods can run at 35 frames per second on NVIDIA RTX TITAN while preserving competitive accuracy.
Dexiang Hong, Guorong Li, Kai Xu 0013, Li Su 0003, Qingming Huang
ICPR2
2020 Generalized Zero-Shot Video Classification via Generative Adversarial Networks
abstract
Zero-shot learning (ZSL) is to classify images according to detailed attribute annotations into new categories that are unseen during the training stage. Generalized zero-shot learning (GZSL) adds seen categories to the test samples. Since the learned classifier has inherent bias against seen categories, GZSL is more challenging than traditional ZSL. However, at present, there is no detailed attribute description dataset for video classification. Therefore, the current zero-shot video classification problem is based on the synthesis of generative adversarial networks trained on seen-class features into unseen-class features for ZSL classification. In order to solve this problem, we propose a description text dataset based on the UCF101 action recognition dataset. To the best of our knowledge, this is the first work to add description of the classes to zero-shot video classification. We propose a new loss function that combines visual features with textual features. We extract text features from the proposed text data set, and constrain the process of generating synthetic features based on the principle that videos with similar text types should be similar. Our method reapplies the traditional zero-shot learning idea to video classification. From the experimental point of view, our proposed dataset and method have a positive impact on the generalized zero-shot video classification.
Mingyao Hong, Guorong Li, Xinfeng Zhang 0001, Qingming Huang
ACM Multimedia2
2020 Video Anomaly Detection Using Open Data Filter and Domain Adaptation
abstract
Video anomaly detection is a very challenging task because of the rarity, openness, and the definition of the anomalies. Researchers pay more attention to the characteristics of anomalies and have proposed a variety of anomaly detection models. However, most existing methods only use normal events to construct anomaly detection models and ignore the diversity and openness of normal events. Actually, because real-world video data often have an open-ended distribution, some normal patterns hardly ever appeared in the training data. In addition, analogous to human experience in identifying anomalies, rare abnormal events can play a certain role in the detection of similar abnormal events in the dataset. Therefore, assuming that a small number of abnormal events are known, we propose a novel supervised anomaly detection model which explicitly detects open normal events and open abnormal events in the dataset and treats open data and seen data with different classifiers. First, we use the training video to train an imbalanced classifier as the seen data classifier. Then, during the testing phase, an open data filter module isused to divide the test data into seen data and open data. Finally, we directly use the seen data classifier to generate anomaly scores for the seen test data. For the open test data, we adopt a domain adaptation method to reduce the distribution difference between it and the training data and train a new classifier to score for it. Extensive experimental results prove the effectiveness of our model.
Chen Zhang 0013, Guorong Li, Li Su 0003, Weigang Zhang, Qingming Huang
VCIP2
2020 CSCNet: A Shallow Single Column Network for Crowd Counting
abstract
Crowd counting in complex scene is an important but challenge task. The scale variation of crowd makes the shallow network hard to extract effective features. In this paper, we propose a shallow single column network named CSCNet for crowd counting. The key component is complementary scale context block (CSCB). It is designed to capture complementary scale context and obtains a high accuracy with limited depth of the network. As far as we know, CSCNet is the shallowest single column network in existing works. We demonstrate our methods on three challenge benchmarks. Compared to state-of-the-art methods, CSCNet achieves comparable accuracy with much less complexity. CSCNet provides an alternative to achieve comparable or even better performance with about 30% of depth and 50% of width decrease. Besides, CSCNet performs more stably on both sparse and congested crowd scenes.
Zhida Zhou, Li Su 0003, Guorong Li, Yifang Yang, Qingming Huang
VCIP3
2020 Real-time Visual Object Tracking with Natural Language Description
abstract
In this work, we argue that conditioning on the natural language (NL) description of a target provides information for longer-term invariance, and thus helps cope with typical tracking challenges. However, deriving a formulation to combine the strengths of appearance-based tracking with the language modality is not straightforward. Therefore, we propose a novel deep tracking-by-detection formulation that can take advantage of NL descriptions. Regions that are related to the given NL description are generated by a proposal network during the detection stage of the tracker. Our LSTM based tracker then predicts the update of the target from regions proposed by the NL based detection stage. Our method runs at over 30 fps on a single GPU. In benchmarks, our method is competitive with state of the art trackers that employ bounding boxes for initialization, while it outperforms all other trackers on targets given unambiguous and precise language annotations. When conditioned on NL descriptions only, our model doubles the performance of the previous best attempt [25].
Qi Feng 0004, Vitaly Ablavsky, Qinxun Bai, Guorong Li, Stan Sclaroff
WACV4
2020 The Unmanned Aerial Vehicle Benchmark: Object Detection, Tracking and Baseline
Hongyang Yu 0001, Guorong Li, Weigang Zhang, Qingming Huang, Dawei Du, Qi Tian 0001, Nicu Sebe
Int. J. Comput. Vis.2
2020 Conditional GAN based individual and global motion fusion for multiple object tracking in UAV videos
Hongyang Yu 0001, Guorong Li, Li Su 0003, Bineng Zhong 0001, Hongxun Yao, Qingming Huang
Pattern Recognit. Lett.2
2019 Spatiotemporal CNN for Video Object Segmentation
abstract
In this paper, we present a unified, end-to-end trainable spatiotemporal CNN model for VOS, which consists of two branches, i.e., the temporal coherence branch and the spatial segmentation branch. Specifically, the temporal coherence branch pretrained in an adversarial fashion from unlabeled video data, is designed to capture the dynamic appearance and motion cues of video sequences to guide object segmentation. The spatial segmentation branch focuses on segmenting objects accurately based on the learned appearance and motion cues. To obtain accurate segmentation results, we design a coarse-to-fine process to sequentially apply a designed attention module on multi-scale feature maps, and concatenate them to produce the final prediction. In this way, the spatial segmentation branch is enforced to gradually concentrate on object regions. These two branches are jointly fine-tuned on video segmentation sequences in an end-to-end manner. Several experiments are carried out on three challenging datasets (i.e., DAVIS-2016, DAVIS-2017 and Youtube-Object) to show that our method achieves favorable performance against the state-of-the-arts. Code is available at https://github.com/longyin880815/STCNN.
Kai Xu 0013, Longyin Wen, Guorong Li, Liefeng Bo, Qingming Huang
CVPR3
2019 Training Efficient Saliency Prediction Models with Knowledge Distillation
abstract
Recently, deep learning-based saliency prediction methods have achieved significant accuracy improvements. However, they are hard to embed in practical multimedia applications due to large memory consumption and running time caused by complicated architectures. In addition, most methods are fine-tuned from pre-trained models for classification tasks, and networks cannot flexibly be transferred for a new task. In this paper, a condensed and randomly initialized student network is employed to achieve higher efficiency by transferring knowledge from complicated and well-trained teacher networks. This is the first use of knowledge distillation for efficient pixel-wise saliency prediction. Instead of directly minimizing Euclidean distance between feature maps, we propose two statistical representations of feature maps (i.e., first-order and second-order statistics) as knowledge. We conduct experiments on three kinds of teacher networks and four benchmark datasets to verify the effectiveness of the proposed method. Compared with the teacher networks, the student networks achieve an acceleration ratio of 4.56-4.73. Compared with state-of-the-art approaches, the proposed model achieves competitive accuracy with faster running speed (up to 4.38 times) and smaller model size (up to 93.27% reduction). We further embedded the proposed saliency prediction model into a video captioning application. The saliency-embedded approaches improve video captioning on all test metrics with a small complexity cost. The student-model embedded approach achieves 25% time saving with similar performance to the teacher embedded one.
Peng Zhang 0024, Li Su 0003, Liang Li 0003, Bing-Kun Bao, Pamela C. Cosman, Guorong Li, Qingming Huang
ACM Multimedia6
2019 Self-balance Motion and Appearance Model for Multi-object Tracking in UAV
abstract
Under the tracking-by-detection framework, multi-object tracking methods try to connect object detections with target trajectories by reasonable policy. Most methods represent objects by the appearance and motion. The inference of the association is mostly judged by a fusion of appearance similarity and motion consistency. However, the fusion ratio between appearance and motion are often determined by subjective setting. In this paper, we propose a novel self-balance method fusing appearance similarity and motion consistency. Extensive experimental results on public benchmarks demonstrate the effectiveness of the proposed method with comparisons to several state-of-the-art trackers.
Hongyang Yu 0001, Guorong Li, Weigang Zhang, Hongxun Yao, Qingming Huang
MMAsia2
2019 Beyond global fusion: A group-aware fusion approach for multi-view image clustering
Zhe Xue, Guorong Li, Shuhui Wang, Jun Huang 0003, Weigang Zhang, Qingming Huang
Inf. Sci.2
2019 Deep Constrained Low-Rank Subspace Learning for Multi-View Semi-Supervised Classification
abstract
Semi-supervised classification receives increasing interests because it can predict class labels based on both limited labeled and sufficient unlabeled data. In this letter, we propose a deep constrained low-rank subspace learning (DCLSL) method for multi-view semi-supervised classification. Specifically, we integrate deep constrained matrix factorization, low-rank subspace learning, and class label learning into a unified objective function to jointly learn data similarity matrices and class label matrix. DCLSL is able to obtain the discriminative subspace representation of each view and effectively aggregate similarity matrices of multiple views, resulting in better classification performance. Experimental results on various datasets demonstrate the effectiveness of our method.
Zhe Xue, Junping Du 0001, Dawei Du, Guorong Li, Qingming Huang, Siwei Lyu
IEEE Signal Process. Lett.4
2018 The Unmanned Aerial Vehicle Benchmark: Object Detection and Tracking
Dawei Du, Yuankai Qi, Hongyang Yu 0001, Kaiwen Duan, Guorong Li, Weigang Zhang, Qingming Huang, Qi Tian 0001
ECCV (10)6
2018 Edge Guided Generation Network for Video Prediction
abstract
Video prediction is a challenging problem due to the highly complex variation of video appearance and motions. Traditional methods that directly predict pixel values often result in blurring and artifacts. Furthermore, cumulative errors can lead to a sharp drop of prediction quality in long-term prediction. To alleviate the above problems, we propose a novel edge guided video prediction network, which firstly models the dynamic of frame edges and predicts the future frame edges, then generates the future frames under the guidance of the obtained future frame edges. Specifically, our network consists of two modules that are ConvLSTM based edge prediction module and the edge guided frames generation module. The whole network is differentiable and can be trained end-to-end without any supervision effort. Extensive experiments on KTH human action dataset and challenging autonomous driving KITTI dataset demonstrate that our method achieves better results than state-of-the-art methods especially in long-term video predictions.
Kai Xu 0013, Guorong Li, Huijuan Xu 0001, Weigang Zhang, Qingming Huang
ICME2
2018 Joint multi-view representation and image annotation via optimal predictive subspace learning
Zhe Xue, Guorong Li, Qingming Huang
Inf. Sci.2
2018 Bilevel Multiview Latent Space Learning
abstract
Different kinds of features describe different aspects of image data, and each feature can be treated as a view when we take it as a particular understanding of images. Leveraging multiple views provides a richer and comprehensive description than using only a single view. However, multiview data are often represented by high-dimensional heterogeneous features, so it is meaningful to find a low-dimensional consensus representation from multiple views. In this paper, we propose an unsupervised multiview dimensionality reduction method for images based on bilevel latent space learning. As different views have different physical meanings and statistical properties, they are not directly comparable. Therefore, we learn the comparable representation for each view in the first level. The shared and the private nature of multiview data are exploited to accurately preserve the information of each view. Then, we fuse different views into a low-dimensional representation by conducting joint matrix factorization in the second level. To guarantee the low-dimensional representation to be compact and discriminative, the intrinsic geometric structure of data is utilized. Besides, our method considers resisting the outliers and noise contained in multiview data, which may influence the learned representation and deteriorate its semantic consistency. We design appropriate optimization objectives to learn the latent spaces in different levels. Compared with the existing methods, our method could provide a more flexible multiview learning strategy that not only accurately captures the information of each view but also is robust to outliers and noise, which can obtain a more discriminative and compact low-dimensional representation. Experiments on two real-world image data sets demonstrate the advantages of our method over the existing multiview dimensionality reduction methods.
Zhe Xue, Guorong Li, Shuhui Wang, Weigang Zhang, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.2
2018 Joint Feature Selection and Classification for Multilabel Learning
abstract
Multilabel learning deals with examples having multiple class labels simultaneously. It has been applied to a variety of applications, such as text categorization and image annotation. A large number of algorithms have been proposed for multilabel learning, most of which concentrate on multilabel classification problems and only a few of them are feature selection algorithms. Current multilabel classification models are mainly built on a single data representation composed of all the features which are shared by all the class labels. Since each class label might be decided by some specific features of its own, and the problems of classification and feature selection are often addressed independently, in this paper, we propose a novel method which can perform joint feature selection and classification for multilabel learning, named JFSC. Different from many existing methods, JFSC learns both shared features and label-specific features by considering pairwise label correlations, and builds the multilabel classifier on the learned low-dimensional data representations simultaneously. A comparative study with state-of-the-art approaches manifests a competitive performance of our proposed method both in classification and feature selection for multilabel learning.
Jun Huang 0003, Guorong Li, Qingming Huang, Xindong Wu 0001
IEEE Trans. Cybern.2
2018 Generalized Semi-supervised and Structured Subspace Learning for Cross-Modal Retrieval
abstract
Motivated by the fact that unlabeled data can be easily collected and help to exploit the correlations among different modalities, this paper proposes a novel method named generalized semi-supervised structured subspace learning (GSS-SL) for the task of cross-modal retrieval. First, to predict more relevant class labels for unlabeled data, we propose a label graph constraint that ensures the intrinsic geometric structures of different feature spaces consistent with that of label space. Second, considering that class labels directly reveal the semantic information of multimedia data, GSS-SL takes the label space as a linkage to model the correlations among different modalities. Concretely, the label graph constraint, label-linked loss function, and regularization are integrated into a joint minimization formulation to learn a discriminative common subspace. Finally, an efficient optimization algorithm is designed to alternately optimize multiple linear transformations for different modalities and update the class indicator matrices for unlabeled data. Furthermore, an arbitrary number of modalities can be solved in the proposed framework. Extensive experiments on three standard benchmark datasets demonstrate that GSS-SL outperforms previous methods on exploiting the correlations among different modalities.
Bingpeng Ma, Guorong Li, Qingming Huang, Qi Tian 0001
IEEE Trans. Multim.3
2017 Metric based on multi-order spaces for cross-modal retrieval
abstract
This paper proposes a novel method for cross-modal retrieval. Different from vector (text)-to-vector (image) framework of the traditional cross-modal methods, we adopt a vector (text)-to-matrix (image) framework. We assume that compared with vectors, matrices can directly represent images and characterize the structure of feature space. Furthermore, we propose a Metric based on Multi-order spaces (MMs). Multi-order statistic features are used to represent images for enriching the semantic information, and metrics among the multi-spaces are jointly learned to measure the similarity between two different modalities. Specifically, there are three steps for MMs. First, we jointly use the bags of visual features (zero-order), mean (first-order) and covariance (second-order) to characterize each image. Second, considering that covariance matrices and vectors lie on a Riemannian manifold and an Euclidean space respectively, we embed multi-order spaces into their corresponding Hilbert spaces to reduce the heterogeneity among the original spaces. Finally, the similarity between two different modalities can be measured by learning multiple transformations from the different Hilbert spaces to a common subspace. The performance of the proposed method over the state-of-the-art has been demonstrated through the experiments on two public datasets.
Bingpeng Ma, Guorong Li, Qingming Huang
ICME3
2017 Adaptively Unified Semi-supervised Learning for Cross-Modal Retrieval
abstract
Motivated by the fact that both relevancy of class labels and unlabeled data can help to strengthen multi-modal correlation, this paper proposes a novel method for cross-modal retrieval. To make each sample moving to the direction of its relevant label while far away from that of its irrelevant ones, a novel dragging technique is fused into a unified linear regression model. By this way, not only the relation between embedded features and relevant class labels but also the relation between embedded features and irrelevant class labels can be exploited. Moreover, considering that some unlabeled data contain specific semantic information, a weighted regression model is designed to adaptively enlarge their contribution while weaken that of the unlabeled data with non-specific semantic information. Hence, unlabeled data can supply semantic information to enhance discriminant ability of classifier. Finally, we integrate the constraints into a joint minimization formulation and develop an efficient optimization algorithm to learn a discriminative common subspace for different modalities. Experimental results on Wiki, Pascal and NUS-WIDE datasets show that the proposed method outperforms the state-of-the-art methods even when we set 20% samples without class labels.
Bingpeng Ma, Guorong Li, Qingming Huang, Qi Tian 0001
IJCAI4
2017 Multi-Networks Joint Learning for Large-Scale Cross-Modal Retrieval
abstract
This paper proposes a novel deep framework of multi-networks joint learning for large-scale cross-modal retrieval. For most existing cross-modal methods, the processes of training and testing don't care about the problem of memory requirement. Hence, they are generally implemented on small-scale data. Moreover, they take feature learning and latent space embedding as two separate steps which cannot generate specific features to accord with the cross-modal task. To alleviate the problems, we first disintegrate the multiplication and inverse of some big matrices, usually involved in existing methods, into that of many sub-matrices. Each sub-matrix is targeted to dispose one pair of image-sentence, for which we further design a novel sampling strategy to select the most representative samples to construct the cross-modal ranking loss and within-modal discriminant loss functions. By this way, the proposed model consumes less memory each time such that it can scale to large-scale data. Furthermore, we apply the proposed discriminative ranking loss to effectively unify two heterogenous networks, deep residual network for images and long short-term memory for sentences, into an end-to-end deep learning architecture. Finally, we can simultaneously achieve specific features adapting to cross-modal task and learn a shared latent space for images and sentences. Extensive evaluations on two large-scale cross-modal datasets show that the proposed method brings substantial improvements over other state-of-the-art ranking methods.
Bingpeng Ma, Guorong Li, Qingming Huang, Qi Tian 0001
ACM Multimedia3
2017 Multi-label classification by exploiting local positive and negative pairwise label correlation
Jun Huang 0003, Guorong Li, Shuhui Wang, Zhe Xue, Qingming Huang
Neurocomputing2
2017 Cross-Modal Retrieval Using Multiordered Discriminative Structured Subspace Learning
abstract
This paper proposes a novel method for cross-modal retrieval. In addition to the traditional vector (text)-to-vector (image) framework, we adopt a matrix (text)-to-matrix (image) framework to faithfully characterize the structures of different feature spaces. Moreover, we propose a novel metric learning framework to learn a discriminative structured subspace, in which the underlying data distribution is preserved for ensuring a desirable metric. Concretely, there are three steps for the proposed method. First, the multiorder statistics are used to represent images and texts for enriching the feature information. We jointly use the covariance (second-order), mean (first-order), and bags of visual (textual) features (zeroth-order) to characterize each image and text. Second, considering that the heterogeneous covariance matrices lie on the different Riemannian manifolds and the other features on the different Euclidean spaces, respectively, we propose a unified metric learning framework integrating multiple distance metrics, one for each order statistical feature. This framework preserves the underlying data distribution and exploits complementary information for better matching heterogeneous data. Finally, the similarity between the different modalities can be measured by transforming the multiorder statistical features to the common subspace. The performance of the proposed method over the previous methods has been demonstrated through the experiments on two public datasets.
Bingpeng Ma, Guorong Li, Qingming Huang, Qi Tian 0001
IEEE Trans. Multim.3
2016 Joint Multi-View Representation Learning and Image Tagging
abstract
Automatic image annotation is an important problem in several machine learning applications such as image search. Since there exists a semantic gap between low-level image features and high-level semantics, the description ability of image representation can largely affect annotation results. In fact, image representation learning and image tagging are two closely related tasks. A proper image representation can achieve better image annotation results, and image tags can be treated as guidance to learn more effective image representation. In this paper, we present an optimal predictive subspace learning method which jointly conducts multi-view representation learning and image tagging. The two tasks can promote each other and the annotation performance can be further improved. To make the subspace to be more compact and discriminative, both visual structure and semantic information are exploited during learning. Moreover, we introduce powerful predictors (SVM) for image tagging to achieve better annotation performance. Experiments on standard image annotation datasets demonstrate the advantages of our method over the existing image annotation methods.
Zhe Xue, Guorong Li, Qingming Huang
AAAI2
2016 Video saliency prediction with optimized optical flow and gravity center bias
abstract
Dynamic videos are viewed fundamentally different from static images. Besides spatial features, motion feature also plays an important role as a temporal factor. Most existing video saliency models usually employ optical flow to represent the motion feature. However, optical flow often suffers from the discontinuity problem. And we also notice that human fixations in one single video frame are much sparser than that in an identical still picture. However, many spatial saliency models take each video frame as static image independently. In this paper, we predict the dynamic visual saliency by fusing spatial and temporal features. In order to construct the temporal relationships among a set of successive frames, we introduce a smoothness operator in optical flow field to obtain more accurate motion feature. Then, considering the sparse property of video saliency, we adapt the weights of the regions surrounding to the saliency gravity center in the final maps. The experiments show that our model is more consistent with humans eye-tracking benchmarks than the state-of-the-art models.
Zhe Wu 0006, Li Su 0003, Qingming Huang, Bo Wu 0016, Guorong Li
ICME6
2016 PL-ranking: A Novel Ranking Method for Cross-Modal Retrieval
abstract
This paper proposes a novel method for cross-modal retrieval named Pairwise-Listwise \textbf{ranking} (PL-ranking) based on the low-rank optimization framework. Motivated by the fact that optimizing the top of ranking is more applicable in practice, we focus on improving the precision at the top of ranked list for a given sample and learning a low-dimensional common subspace for multi-modal data. Concretely, there are three constraints in PL-ranking. First, we use a pairwise ranking loss constraint to optimize the top of ranking. Then, considering that the pairwise ranking loss constraint ignores class information, we further adopt a listwise constraint to minimize the intra-neighbors variance and maximize the inter-neighbors separability. By this way, class information is preserved while the number of iterations is reduced. Finally, low-rank based regularization is applied to exploit the correlations between features and labels so that the relevance between the different modalities can be enhanced after mapping them into the common subspace. We design an efficient low-rank stochastic subgradient descent method to solve the proposed optimization problem. The experimental results show that the average MAP scores of PL-ranking are improved 5.1%, 9.2%, 4.7% and 4.8% than those of the state-of-the-art methods on the Wiki, Flickr, Pascal and NUS-WIDE datasets, respectively.
Bingpeng Ma, Guorong Li, Qingming Huang, Qi Tian 0001
ACM Multimedia3
2016 Beyond appearance model: Learning appearance variations for object tracking
Guorong Li, Bingpeng Ma, Jun Huang 0003, Qingming Huang, Weigang Zhang
Neurocomputing1
2016 Online web video topic detection and tracking with semi-supervised learning
Guorong Li, Shuqiang Jiang, Weigang Zhang, Junbiao Pang, Qingming Huang
Multim. Syst.1
2016 Effective Multimodality Fusion Framework for Cross-Media Topic Detection
abstract
Due to the prevalence of We-Media, information is quickly published and received in various forms anywhere and anytime through the Internet. The rich cross-media information carried by the multimodal data in multiple media has a wide audience, deeply reflects the social realities, and brings about much greater social impact than any single media information. Therefore, automatically detecting topics from cross media is of great benefit for the organizations (i.e., advertising agencies and governments) that care about the social opinions. However, cross-media topic detection is challenging from the following aspects: 1) the multimodal data from different media often involve distinct characteristics and 2) topics are presented in an arbitrary manner among the noisy web data. In this paper, we propose a multimodality fusion framework and a topic recovery (TR) approach to effectively detect topics from cross-media data. The multimodality fusion framework flexibly incorporates the heterogeneous multimodal data into a multimodality graph, which takes full advantage from the rich cross-media information to effectively detect topic candidates (T.C.). The TR approach solidly improves the entirety and purity of detected topics by: 1) merging the T.C. that are highly relevant themes of the same real topic and 2) filtering out the less-relevant noise data in the merged T.C. Extensive experiments on both single-media and cross-media data sets demonstrate the promising flexibility and effectiveness of our method in detecting topics from cross media.
Lingyang Chu, Guorong Li, Shuhui Wang, Weigang Zhang, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.3
2016 Learning Label-Specific Features and Class-Dependent Labels for Multi-Label Classification
abstract
Binary Relevance is a well-known framework for multi-label classification, which considers each class label as a binary classification problem. Many existing multi-label algorithms are constructed within this framework, and utilize identical data representation in the discrimination of all the class labels. In multi-label classification, however, each class label might be determined by some specific characteristics of its own. In this paper, we seek to learn label-specific data representation for each class label, which is composed of label-specific features. Our proposed method LLSF can not only be utilized for multi-label classification directly, but also be applied as a feature selection method for multi-label learning and a general strategy to improve multi-label classification algorithms comprising a number of binary classifiers. Inspired by the research works on modeling high-order label correlations, we further extend LLSF to learn class-Dependent Labels in a sparse stackingway, denoted as LLSF-DL. It incorporates both second-order- and high-order label correlations. A comparative study with the state-of-the-art approaches manifests the effectiveness and efficiency of our proposed methods.
Jun Huang 0003, Guorong Li, Qingming Huang, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.2
2015 Learning Label Specific Features for Multi-label Classification
abstract
Binary relevance (BR) is a well-known framework for multi-label classification. It decomposes multi-label classification into binary (one-vs-rest) classification subproblems, one for each label. The BR approach is a simple and straightforward way for multi-label classification, but it still has several drawbacks. First, it does not consider label correlations. Second, each binary classifier may suffer from the issue of class-imbalance. Third, it can become computationally unaffordable for data sets with many labels. Several remedies have been proposed to solve these problems by exploiting label correlations between labels and performing label space dimension reduction. Meanwhile, inconsistency, another potential drawback of BR, is often ignored by researchers when they construct multi-label classification models. Inconsistency refers to the phenomenon that if an example belongs to more than one class label, then during the binary training stage, it can be considered as both positive and negative example simultaneously. This will mislead binary classifiers to learn suboptimal decision boundaries. In this paper, we seek to solve this problem by learning label specific features for each label. We assume that each label is only associated with a subset of features from the original feature set, and any two strongly correlated class labels can share more features with each other than two uncorrelated or weakly correlated ones. The proposed method can be applied as a feature selection method for multi-label learning and a general strategy to improve multi-label classification algorithms comprising a number of binary classifiers. Comparison with the state-of-the-art approaches manifests competitive performance of our proposed method.
Jun Huang 0003, Guorong Li, Qingming Huang, Xindong Wu 0001
ICDM2
2015 Group sensitive Classifier Chains for multi-label classification
abstract
In multi-label classification, labels often have correlations with each other. Exploiting label correlations can improve the performances of classifiers. Current multi-label classification methods mainly consider the global label correlations. However, the label correlations may be different over different data groups. In this paper, we propose a simple and efficient framework for multi-label classification, called Group sensitive Classifier Chains. We assume that similar examples not only share the same label correlations, but also tend to have similar labels. We augment the original feature space with label space and cluster them into groups, then learn the label dependency graph in each group respectively and build the classifier chains on each group specific label dependency graph. The group specific classifier chains which are built on the nearest group of the test example are used for prediction. Comparison results with the state-of-the-art approaches manifest competitive performances of our method.
Jun Huang 0003, Guorong Li, Shuhui Wang, Weigang Zhang, Qingming Huang
ICME2
2015 GOMES: A group-aware multi-view fusion approach towards real-world image clustering
abstract
Different features describe different views of visual appearance, multi-view based methods can integrate the information contained in each view and improve the image clustering performance. Most of the existing methods assume that the importance of one type of feature is the same to all the data. However, the visual appearance of images are different, so the description abilities of different features vary with different images. To solve this problem, we propose a group-aware multi-view fusion approach. Images are partitioned into groups which consist of several images sharing similar visual appearance. We assign different weights to evaluate the pairwise similarity between different groups. Then the clustering results and the fusion weights are learned by an iterative optimization procedure. Experimental results indicate that our approach achieves promising clustering performance compared with the existing methods.
Zhe Xue, Guorong Li, Shuhui Wang, Chunjie Zhang 0001, Weigang Zhang, Qingming Huang
ICME2
2015 Online learning affinity measure with CovBoost for multi-target tracking
Guorong Li, Qingming Huang, Shuqiang Jiang, Yingkun Xu, Weigang Zhang
Neurocomputing1
2015 Fusing cross-media for topic detection by dense keyword groups
Weigang Zhang, Tianlong Chen 0003, Guorong Li, Junbiao Pang, Qingming Huang, Wen Gao 0001
Neurocomputing3
2014 Web topic detection using a ranked clustering-like pattern across similarity cascades
abstract
In multi-media and social media communities, web topic detection poses two main difficulties that conventional approaches can barely handle: 1) there are large inter-topic variations among web topics; 2) supervised information is rare to identify the real topics. In this paper, we address these problems from the similarity diffusion perspective among objects on web, and present a clustering-like pattern across similarity cascades (SCs). SCs are a series of subgraphs generated by truncating a weighted graph with a set of thresholds, and then maximal cliques are used to describe the topic candidates. Poisson deconvolution is adopted to efficiently identify the real topics from these topic candidates. Experiments demonstrate that our approach outperforms the state-of-the-arts on two datasets. In addition, we report accuracy v.s. false positives per topic (FPPT) curves for performance evaluation. To our knowledge, this is the first complete evaluation of web topic detection at the topic-wise level, and it establishes a new benchmark for this problem.
Fei Jia, Junbiao Pang, Weigang Zhang, Guorong Li, Chunjie Zhang 0001, Qingming Huang, Yugui Liu
ICME4
2014 Web video thumbnail recommendation with content-aware analysis and query-sensitive matching
Weigang Zhang, Chunxi Liu, Zhenjun Wang, Guorong Li, Qingming Huang, Wen Gao 0001
Multim. Tools Appl.4
2014 Face Distortion Recovery Based on Online Learning Database for Conversational Video
abstract
With the real-time requirement for video conversation, the coding system needs to adopt low delay and low complexity strategy to encode conversational videos, which may result in a significant decline of the video quality under the constrained bandwidth of network. In conversational videos, the face region attracts most human attentions. Therefore, recovering the distortion of face region will effectively improve the visual quality of conversational video. Actually, the participants in a conversation are usually unchanged in a relative long period, and similar facial expressions of the participants would be often repetitive. However, conventional video coding methods just consider the correlation of several neighboring frames while the long-range correlation of similar face regions in the whole conversational video has not been fully used. In this paper, we propose a face distortion recovery system to improve the visual quality of decoded conversational video by online learning an own face feature database for each user. First, at the sender side, the face feature database is established and online updated to include different facial expressions of the person. Then, at the receiver side, the low quality face regions in decoded video are recovered with the face patches in the database. Experimental results show that, under low bits rates the proposed method achieves average 5.22 dB gain with small burden to update the database.
Xi Wang 0014, Li Su 0003, Honggang Qi, Qingming Huang, Guorong Li
IEEE Trans. Multim.5
2013 Online Learning Based Face Distortion Recovery for Conversational Video Coding
abstract
In a video conversation, the participants usually remain the same. As the conversation continues, similar facial expressions of the same person would occur intermittently. However, the correlation of similar face features has not been fully used since the conventional methods only focus on independent frames. We set up a face feature database and updated it online to include new facial expressions during the whole conversation. At the receiver side, the database is used to recover the face distortion and thus improve the visual quality. Additionally, the proposed method brings small burden to update the database and is generic to various CODEC.
Xi Wang 0014, Li Su 0003, Qingming Huang, Guorong Li, Honggang Qi
DCC4
2013 An efficient occlusion detection method to improve object trackers
abstract
Occlusion is one of the challenging problems in visual tracking. Most of the previous works alleviate this problem by randomly sampled weak features, or analyze it by methods closely related with specific trackers. In this paper, we propose an effective mechanism to detect the occlusion status by random forests, and embed this method into object trackers with a common strategy. We divide the target region into some regular parts, and extract the pairing features within and outside the parts to encode the structure information of the tracked target. The random forests are online trained to discriminate the occlusion status of the parts using occlusion-dependent samples. Several challenging video sequences are used to verify the model, and it proves that our model is capable of recognizing the occlusion status during tracking. The performance of two typical state-of-the-art object trackers is improved by embedding this occlusion detection method.
Yingkun Xu, Guorong Li, Qingming Huang
ICIP3
2013 Cross-media topic detection: A multi-modality fusion framework
abstract
Detecting topics from Web data attracts increasing attention in recent years. Most previous works on topic detection mainly focus on the data from single medium, however, the rich and complementary information carried by multiple media can be used to effectively enhance the topic detection performance. In this paper, we propose a flexible data fusion framework to detect topics that simultaneously exist in different mediums. The framework is based on a multi-modality graph (MMG), which is obtained by fusing two single-modality graphs together: a text graph and a visual graph. Each node of MMGrepresents a multi-modal data and the edge weight between two nodes jointly measures their content and upload-time similarities. Since the data about the same topic often have similar content and are usually uploaded in a similar period of time, they would naturally form a dense (namely, strongly connected) subgraph in MMG. Such dense subgraph is robust to noise and can be efficiently detected by pair-wise clustering methods. The experimental results on single-medium and cross-media datasets demonstrate the flexibility and effectiveness of our method.
Guorong Li, Lingyang Chu, Shuhui Wang, Weigang Zhang, Qingming Huang
ICME2
2013 SSOCBT: A Robust Semisupervised Online CovBoost Tracker That Uses Samples Differently
abstract
Most existing feature selection methods for object tracking assume that the samples in the previous frames are governed by the same distribution of the labeled samples obtained in the current frame and unlabeled samples collected in the next frame. However, according to our statistical analysis on very common videos, this assumption is not true in many scenarios. As a result, the selected features are not suitable for discriminating between the target from the background in the next frame. A tracking error accumulates and finally the drift problem happens. In this paper, we consider data distribution in tracking from a new perspective to adapt to target's and background's changes. We classify the samples into three categories: auxiliary samples (samples in the previous frames), target samples (samples collected in the current frame), and unlabeled samples (samples obtained in the next frame). To make the best use of them for tracking, we propose a novel semisupervised transfer learning approach that treats samples differently. Specifically, we assume that only target samples follow the same distribution as the unlabeled samples that we want to classify. Then, a novel and interesting semisupervised CovBoost method is developed utilizing the information provided by the three kinds of samples effectively when training the best strong classifier for tracking. Furthermore, we develop a new online updating algorithm for semisupervised CovBoost, making our tracker handle with significant variations of the tracked target and background successfully. Our experimental results demonstrate the advantages of treating samples differently during tracking. Our tracker outperforms state-of-the-art trackers on the benchmark datasets.
Guorong Li, Qingming Huang, Shuqiang Jiang
IEEE Trans. Circuits Syst. Video Technol.1
2012 Online selection of the best k-feature subset for object tracking
Guorong Li, Qingming Huang, Junbiao Pang, Shuqiang Jiang
J. Vis. Commun. Image Represent.1
2012 A Multiple Targets Appearance Tracker Based on Object Interaction Models
abstract
Kernel-based method has been proved to be effective in solving single-target tracking problem. However, facing more complicated multitarget tracking task, most classic kernel-based multitarget trackers do not model the interaction among targets successfully and simply track each target independently. Thus, they usually could not deal with “singularity” problem and fail in tracking the target when occlusions occur or distracters appear. Although multikernel methods may improve the performance by introducing more constraints, how to simulate the relationship and interaction among the tracked targets is still not fully investigated. In this paper, we discuss a very common scenario of multitarget tracking, in which a moving object's motion is not only determined by its virtual destination but also impacted by other neighboring objects. This phenomenon exists in many usual tracking applications such as human tracking, traffic monitoring, video surveillance, and so on, where an object usually moves toward a particular direction but meanwhile detours when close to others to avoid collision. Specifically, we propose a novel interaction model to explain the above phenomenon. Then by defining a new cost function, we embed this interaction model into a kernel-based tracker and further derive our interactive kernel-based multitarget tracker. Experimental results on various datasets demonstrate that our interaction model can alleviate “singularity” problem and, thus, the proposed tracking method could achieve superior performance in multitarget tracking.
Guorong Li, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.1
2011 Treat samples differently: Object tracking with semi-supervised online CovBoost
abstract
Most feature selection methods for object tracking assume that the labeled samples obtained in the next frames follow the similar distribution with the samples in the previous frame. However, this assumption is not true in some scenarios. As a result, the selected features are not suitable for tracking and the “drift” problem happens. In this paper, we consider data's distribution in tracking from a new perspective. We classify the samples into three categories: auxiliary samples (samples in the previous frames), target samples (collected in the current frame) and unlabeled samples (obtained in the next frame). To make the best use of them for tracking, we propose a novel semi-supervised transfer learning approach. Specifically, we assume only target samples follow the same distribution as the unlabeled samples and develop a novel semi-supervised CovBoost method. It could utilize auxiliary samples and unlabeled samples effectively when training the best strong classifier for tracking. Furthermore, we develop a new online updating algorithm for semi-supervised CovBoost, making our tracker handle with significant variations of the tracked target and background successfully. We demonstrate the excellent performance of the proposed tracker on several challenging test videos.
Guorong Li, Qingming Huang, Junbiao Pang, Shuqiang Jiang
ICCV1
2010 Real-time interactive multi-target tracking using kernel-based trackers
abstract
Although kernel-based methods have been demonstrated effectively in solving single-target tracking problem, facing more complicated multi-target tracking task, most of them still suffer from `singularity' problem because they usually simply track each target independently ignoring their interactions and thus fail when occlusions occur or distracters appear. In this paper, we discuss a very common scenario of multi-target tracking, in which a moving object has its own destination in a short time and pays attention on some of its neighboring objects. This phenomenon exists in many practical tracking applications such as human tracking, traffic monitoring, video surveillance, etc., where an object usually moves towards a particular direction but meanwhile detours when close to others to avoid potential collision. In particular, we propose two novel interaction models and embed them into a kernel-based tracking framework to derive our interactive kernel-based multi-target tracker. Experimental results on challenging videos demonstrate its superiority.
Guorong Li, Qingming Huang
ICIP1
2010 Memory matrix: a novel user experience for home video
abstract
Nowadays, various efforts have sprung up aiming to automatically analyze home videos and provide users satisfactory experiences. In this paper, we present a novel user experience for home video called Memory Matrix, which could facilitate users to re-experience the joy of their memories, travelling along not only the time axis but also the space axis. In other words, the video clips (sub-shots) are organized both by taken times and taken locations, which further allows the user to browse home videos taken at similar locations. Moreover, given a specific query in Memory Matrix (row, column), it can also provide the user optional summaries along the time axis or space axis. The summarization scheme in this paper is based on a top-down interest score generation algorithm which automatically propagates the pre-labeled video level interest scores to sub-shot level interest scores. Firstly, the user is asked to provide interest scores to all the video sequences in the home video collection. Then, the video sequences are decomposed into sub-shots which are represented by keyframes. Consequently, we employ multi-scale spatial saliency analysis to remove the foregrounds and model the background scenes based on histogram of visual words. Finally, the interest scores are propagated from video level to sub-shot level by using gradient descent algorithm. Experimental results demonstrate the effectiveness, efficiency, and robustness of our framework.
Qianqian Xu 0001, Guorong Li, Shuqiang Jiang, Qingming Huang
ACM Multimedia3
2008 Object tracking using incremental 2D-LDA learning and Bayes inference
abstract
The appearances of the tracked object and its surrounding background usually change during tracking. As for tracking methods using subspace analysis, fixed subspace basis tends to cause tracking failure. In this paper, a novel tracking method is proposed by using incremental 2D-LDA learning and Bayes inference. Incremental 2D-LDA formulates object tracking as online classification between foreground and background. It updates the row- or/and column-projected matrix efficiently. Based on the current object location and the prior knowledge, the possible locations of the object (candidates) in the next frame are predicted using simple sampling method. Applying 2D-LDA projection matrix and Bayes inference, candidate that maximizes the posterior probability is selected as the target object. Moreover, informative background samples are selected to update the subspace basis. Experiments are performed on image sequences with the object’s appearance variations due to pose, lighting, etc. We also make comparison to incremental 2D-PCA and incremental FDA. The experimental results demonstrate that the proposed method is efficient and outperforms both the compared methods.
Guorong Li, Dawei Liang, Qingming Huang, Shuqiang Jiang, Wen Gao 0001
ICIP1