Peng Zhang 0005

dblp:21/1048-5 · DBLP profile ↗
← Back
75ranked-venue papers
14as first author
35since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 9 first-author · 19 since 2021Artificial intelligence and machine learning · 24 · 5 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FiMLink: Enhancing sparse feature matching via FFT-in-Mamba and dynamic learnable Fourier encoding
Guancheng Jia, Boxiong Sun, Yufei Zha, Peng Zhang 0005
Knowl. Based Syst.6
2026 MiA: A plug-and-play hyperparameter-free Mamba in Attention module for spatial-temporal consistent visual tracking
Guancheng Jia, Yufei Zha, Peng Zhang 0005
Pattern Recognit.5
2026 FIST: Flow-inspired spatio-temporal target state modeling for visual tracking
Boxiong Sun, Tonghao Han, Yufei Zha, Peng Zhang 0005
Pattern Recognit.5
2026 WmLSTM: A Plug-and-Play Window-Level mLSTM-Based Temporal Encoder for Robust Visual Tracking
abstract
Temporal context modeling constitutes a fundamental issue for robust visual tracking. However, existing approaches are plagued by a critical granularity trade-off: frame-level template update mechanisms inevitably introduce background redundancy due to global frame information aggregation, while token-level propagation mechanisms undermine inherent local spatial correlations via independent feature transmission. To resolve this challenge, we propose WmLSTM, a plug-and-play Window-level mLSTM-based temporal encoder that reconfigures the temporal modeling paradigm for visual tracking. First, window-centric modeling retains intra-window spatial correlations while adaptively suppressing background clutter. Second, we pioneer the application of mLSTM in visual tracking, exploiting its explicit memory architecture that outperforms implicit sequence modeling alternatives (e.g., Mamba). Third, our plug-and-play design enables seamless integration with state-of-the-art trackers with minor computational overhead. Extensive experiments on seven benchmark datasets validate that WmLSTMTrack achieves an excellent balance among accuracy, speed, and parameter efficiency, attaining state-of-the-art accuracy on five benchmarks, superior real-time speed (GPU: 201fps, CPU: 47fps), and compact model size (8.22 M parameters). Moreover, the WMLSTM module consistently enhances the performance of diverse trackers, e.g., real-time FERMT-256: +2.3 points SR75 on GOT-10k, non-real-time EVPTrack-224: +1.8 points P on LaSOText, with merely 30 training epochs. The source code is available at https://github.com/Xiaochen918/WmLSTM.
Guancheng Jia, Yufei Zha, Peng Zhang 0005
IEEE Trans. Circuits Syst. Video Technol.5
2026 PromptTrack: Streaming Spatial-Temporal Prompt Learning for RGB-T Tracking
abstract
This paper presents a novel video-level RGB-T tracking paradigm based on prompt learning, termed PromptTrack, which establishes dense spatial-temporal associations through cross-modal interactions. The method introduces streaming temporal prompts to capture continuous target dynamics (e.g., appearance changes and motion trajectories), while leveraging multimodal spatial prompts to utilize complementary RGB and thermal-infrared (TIR) features dynamically. By propagating temporal prompts through consecutive frames and integrating bidirectional spatial interactions between modalities, PromptTrack achieves superior tracking performance in complex scenarios such as occlusion, low illumination, and distractors. The proposed framework employs a unified multimodal encoder with spatial-temporal modeling via multimodal spatial prompt blocks, enabling efficient fusion of RGB-TIR features without requiring domain-specific structure modifications. Extensive experiments on three RGB-T benchmarks (LasHeR, RGBT210, RGBT234) demonstrate that PromptTrack achieves new state-of-the-art performance, with 76.2% in precision rate on LasHeR and outperforming existing methods by ++1.9% in precision rate and +0.5% in success rate on RGBT234. Notably, its modality-agnostic design facilitates seamless generalization to RGB-D and RGB-E tracking domains achieving new benchmarks on DepthTrack, VOT-RGBD2022, and VisEvent datasets.
Hangfei Li, Guangbin Liu, Yufei Zha, Peng Zhang 0005
IEEE Trans. Multim.5
2025 Audio-visual correspondences based joint learning for instrumental playing source separation
Peng Zhang 0005, Siliang Wang, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001
Neurocomputing2
2025 LINR: A Plug-and-Play Local Implicit Neural Representation Module for Visual Object Tracking
abstract
Current one-stream trackers suffer from limitations in distinguishing targets from complex backgrounds owing to their uniform token division strategy. By treating all regions equally, these methods allocate inadequate attention to crucial target details while overemphasizing redundant background information. Consequently, their performance deteriorates significantly in scenarios involving similar distractors or background clutter. In this work, we propose a Local Implicit Neural Representation (LINR) module specifically designed for local fine-grained object modeling. It consists of two key modules: (1) Local Window Selection: Leveraging template-guided CNN-based cross-correlation, it accurately identify crucial target-relevant regions, reducing background information redundant and computation burden. (2) INR-based Window Refinement: Using implicit neural networks, it optimizes token density and spatial continuity to improve local fine-grained instance-level representations, facilitating the discriminative ability between the target and the background. Moreover, the LINR module exhibits three remarkable advantages as a generalized enhancement for visual tracking. Firstly, it is plug-and-play, seamlessly integrating into existing one-stream trackers, both non-real-time and real-time ones, without architectural modifications, achieving significant performance improvements. Secondly, it is highly portable since it does not introduce new loss functions, additional training strategies or data. Thirdly, it is efficiency-friendly, having minimal impact on model parameters and tracking speed,e.g., AQATrack-LINR increases only 1.9% of the parameters and reduces the tracking speed by only 6fps. We incorporate the LINR module into two non-real-time trackers, OSTrack based on ViT-B and AQATrack based on HiViT-B, and one real-time tracker, FERMT based on ViT-tiny, respectively. The resultant OSTrack-LINR, AQATrack-LINR, and FERMT-LINR achieve state-of-the-art performance across seven widely utilized datasets, such as TrackingNet, LaSOT, and NFS30. The source code is available at https://github.com/Xiaochen918/LINR.
Guancheng Jia, Yufei Zha, Peng Zhang 0005, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Flexible Temperature Parallel Distillation for Dense Object Detection: Make Response-Based Knowledge Distillation Great Again
abstract
Feature-based approaches have been the focal point of previous research on knowledge distillation (KD) for dense object detection. These methods employ feature imitation and result in competitive performance. Despite being able to achieve comparable performance in image recognition, response-based KD methods can not reach the same level in dense object detection. Inspired by improving distillation performance from two key aspects: where to distill and how to distill, in this paper, a parallel distillation (PD) is introduced to fully utilize the sophisticated detection head and transfer all the output responses from the teacher to the student efficiently. In particular, the proposed PD takes an important consideration of the specific location of distillation, which is crucial for effective knowledge transfer. Regarding the discrepancies in output responses between the localization branch and the classification branch, we propose a novel Dynamic Localization Temperature (DLT) module to enhance the precision of distilling localization information. As for the classification branch, a Classification Temperature-Free (CTF) module is also designed to increase the robustness of distillation in heterogeneous networks. By incorporating the DLT and CTF into the PD framework to avoid setting temperature values manually, the Flexible Temperature Parallel Distillation (FTPD) is proposed to achieve a state-of-the-art (SOTA) performance, which can also be further combined with mainstream feature-based methods for better results. In terms of accuracy and robustness with extensive experiments, the proposed FTPD outperforms other KD methods in the task of dense object detection.
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Toward Unifying Saliency Transformer for Video Saliency Prediction and Detection
abstract
Video saliency prediction and detection are thriving research domains that enable computers to simulate the distribution of visual attention akin to how humans perceive dynamic scenes. While many approaches have crafted task-specific training paradigms for either video saliency prediction or video salient object detection tasks, few attention has been devoted to devising a generalized saliency modeling framework that seamlessly bridges both these distinct tasks. In this study, we introduce the Unified Saliency Transformer (UniST) framework, which comprehensively utilizes the essential attributes of video saliency prediction and video salient object detection. In addition to extracting representations of frame sequences, a saliency-aware transformer is designed to learn the spatio-temporal representations at progressively increased resolutions, while incorporating effective cross-scale saliency information to produce a robust representation. Furthermore, task-specific decoders are proposed to perform the final prediction for each task. To the best of our knowledge, this is the first work to explore the design of a unified framework for both saliency modeling tasks. Convincible experiments demonstrate that the proposed UniST achieves superior performance across eight challenging benchmarks for two tasks, outperforming other state-of-the-art methods in most metrics. The project page ishttps://junwenxiong.github.io/UniST.
Junwen Xiong, Chuanyue Li, Peng Zhang 0005, Yue Huo, Wei Huang 0013, Yufei Zha
IEEE Trans. Circuits Syst. Video Technol.4
2024 DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
abstract
Audio-visual saliency prediction can draw support from diverse modality complements, but further performance enhancement is still challenged by customized architectures as well as task-specific loss functions. In recent studies, denoising diffusion models have shown more promising in unifying task frameworks owing to their inherent ability of generalization. Following this motivation, a novel Diffusion architecture for generalized audio-visual Saliency prediction (DiffSal) is proposed in this work, which formulates the prediction problem as a conditional generative task of the saliency map by utilizing input audio and video as the conditions. Based on the spatiotemporal audio-visual features, an extra network Saliency-UNet is designed to perform multimodal attention modulation for progressive refinement of the ground-truth saliency map from the noisy map. Extensive experiments demonstrate that the proposed DiffSal can achieve excellent performance across six challenging audio-visual benchmarks, with an average relative improvement of 6.3% over the previous state-of-the-art results by six metrics. The project url is htt ps: //junwenxiong. github.io/DiffSal.
Junwen Xiong, Peng Zhang 0005, Tao You, Chuanyue Li, Wei Huang 0013, Yufei Zha
CVPR2
2024 Unidirectional Cross-Modal Fusion for RGB-T Tracking
abstract
The key issue of RGB-T tracking is to obtain an effective multimodal representation of targets by utilizing complementary RGB and TIR modality information. Previous methods of template fusion or bidirectional search-template interaction potentially diminish the target representation, resulting from noise information of both templates and search regions. Meanwhile, the direct fusion of sole search features without interacting with templates cannot fully utilize target-relevant contextual information. To mitigate these issues, we present UCTrack, which fuses complementary multimodal search features conditioned on undisturbed RGB and TIR template features. Specifically, we design a Unidirectional Cross-modal Fusion (UCF) module to effectively minimize the influence of background noise on templates by pruning the unnecessary template-to-search cross-modal interaction and to mutually enhance RGB and TIR search features with target-relevant information through multimodal spatial fusion. Furthermore, this module is seamlessly integrated into different layers of a ViT backbone to facilitate feature extraction and cross-modal fusion for RGB-T tracking. Benefiting from the UCF module, UCTrack can effectively and accurately represent multimodal target features without unnecessary template-to-search interaction flow and direct template fusion, making the first proposal of unidirectional cross-modal fusion paradigm for RGB-T tracking to our best knowledge. Extensive experiments on three popular RGB-T tracking benchmarks demonstrate that our method achieves state-of-the-art performance.
Hangfei Li, Yufei Zha, Peng Zhang 0005
ECAI4
2024 How does Layer Normalization improve Batch Normalization in self-supervised sound source localization?
Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
Neurocomputing2
2024 Closed-loop unified knowledge distillation for dense object detection
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
Pattern Recognit.2
2024 Facial Action Unit Representation Based on Self-Supervised Learning With Ensembled Priori Constraints
abstract
Facial action units (AUs) focus on a comprehensive set of atomic facial muscle movements for human expression understanding. Based on supervised learning, discriminative AU representation can be achieved from local patches where the AUs are located. Unfortunately, accurate AU localization and characterization are challenged by the tremendous manual annotations, which limits the performance of AU recognition in realistic scenarios. In this study, we propose an end-to-end self-supervised AU representation learning model (SsupAU) to learn AU representations from unlabeled facial videos. Specifically, the input face is decomposed into six components using auto-encoders: five photo-geometric meaningful components, together with 2D flow field AUs. By constructing the canonical neutral face, posed neutral face, and posed expressional face gradually, these components can be disentangled without supervision, therefore the AU representations can be learned. To construct the canonical neutral face without manually labeled ground truth of emotion state or AU intensity, two priori knowledge based assumptions are proposed: 1) identity consistency, which explores the identical albedos and depths of different frames in a face video, and helps to learn the camera color mode as an extra cue for canonical neutral face recovery. 2) average face, which enables the model to discover a 'neutral facial expression' of the canonical neutral face and decouple the AUs in representation learning. To the best of our knowledge, this is the first attempt to design self-supervised AU representation learning method based on the definition of AUs. Substantial experiments on benchmark datasets have demonstrated the superior performance of the proposed work in comparison to other state-of-the-art approaches, as well as an outstanding capability of decomposing input face into meaningful factors for its reconstruction. The code is made available at https://github.com/Sunner4nwpu/SsupAU.
Peng Zhang 0005, Chujia Guo, Ke Lu 0002, Dongmei Jiang
IEEE Trans. Image Process.2
2024 Auto Diagnosis of Parkinson's Disease Via a Deep Learning Model Based on Mixed Emotional Facial Expressions
abstract
Parkinson's disease (PD) is a common degenerative disease of the nervous system in the elderly. The early diagnosis of PD is very important for potential patients to receive prompt treatment and avoid the aggravation of the disease. Recent studies have found that PD patients always suffer from emotional expression disorder, thus forming the characteristics of "masked faces". Based on this, we thus propose an auto PD diagnosis method based on mixed emotional facial expressions in the paper. Specifically, the proposed method is cast into four steps: Firstly, we synthesize virtual face images containing six basic expressions (i.e., anger, disgust, fear, happiness, sadness, and surprise) via generative adversarial learning, in order to approximate the premorbid expressions of PD patients; Secondly, we design an effective screening scheme to assess the quality of the above synthesized facial expression images and then shortlist the high-quality ones; Thirdly, we train a deep feature extractor accompanied with a facial expression classifier based on the mixture of the original facial expression images of the PD patients, the high-quality synthesized facial expression images of PD patients, and the normal facial expression images from other public face datasets; Finally, with the well-trained deep feature extractor, we thus adopt it to extract the latent expression features for six facial expression images of a potential PD patient to conduct PD/non-PD prediction. To show real-world impacts, we also collected a new facial expression dataset of PD patients in collaboration with a hospital. Extensive experiments are conducted to validate the effectiveness of the proposed method for PD diagnosis and facial expression recognition.
Wei Huang 0013, Renjie Wan, Peng Zhang 0005, Yufei Zha
IEEE J. Biomed. Health Informatics4
2023 CASP-Net: Rethinking Video Saliency Prediction from an Audio-Visual Consistency Perceptual Perspective
abstract
Incorporating the audio stream enables Video Saliency Prediction (VSP) to imitate the selective attention mechanism of human brain. By focusing on the benefits of joint auditory and visual information, most VSP methods are capable of exploiting semantic correlation between vision and audio modalities but ignoring the negative effects due to the temporal inconsistency of audio-visual intrinsics. Inspired by the biological inconsistency-correction within multi-sensory information, in this study, a consistency-aware audio-visual saliency prediction network (CASP-Net) is proposed, which takes a comprehensive consideration of the audio-visual semantic interaction and consistent perception. In addition a two-stream encoder for elegant association between video frames and corresponding sound source, a novel consistency-aware predictive coding is also designed to improve the consistency within audio and visual representations iteratively. To further aggregate the multi-scale audio-visual information, a saliency decoder is introduced for the final saliency map generation. Substantial experiments demonstrate that the proposed CASP-Net outperforms the other state-of-the-art methods on six challenging audio-visual eye-tracking datasets. For a demo of our system please see our project webpage.
Junwen Xiong, Ganglai Wang, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Guangtao Zhai
CVPR3
2023 Semi-Supervised Multimodal Emotion Recognition with Class-Balanced Pseudo-labeling
abstract
This paper presents our solution for the Semi-Supervised Multimodal Emotion Recognition Challenge (MER2023-SEMI), addressing the issue of limited annotated data in emotion recognition. Recently, the self-training-based Semi-Supervised Learning~(SSL) method has demonstrated its effectiveness in various tasks, including emotion recognition. However, previous studies focused on reducing the confirmation bias of data without adequately considering the issue of data imbalance, which is of great importance in emotion recognition. Additionally, previous methods have primarily focused on unimodal tasks and have not considered the inherent multimodal information in emotion recognition tasks. We propose a simple yet effective semi-supervised multimodal emotion recognition method to address the above issues. We assume that the pseudo-labeled samples with consistent results across unimodal and multimodal classifiers have a more negligible confirmation bias. Based on this assumption, we suggest using a class-balanced strategy to select top-k high-confidence pseudo-labeled samples from each class. The proposed method is validated to be effective on the MER2023-SEMI Grand Challenge, with the weighted F1 score reaching 88.53% on the test set.
Chujia Guo, Yan Li 0121, Peng Zhang 0005, Dongmei Jiang
ACM Multimedia4
2023 Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source Localization
abstract
Self-supervised sound source localization is usually challenged by the modality inconsistency. In recent studies, contrastive learning based strategies have shown promising to establish such a consistent correspondence between audio and sound sources in visual scenarios. Unfortunately, the insufficient attention to the heterogeneity influence in the different modality features still limits this scheme to be further improved, which also becomes the motivation of our work. In this study, an Induction Network is proposed to bridge the modality gap more effectively. By decoupling the gradients of visual and audio modalities, the discriminative visual representations of sound sources can be learned with the designed Induction Vector in a bootstrap manner, which also enables the audio modality to be aligned with the visual modality consistently. In addition to a visual weighted contrastive loss, an adaptive threshold selection strategy is introduced to enhance the robustness of the Induction Network. Substantial experiments conducted on SoundNet-Flickr and VGG-Sound Source datasets have demonstrated a superior performance compared to other state-of-the-art works in different challenging scenarios. The code is available at https://github.com/Tahy1/AVIN.
Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
ACM Multimedia2
2023 Efficient thermal infrared tracking with cross-modal compress distillation
Hangfei Li, Yufei Zha, Huanyu Li 0003, Peng Zhang 0005, Wei Huang 0013
Eng. Appl. Artif. Intell.4
2023 Object detection based on cortex hierarchical activation in border sensitive mechanism and classification-GIou joint representation
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
Pattern Recognit.2
2023 Conditional invertible image re-scaling
Yufei Zha, Peng Zhang 0005, Wei Huang 0013
Pattern Recognit.3
2023 A novel evolutionary algorithm inspired from triangle search and its applications on parameters identification of photovoltaic models
Zhenglei Wei, Huan Zhou 0004, Fei Cen, Lei Xie 0001, Peng Zhang 0005, Qinzhi Hao
Soft Comput.6
2023 Facial Expression Guided Diagnosis of Parkinson's Disease via High-Quality Data Augmentation
abstract
Parkinson's disease (PD) is a neurodegenerative disease which is prevalent among the elder population and severely affects the life quality of patients and their families. Therefore, it is important to conduct an early diagnosis for potential patients with PD, so as to promote prompt treatment and avoid the aggravation of the disease. Recently, the in-vitro PD diagnosis based on facial expressions has received increasing attention because of its distinguishability (i.e., PD patients always possess the characteristics of “masked face”) and affordability. However, the performance of the existing facial expression-based PD diagnosis approaches is limited by: 1) the small-scale training data on PD patients' facial expressions, and 2) the weak prediction model. To address these two problems, we propose a new facial expression guided PD diagnosis method based on high-quality training data augmentation and deep neural network prediction. Specifically, the proposed method consists of three stages: Firstly, we synthesize virtual facial expression images with 6 basic emotions (i.e., anger, disgust, fear, happiness, sadness, and surprise) based on multi-domain adversarial learning to approximate the premorbid expressions of PD patients. Secondly, we introduce three facial image quality assessment (FIQA) criteria to measure the quality of these synthesized facial expression images and design a fusion screening strategy that shortlists the high-quality ones to augment the training data. Finally, we train a deep neural network prediction model based on the original and synthesized high-quality facial expression images for PD diagnosis. To show real-world impacts and evaluate the proposed method under different facial expressions, we also create a (currently largest) multiple facial expressions-based PD face dataset in collaboration with a hospital. Extensive experiments are performed to demonstrate the effectiveness of the multi-domain adversarial learning-based facial expression synthesis and the fusion screening strategy, particularly the superior performance of the proposed method for PD diagnosis.
Wei Huang 0013, Yintao Zhou, Yiu-Ming Cheung, Peng Zhang 0005, Yufei Zha
IEEE Trans. Multim.4
2023 Look&listen: Multi-Modal Correlation Learning for Active Speaker Detection and Speech Enhancement
abstract
Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed architecture has been widely used in correspondence to each single task. This may lead to the representation learned by the model being task-specific, and inevitably result in the lack of generalization ability of the feature based on multi-modal modeling. More recent studies have shown that establishing cross-modal relationship between auditory and visual stream is a promising solution for the challenge of audio-visual multi-task learning. Therefore, as a motivation to bridge the multi-modal associations in audio-visual tasks, a unified framework is proposed to achieve target speaker detection and speech enhancement with joint learning of audio-visual modeling in this study. With the assistance of audio-visual channels of videos in challenging real-world scenarios, the proposed method is able to exploit inherent correlations in both audio and visual signals, which is used to further anticipate and model the temporal audio-visual relationships across spatial-temporal space via a cross-modal conformer. In addition, a plug-and-play multi-modal layer normalization is introduced to alleviate the distribution misalignment of multi-modal features. Based on cross-modal circulant fusion, the proposed model is capable to learned all audio-visual representations in a holistic process. Substantial experiments demonstrate that the correlations between different modalities and the associations among diverse tasks can be learned by the optimized model more effectively. In comparison to other state-of-the-art works, the proposed work shows a superior performance for active speaker detection and audio-visual speech enhancement on three benchmark datasets, also with a favorable generalization in diverse challenges.
Junwen Xiong, Peng Zhang 0005, Lei Xie 0001, Wei Huang 0013, Yufei Zha
IEEE Trans. Multim.3
2022 Information Lossless Multi-modal Image Generation for RGB-T Tracking
Yufei Zha, Lichao Zhang 0001, Peng Zhang 0005, Lang Chen
PRCV (4)4
2022 Semantic-aware spatial regularization correlation filter for visual tracking
abstract
Abstract Correlation filters with convolutional neural network (CNN) features have been successfully applied to visual tracking owing to their impressive combined capability for object representation. Unfortunately, further performance improvement is limited due to unwanted boundary effects of the circular structure. In this work, through an in‐depth study of the features’ characteristics, the authors propose a novel tracking strategy to achieve simultaneous filter matching and regularization with CNN features when tracking is on the fly. With a feature decomposed transform matrix, a spatial semantic regularization is generated to reduce the boundary effect effectively during filter optimization. Before each output, the regularized filter is then back performed to match with the extracted features of a search region to find the optimum candidate. Specifically, the most important advantage of the proposed spatial semantic map is to initialize only in the first frame as all the other tracking strategies. Besides, the authors design a novel updating strategy to tackle the cases where the object is occluded or disappeared in the scene. At this time, the maximum of the map is small, even negative. A substantial experiment has been carried out on the popular benchmark tracking datasets; the reliable results have demonstrated that the authors’ method is able to outperform most of the state‐of‐the‐art tracking works in both accuracy and robustness.
Yufei Zha, Peng Zhang 0005, Lei Pu, Lichao Zhang 0001
IET Comput. Vis.2
2022 One-shot Video Graph Generation for Explainable Action Reasoning
Tao Zhuo, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001, Mohan Kankanhalli
Neurocomputing3
2022 A novel locally-constrained GAN-based ensemble to synthesize arterial spin labeling images
Wei Huang 0013, Mingyuan Luo, Jing Li 0027, Peng Zhang 0005, Yufei Zha
Inf. Sci.4
2022 Identity-Aware Facial Expression Recognition Via Deep Metric Learning Based on Synthesized Images
abstract
Person-dependent facial expression recognition has received considerable research attention in recent years. Unfortunately, different identities can adversely influence recognition accuracy, and the recognition task becomes challenging. Other adverse factors, including limited training data and improper measures of facial expressions, can further contribute to the above dilemma. To solve these problems, a novel identity-aware method is proposed in this study. Furthermore, this study also represents the first attempt to fulfill the challenging person-dependent facial expression recognition task based on deep metric learning and facial image synthesis techniques. Technically, a StarGAN is incorporated to synthesize facial images depicting different but complete basic emotions for each identity to augment the training data. Then, a deep-convolutional-neural-network-based network is employed to automatically extract latent features from both real facial images and all synthesized facial images. Next, a Mahalanobis metric network trained based on extracted latent features outputs a learned metric that measures facial expression differences between images, and the recognition task can thus be realized. Extensive experiments based on several well-known publicly available datasets are carried out in this study for performance evaluations. Person-dependent datasets, including CK+, Oulu (all 6 subdatasets), MMI, ISAFE, ISED, etc., are all incorporated. After comparing the new method with several popular or state-of-the-art facial expression recognition methods, its superiority in person-dependent facial expression recognition can be proposed from a statistical point of view.
Wei Huang 0013, Peng Zhang 0005, Yufei Zha, Yuming Fang 0001, Yanning Zhang 0001
IEEE Trans. Multim.3
2021 Unsupervised Cross-Modal Distillation for Thermal Infrared Tracking
abstract
The target representation learned by convolutional neural networks plays an important role in Thermal Infrared (TIR) tracking. Currently, most of the top-performing TIR trackers are still employing representations learned by the model trained on the RGB data. However, this representation does not take into account the information in the TIR modality itself, limiting the performance of TIR tracking.
Jingxian Sun 0003, Lichao Zhang 0001, Yufei Zha, Abel Gonzalez-Garcia, Peng Zhang 0005, Wei Huang 0013, Yanning Zhang 0001
ACM Multimedia5
2021 Multiple object tracking based on multi-task learning with strip attention
abstract
Abstract Multiple object tracking (MOT) framework based on bifurcate strategy was usually challenged by data association of different model path, which work for object localisation and appearance embedding independently. By incorporating the re‐identification (re‐ID) as appearance embedding model, more recent studies on task combination of a single network have made a great progress in tracking performance. Unfortunately, the contributive improvement from re‐ID model is hard to balance the accuracy and efficiency for the whole framework. For more effective enhancement of the overall tracking performance, a real‐time detection needs to be taken into consideration with other auxiliary means for MOT modelling. Therefore, in this study, a one‐shot multiple object tracking is proposed based on multi‐task learning to obtain satisfactory performance in both speed and robustness. With updated re‐training strategy for the backbone model of detection, a D2LA network is proposed to achieve more characteristic fine‐grained feature extraction in branching task of pedestrian recognition. Additionally, a strip attention module is also introduced to further strengthen the feature discriminative capability of the tracking framework in occlusion. Experiments on the 2DMOT15, MOT16, MOT17, and MOT20 benchmark data sets have shown a superior performance in comparison to other state‐of‐the‐art tracking approaches.
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001
IET Image Process.2
2021 Learning spatial-channel regularization jointly with correlation filter for visual tracking
Yufei Zha, Zhuling Qiu, Jingxian Sun 0003, Peng Zhang 0005, Wei Huang 0013
Neurocomputing4
2021 Full-scaled deep metric learning for pedestrian re-identification
Wei Huang 0013, Mingyuan Luo, Peng Zhang 0005, Yufei Zha
Multim. Tools Appl.3
2021 A novel multi-loss-based deep adversarial network for handling challenging cases in semi-supervised image semantic segmentation
Wei Huang 0013, Zhanfei Shao, Mingyuan Luo, Peng Zhang 0005, Yufei Zha
Pattern Recognit. Lett.4
2021 Multiple Instance Models Regression for Robust Visual Tracking
abstract
In comparison to single-model based trackers, the model-ensembled tracking strategy has shown a substantial adaptivity in handling various tracking challenges. As the performance of the tracker has been improved by combining different model outputs linearly, the insufficient consideration of each ensemble member’s contribution still limits the tracking performance to be further enhanced. As the performance of the tracker has been improved by combining different model outputs linearly, the insufficient consideration of each ensemble member’s contribution still limits the tracking performance to be further enhanced. In this paper, a tracking strategy based on multiple instance models regression (MIMRT) is proposed with a unified ensembling scheme. By formulating the tracking initialization with an instance model, the encoding process for an object’s specific detail is performed corresponding to the samples in each frame. The advantage of this operation is to guarantee the model frame-wise discrimination of short-term training, as well as to evaluate the reliability of each instance model by utilizing the long-lifetime samples obtained throughout the whole tracking procedure. To finalize the proposed tracking, all the independent instance models attached to the learned regression coefficients are ensembled with respect to the long-lifetime samples. This also effectively bridges the instance model as a latent variable to investigate a semantic association between the tracking model and the overall samples. A comprehensive experiment has shown that the proposed tracker is able to achieve superior performances compared to the state-of-art tracking approaches on both short-term datasets (e.g., OTB2013, OTB100, VOT2016, UAV123) and long-term dataset (UAV20L).
Yufei Zha, Yuanqiang Zhang, Tao Ku, Hanqiao Huang, Wei Huang 0013, Peng Zhang 0005
IEEE Trans. Circuits Syst. Video Technol.6
2020 Robust Visual Tracking based on Adversarial Unlabeled Instance Generation with Label Smoothing Loss Regularization
Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Garth Douglas Cooper, Yanning Zhang 0001
Pattern Recognit.2
2020 Unsupervised Online Video Object Segmentation With Motion Property Understanding
abstract
Unsupervised video object segmentation aims to automatically segment moving objects over an unconstrained video without any user annotation. So far, only few unsupervised online methods have been reported in the literature, and their performance is still far from satisfactory because the complementary information from future frames cannot be processed under online setting. To solve this challenging problem, in this paper, we propose a novel unsupervised online video object segmentation (UOVOS) framework by construing the motion property to mean moving in concurrence with a generic object for segmented regions. By incorporating the salient motion detection and the object proposal, a pixel-wise fusion strategy is developed to effectively remove detection noises, such as dynamic background and stationary objects. Furthermore, by leveraging the obtained segmentation from immediately preceding frames, a forward propagation algorithm is employed to deal with unreliable motion detection and object proposals. Experimental results on several benchmark datasets demonstrate the efficacy of the proposed method. Compared to state-of-the-art unsupervised online segmentation algorithms, the proposed method achieves an absolute gain of 6.2%. Moreover, our method achieves better performance than the best unsupervised offline algorithm on the DAVIS-2016 benchmark dataset. Our code is available on the project website: https://www.github.com/visiontao/uovos.
Tao Zhuo, Zhiyong Cheng 0001, Peng Zhang 0005, Yongkang Wong, Mohan Kankanhalli
IEEE Trans. Image Process.3
2020 Ensemble Tracking Based on Diverse Collaborative Framework With Multi-Cue Dynamic Fusion
abstract
Tracking with deep neural networks has been verified to arrive at a new level accuracy in many challenging scenarios, but the tracking robustness has been still challenged by model singularity and self-learning loop mechanism. As a promising solution for the limitations, to ensemble diverse tracking strategies into a highly-interactive framework has shown a potential effectiveness in recent studies. In this work, a collaborative tracking framework is proposed by exploiting both discriminative correlation filters and deep classifiers into an ensembling framework. With a multi-cue dynamic fusion scheme performed on all the ensembled members’ outputs, a robust long-term tracking can be achieved by calculating the optimal robustness scores based on a dynamic weighted sum of multi-cue metrics. Meanwhile, the obtained reliable and diverse training samples are also utilized to adaptively update the tracker in each branch with heuristic frequency, which is able to alleviate the training samples’ contamination and model corruption. Experiments on the OTB-2015, Temple color 128, UAV123, VOT2016, and VOT2018 benchmark datasets have shown superior performance in comparison to other state-of-the-art tracking approaches.
Peng Zhang 0005, Tao Zhuo, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001
IEEE Trans. Multim.2
2020 Deep Position-Sensitive Tracking
abstract
Classification-based tracking strategies often face more challenges from intra-class discrimination than from inter-class separability. Even for deep convolutional neural networks that have been widely proven to be effective in various vision tasks, their intra-class discriminative capability is still limited by the weakness of softmax loss, especially for targets not seen in the training dataset. By taking intrinsic attributes of training samples into account, in this paper, we propose a position-sensitive loss coupled with softmax loss to achieve intra-class compactness and inter-class explicitness. Particularly, two additive margins are introduced to encode the position attribute for decision boundary maximization, which is also utilized with the proposed loss to supervise the fine-tuned features on the pre-trained model. With the nearest neighbor ranking measurement in the feature embedding domain, the whole scheme is able to reach an optimized balance between the feature-level inter-class semantic separability and instance-level intra-class relative distance ranking. We evaluate the proposed work on different popular benchmarks, and experimental results demonstrate that our tracking strategy performs favorably against most of the state-of-the-art trackers in the comparison of accuracy and robustness.
Yufei Zha, Tao Ku, Yunqiang Li, Peng Zhang 0005
IEEE Trans. Multim.4
2019 Deep Audio-visual System for Closed-set Word-level Speech Recognition
abstract
Audio-visual understanding is usually challenged by the complementary gap between audio and visual informative bridging. Motivated by the recent audio-visual studies, a closed-set word-level speech recognition scheme is proposed for the Mandarin Audio-Visual Speech Recognition (MAVSR) Challenge in this study. To achieve respective audio and visual encoder initialization more effectively, a 3-dimensional convolutional neural network (CNN) and an attention-based bi-directional long short-term memory (Bi-LSTM) network are trained. With two fully connected layers in addition to the concatenated encoder outputs for the audio-visual joint training, the proposed scheme won the first place with a relative word accuracy improvement of 7.9% over the solitary audio system. Experiments on LRW-1000 dataset have substantially demonstrated that the proposed joint training scheme by audio-visual incorporation is capable of enhancing the recognition performance of relatively short duration samples, unveiling the multi-modal complementarity.
Yougen Yuan, Minhao Fan, Peng Zhang 0005, Lei Xie 0001
ICMI5
2019 Arterial Spin Labeling Images Synthesis via Locally-Constrained WGAN-GP Ensemble
Wei Huang 0013, Mingyuan Luo, Xi Liu 0008, Peng Zhang 0005, Huijun Ding, Dong Ni 0001
MICCAI (4)4
2019 Explainable Video Action Reasoning via Prior Knowledge and State Transitions
abstract
Human action analysis and understanding in videos is an important and challenging task. Although substantial progress has been made in past years, the explainability of existing methods is still limited. In this work, we propose a novel action reasoning framework that uses prior knowledge to explain semantic-level observations of video state changes. Our method takes advantage of both classical reasoning and modern deep learning approaches. Specifically, prior knowledge is defined as the information of a target video domain, including a set of objects, attributes and relationships in the target video domain, as well as relevant actions defined by the temporal attribute and relationship changes (i.e. state transitions). Given a video sequence, we first generate a scene graph on each frame to represent concerned objects, attributes and relationships. Then those scene graphs are associated by tracking objects across frames to form a spatio-temporal graph (also called video graph), which represents semantic-level video states. Finally, by sequentially examining each state transition in the video graph, our method can detect and explain how those actions are executed with prior knowledge, just like the logical manner of thinking by humans. Compared to previous works, the action reasoning results of our method can be explained by both logical rules and semantic-level observations of video content changes. Besides, the proposed method can be used to detect multiple concurrent actions with detailed information, such as who (particular objects), when (time), where (object locations) and how (what kind of changes). Experiments on a re-annotated dataset CAD-120 show the effectiveness of our method.
Tao Zhuo, Zhiyong Cheng 0001, Peng Zhang 0005, Yongkang Wong, Mohan Kankanhalli
ACM Multimedia3
2019 ACFT: adversarial correlation filter for robust tracking
abstract
Tracking based on correlation filters has demonstrated outstanding performance in recent visual object tracking studies and competitions. However, the performance is limited since the boundary effects are introduced by the intrinsic circular structure. In this study, a tracker, called adversarial correlation filter tracker (ACFT), is proposed to solve the above problem through Generative Adversarial Networks (GANs) that is specifically strong at producing realistic‐looking data from noise circumstances. Especially, a mask is generated by the GANs to assist the conventional correlation filter for the spatial regularisation. By overcoming the feature independence of current regularisation in another tracker, the GANs’ mask can be effectively used to identify the robust features for the target variations representation in the temporal domain. Also in the spatial domain, the background features can be substantially suppressed to obtain the optimisation filter for more reliable matching and updating. In verification, the authors evaluate the proposed tracker on the standard tracking benchmarks, and the experimental results show that their tracker outperforms favourably against other state‐of‐the‐art trackers in the measurements of accuracy and robustness.
Hanqiao Huang, Yufei Zha, Meiyun Zheng, Peng Zhang 0005
IET Image Process.4
2019 Robust Hyperspectral Image Domain Adaptation With Noisy Labels
abstract
In hyperspectral image (HSI) classification, domain adaptation (DA) methods have been proved effective to address unsatisfactory classification results caused by the distribution difference between training (i.e., source domain) and testing (i.e., target domain) pixels. However, these methods rely on accurate labels in source domain, and seldom consider the performance drop resulted by noisy label, which often happens since labeling pixel in HSI is a challenging task. To improve the robustness of DA method to label noise, we propose a new unsupervised HSI DA method, which is constructed from both feature-level and classifier-level. First, a linear transformation function is learned in feature-level to align the source (domain) subspace with the target (domain) subspace. Then, a robust low-rank representation based classifier is developed to well cope with the features obtained from the aligned subspace. Since both subspace alignment and the classifier are immune to noisy labels, the proposed method obtains good classification results when confronting with noisy labels in source domain. Experimental results on two DA benchmarks demonstrate the effectiveness of the proposed method.
Wei Wei 0008, Wei Li 0219, Lei Zhang 0054, Cong Wang 0013, Peng Zhang 0005, Yanning Zhang 0001
IEEE Geosci. Remote. Sens. Lett.5
2019 Arterial Spin Labeling Images Synthesis From sMRI Using Unbalanced Deep Discriminant Learning
abstract
Adequate medical images are often indispensable in contemporary deep learning-based medical imaging studies, although the acquisition of certain image modalities may be limited due to several issues including high costs and patients issues. However, thanks to recent advances in deep learning techniques, the above tough problem can be substantially alleviated by medical images synthesis, by which various modalities including T1/T2/DTI MRI images, PET images, cardiac ultrasound images, retinal images, and so on, have already been synthesized. Unfortunately, the arterial spin labeling (ASL) image, which is an important fMRI indicator in dementia diseases diagnosis nowadays, has never been comprehensively investigated for the synthesis purpose yet. In this paper, ASL images have been successfully synthesized from structural magnetic resonance images for the first time. Technically, a novel unbalanced deep discriminant learning-based model equipped with new ResNet sub-structures is proposed to realize the synthesis of ASL images from structural magnetic resonance images. The extensive experiments have been conducted. Comprehensive statistical analyses reveal that: 1) this newly introduced model is capable to synthesize ASL images that are similar towards real ones acquired by actual scanning; 2) synthesized ASL images obtained by the new model have demonstrated outstanding performance when undergoing rigorous tests of region-based and voxel-based corrections of partial volume effects, which are essential in ASL images processing; and 3) it is also promising that the diagnosis performance of dementia diseases can be significantly improved with the help of synthesized ASL images obtained by the new model, based on a multi-modal MRI dataset containing 355 demented patients in this paper.
Wei Huang 0013, Mingyuan Luo, Xi Liu 0008, Peng Zhang 0005, Huijun Ding, Wufeng Xue, Dong Ni 0001
IEEE Trans. Medical Imaging4
2018 Pixel-wise partial volume effects correction on arterial spin labeling magnetic resonance images
Wei Huang 0013, Chuyu Wan, Huijun Ding, Peng Zhang 0005, Guang Chen 0004
Multim. Tools Appl.5
2018 Single-target localization in video sequences using offline deep-ranked metric learning and online learned models updating
Wei Huang 0013, Peng Zhang 0005, Guang Chen 0004, Huijun Ding
Multim. Tools Appl.3
2018 Robust tracking based on H-CNN with low-resource sampling and scaling by frame-wise motion localization
Peng Zhang 0005, Tao Zhuo, Hanqiao Huang, Kangli Chen, Mohan Kankanhalli
Multim. Tools Appl.1
2018 Going deeper with two-stream ConvNets for action recognition in video surveillance
Peng Zhang 0005, Tao Zhuo, Wei Huang 0013, Yanning Zhang 0001
Pattern Recognit. Lett.2
2018 Saliency flow based video segmentation via motion guided contour refinement
Peng Zhang 0005, Tao Zhuo, Hanqiao Huang, Mohan Kankanhalli
Signal Process.1
2017 Online object tracking based on CNN with spatial-temporal saliency guided sampling
Peng Zhang 0005, Tao Zhuo, Wei Huang 0013, Kangli Chen, Mohan Kankanhalli
Neurocomputing1
2016 Deformable object tracking with spatiotemporal segmentation in big vision surveillance
Peng Zhang 0005, Tao Zhuo, Lei Xie 0001, Yanning Zhang 0001
Neurocomputing1
2016 Online tracking based on efficient transductive learning with sample matching costs
Peng Zhang 0005, Tao Zhuo, Yanning Zhang 0001, Dapeng Tao, Jun Cheng 0002
Neurocomputing1
2016 A novel dementia diagnosis strategy on arterial spin labeling magnetic resonance images via pixel-wise partial volume correction and ranking
Wei Huang 0013, Peng Zhang 0005, Minmin Shen
Multim. Tools Appl.2
2016 Guest Editorial: Immersive Audio/Visual Systems
Lei Xie 0001, Longbiao Wang, Janne Heikkilä, Peng Zhang 0005
Multim. Tools Appl.4
2016 Bayesian tracking fusion framework with online classifier ensemble for immersive visual applications
Peng Zhang 0005, Tao Zhuo, Yanning Zhang 0001, Hanqiao Huang, Kangli Chen
Multim. Tools Appl.1
2016 Real-time tracking-by-learning with high-order regularization fusion for big video abstraction
Peng Zhang 0005, Tao Zhuo, Yanning Zhang 0001, Lei Xie 0001, Dapeng Tao
Signal Process.1
2015 Superframe segmentation based on content-motion correspondence for social video summarization
abstract
The goal of video summarization is to turn large volume of video data into a compact visual summary that can be easily interpreted by users in a while. Existing summarization strategies employed the point based feature correspondence for the superframe segmentation. Unfortunately, the information carried by those sparse points is far from sufficiency and stability to describe the change of interesting regions of each frame. Therefore, in order to overcome the limitations of point feature, we propose a region correspondence based superframe segmentation to achieve more effective video summarization. Instead of utilizing the motion of feature points, we calculate the similarity of content-motion to obtain the strength of change between the consecutive frames. With the help of circulant structure kernel, the proposed method is able to perform more accurate motion estimation efficiently. Experimental testing on the videos from benchmark database has demonstrate the effectiveness of the proposed method.
Tao Zhuo, Peng Zhang 0005, Kangli Chen, Yanning Zhang 0001
ACII2
2015 Online Object Tracking Based on CNN with Metropolis-Hasting Re-Sampling
abstract
Tracking-by-learning strategies have been effective in solving many challenging problems in visual tracking, in which the learning sample generation and labeling play important roles for final performance. Since the concern of deep learning based approaches has shown an impressive performance in different vision tasks, how to properly apply the learning model, such as CNN, to an online tracking framework is still challenging. In this paper, to overcome the overfitting problem caused by straight-forward incorporation, we propose an online tracking framework by constructing a CNN based adaptive appearance model to generate more reliable training data over time. With a reformative Metropolis-Hastings re-sampling scheme to reshape particles for a better state posterior representation during online learning, the proposed tracking outperforms most of the state-of-art trackers on challenging benchmark video sequences.
Xiangzeng Zhou, Lei Xie 0001, Peng Zhang 0005, Yanning Zhang 0001
ACM Multimedia3
2015 Empirical mode decomposition based blind audio watermarking
Zhaoyang Fu, Peng Zhang 0005, Wei Huang 0013, Liang Wang 0001, Sabu Emmanuel, Guang Chen 0004
Multim. Tools Appl.2
2015 A novel marker-less lung tumor localization strategy on low-rank fluoroscopic images with similarity learning
Wei Huang 0013, Jing Li 0027, Peng Zhang 0005, Can Fang, Minmin Shen
Multim. Tools Appl.3
2015 Multiple pedestrian tracking based on couple-states Markov chain with semantic topic learning for video surveillance
Peng Zhang 0005, Liang Wang 0001, Wei Huang 0013, Lei Xie 0001, Guang Chen 0004
Soft Comput.1
2014 An ensemble of deep neural networks for object tracking
abstract
Object tracking in complex backgrounds with dramatic appearance variations is a challenging problem in computer vision. We tackle this problem by a novel approach that incorporates a deep learning architecture with an on-line AdaBoost framework. Inspired by its multi-level feature learning ability, a stacked denoising autoencoder (SDAE) is used to learn multi-level feature descriptors from a set of auxiliary images. Each layer of the SDAE, representing a different feature space, is subsequently transformed to a discriminative object/background deep neural network (DNN) classifier by adding a classification layer. By an on-line AdaBoost feature selection framework, the ensemble of the DNN classifiers is then updated on-line to robustly distinguish the target from the background. Experiments on an open tracking benchmark show promising results of the proposed tracker as compared with several state-of-the-art approaches.
Xiangzeng Zhou, Lei Xie 0001, Peng Zhang 0005, Yanning Zhang 0001
ICIP3
2014 Object Tracking using Reformative Transductive Learning with Sample Variational Correspondence
abstract
Tracking-by-learning strategies have effectively solved many challenging problems for visual tracking. When labeled samples are limited, the learning performance can be improved by exploiting unlabeled ones. Thus, a key issue for semi-supervised learning is the label assignment of the unlabeled samples, which is the principal focus of transductive learning. Unfortunately, the optimization scheme employed by the transductive learning is hard to be applied to online tracking because of its large amount of computation for sample labeling. In this paper, a reformative transductive learning was proposed with the variational correspondence between the learning samples, which are utilized to build an effective matching cost function for more efficient label assignment during the learning of representative separators. By using a weighted accumulative average to update the coefficients via a fixed budget of support vectors, the proposed tracking has been demonstrated to outperform most of the state-of-art trackers.
Tao Zhuo, Peng Zhang 0005, Yanning Zhang 0001, Wei Huang 0013, Hichem Sahli
ACM Multimedia2
2014 Coverage enhancement by using the mobility of mobile sensor nodes
Can Fang, Peng Zhang 0005, Zili Zhang 0001
Multim. Tools Appl.2
2014 Moving people tracking with detection by latent semantic analysis for visual surveillance applications
Peng Zhang 0005, Yanning Zhang 0001, Tony Thomas, Sabu Emmanuel
Multim. Tools Appl.1
2013 A novel marker-less tumor tracking strategyonlow-rank fluoroscopic images for image-guided lung cancer radiotherapy
abstract
Fluoroscopic images recording the real-time motion of lung tumor lesion play an important role on lung cancer radiotherapy, as these images help to facilitate the accurate delivery of radiation dose on target tumor lesion. Derivation of tumor position in conventional lung tumor tracking strategies is realized via either placing external surrogates on patients or implanting internal fiducial markers in patients. Inaccurate tumor tracking and patient safety problems are often inevitable for these strategies. In this study, a novel marker-less tumor tracking strategy is presented for image-guided lung cancer radiotherapy. A fluoroscopic image is first decomposed into low-rank and sparse components based on robust-PCA via a split Bregman method. Then, a series of techniques, including K-means clustering, morphological processing, connected component analysis, etc are employed on obtained low-rank fluoroscopic images for tumor tracking. Clinical data obtained from 45 patients is incorporated for experimental evaluation. Promising results are demonstrated from the introduced strategy.
Wei Huang 0013, Jing Li 0027, Peng Zhang 0005
ICIP3
2013 Non-rigid target tracking based on 'flow-cut' in pair-wise frames with online hough forests
abstract
In conventional online learning based tracking studies, fixed-shape appearance modeling is often incorporated for training samples generation, as it is simple and convenient to be applied. However, for more general non-rigid and articulated object, this strategy may regard some background areas as foreground, which is likely to deteriorate the learning process. Recently published works utilize more than one patches to represent non-rigid object with foreground object segmentation, but most of these segmentation for target representation are performed only in single frame manner. Since the motion information between the consecutive frames was not considered by these approaches, when the backgrounds are similar to the target, accurate segmentation is hard to be achieved. In this work, we propose a novel model for non-rigid object segmentation by incorporating consecutive gradients flow between pair-wise frames into a Gibbs energy function. With help from motion information, the irregular target areas can be segmented more accurately during precise boundary convergence. The proposed segmentation model is incorporated into a semi-supervised online tracking framework for training samples generation. We test the proposed tracking on challenging videos involving heavy intrinsic variations and occlusions. As a result, the experiments demonstrate a significant improvement in tracking accuracy and robustness in comparison with other state-of-art tracking works.
Yanning Zhang 0001, Peng Zhang 0005, Wei Huang 0013, Hichem Sahli
ACM Multimedia3
2012 Privacy enabled video surveillance using a two state Markov tracking algorithm
Peng Zhang 0005, Tony Thomas, Sabu Emmanuel
Multim. Syst.1
2011 Pedestrian Tracking Based on Hidden-Latent Temporal Markov Chain
Peng Zhang 0005, Sabu Emmanuel, Mohan Kankanhalli
MMM (2)1
2010 An Authentication Mechanism Using Chinese Remainder Theorem for Efficient Surveillance Video Transmission
abstract
Now-a-days, surveillance cameras have been widely deployed in various security applications. In many surveillance applications, the background changes very slowly and the foreground objects occupy only a relatively small portion of a video frame. In these type of applications, an efficient solution for transmissions over bandwidth-limited networks is to send only the foreground objects for every frame in real time while the background is sent occasionally. At the receiving end of the transmission, the objects and the most recent background can be fused together and the original frame can be reconstructed. However, protecting the authenticity of the video becomes more challenging in this case as a malicious entity can modify/replace/remove the individual foreground objects and background in the video. In this paper, we propose a Chinese remainder theorem based watermarking mechanism for protecting the authenticity of videos transmitted or stored as objects and background. Our mechanism ensures the authenticity between video objects and their associated background.
Tony Thomas, Sabu Emmanuel, Peng Zhang 0005, Mohan Kankanhalli
AVSS3
2009 Auto-scaled Incremental Tensor Subspace Learning for Region Based Rate Control Application
Peng Zhang 0005, Sabu Emmanuel, Yanning Zhang 0001, Xuan Jing
ACCV (3)1
2009 Spatiotemporal latent semantic cues for moving people tracking
abstract
Effective and robust visual tracking is one of the most important tasks for the intelligent visual surveillance. In this paper, we proposed a novel method for detecting and tracking moving people using the spatiotemporal latent semantic cues and the incremental eigenspace tracking techniques. During tracking process, the target appearance model is incrementally learned in low dimensional tensor eigenspace by adaptively updating the eigenbasis and sample mean. At the same time, the spatiotemporal latent semantic cues calibrate the estimation of tracking and detect new moving people coming in the same surveillance scene. Experiment results show that with the calibration based on spatiotemporal latent semantic cues, the proposed method can track the moving people automatically and effectively.
Peng Zhang 0005, Sabu Emmanuel, Pradeep K. Atrey, Mohan Kankanhalli
ICASSP1
2008 Zoomed Object Segmentation from Dynamic Scene Containing a Door
abstract
Accurately segmenting the moving objects from a sequence of captured video frames is a significant pre-condition for tracking and recognition of these moving objects. The challenge of segmenting the moving object is even harder when the background is dynamic and the camera used can change its zoom dynamically. Here, in this paper, we propose a new method to detect and segment moving object from a dynamic background, which contains moving multiple-leaf doors. In addition the proposed algorithm also takes care of dynamic zoom changes that can occur while shooting a scene. The proposed algorithm uses background-rebuilding with discrete door's position to tackle moving multiple-leaf door backgrounds and image feature comparison to tackle changes in zoom. We have obtained good image sequence segmentation results with high processing speed.
Peng Zhang 0005, Sabu Emmanuel
HPCC1
2008 Spatial and temporal sampling control for visual surveillance application
abstract
In this paper a novel way to control the amount of generated video surveillance data by controlling the spatial and temporal samplings of video is proposed. The samplings are controlled adaptively using the speed, distance and dimension of the object extracted dynamically from the surveillance video. We also present a method of estimating the actual 2D dimension, location and speed of moving objects especially from surveillance video of indoor environments such as indoor car parks, corridors of buildings, shopping malls etc. Experiments were conducted to study the generated data size reduction and the reduction is found to be substantial and also to find the accuracy of the estimated 2D dimension, location and speed of the moving object and the accuracy is found to be high.
Sabu Emmanuel, Peng Zhang 0005, Agus Sugama
SMC2