EDBT 2026 Demo / reviewers in the wild / expert
Xiaotang Chen
dblp:22/10700
· DBLP profile ↗
28ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0002-3362-1431ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual CuesabstractVision-Language Tracking (VLT) aims to localize a target in video sequences using a visual template and language description. While textual cues enhance tracking potential, current datasets typically contain much more image data than text, limiting the ability of VLT methods to align the two modalities effectively. To address this imbalance, we propose a novel plug-and-play method named CTVLT that leverages the strong text-image alignment capabilities of foundation grounding models. CTVLT converts textual cues into interpretable visual heatmaps, which are easier for trackers to process. Specifically, we design a textual cue mapping module that transforms textual cues into target distribution heatmaps, visually representing the location described by the text. Additionally, the heatmap guidance module fuses these heatmaps with the search image to guide tracking more effectively. Extensive experiments on mainstream benchmarks demonstrate the effectiveness of our approach, achieving state-of-the-art performance and validating the utility of our method for enhanced VLT. Xiaokun Feng, Dailing Zhang, Xuchen Li 0001, Meiqi Wu, Jing Zhang 0110, Xiaotang Chen, Kaiqi Huang |
ICASSP | 7 |
| 2025 | ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language TrackingabstractVision-language tracking aims to locate the target object in the video sequence using a template patch and a language description provided in the initial frame. To achieve robust tracking, especially in complex long-term scenarios that reflect real-world conditions as recently highlighted by MGIT, it is essential not only to characterize the target features but also to utilize the context features related to the target. However, the visual and textual target-context cues derived from the initial prompts generally align only with the initial target state. Due to their dynamic nature, target states are constantly changing, particularly in complex long-term sequences. It is intractable for these cues to continuously guide Vision-Language Trackers (VLTs). Furthermore, for the text prompts with diverse expressions, our experiments reveal that existing VLTs struggle to discern which words pertain to the target or the context, complicating the utilization of textual cues. In this work, we present a novel tracker named ATCTrack, which can obtain multimodal cues Aligned with the dynamic target states through comprehensive Target-Context feature modeling, thereby achieving robust tracking. Specifically, (1) for the visual modality, we propose an effective temporal visual target-context modeling approach that provides the tracker with timely visual cues. (2) For the textual modality, we achieve precise target words identification solely based on textual content, and design an innovative context words calibration method to adaptively utilize auxiliary context words. (3) We conduct extensive experiments on mainstream benchmarks and ATCTrack achieves a new SOTA performance. The code and models will be released at: https://github.com/XiaokunFeng/ATCTrack. Xiaokun Feng, Xuchen Li 0001, Dailing Zhang, Meiqi Wu, Jing Zhang 0110, Xiaotang Chen, Kaiqi Huang |
ICCV | 7 |
| 2025 | CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal FeaturesabstractEffectively modeling and utilizing spatiotemporal features from RGB and other modalities (e.g., depth, thermal, and event data, denoted as X) is the core of RGB-X tracker design. Existing methods often employ two parallel branches to separately process the RGB and X input streams, requiring the model to simultaneously handle two dispersed feature spaces, which complicates both the model structure and computation process. More critically, intra-modality spatial modeling within each dispersed space incurs substantial computational overhead, limiting resources for inter-modality spatial modeling and temporal modeling. To address this, we propose a novel tracker, CSTrack, which focuses on modeling Compact Spatiotemporal features to achieve simple yet effective tracking. Specifically, we first introduce an innovative Spatial Compact Module that integrates the RGB-X dual input streams into a compact spatial feature, enabling thorough intra- and inter-modality spatial modeling. Additionally, we design an efficient Temporal Compact Module that compactly represents temporal features by constructing the refined target distribution heatmap. Extensive experiments validate the effectiveness of our compact spatiotemporal modeling method, with CSTrack achieving new SOTA results on mainstream RGB-X benchmarks. The code and models will be released at: https://github.com/XiaokunFeng/CSTrack. Xiaokun Feng, Dailing Zhang, Xuchen Li 0001, Meiqi Wu, Jing Zhang 0110, Xiaotang Chen, Kaiqi Huang |
ICML | 7 |
| 2025 | Uncertainty-Aware Opponent Modeling for Deep Reinforcement Learning
Likun Yang, Pei Xu 0003, Shiyue Cao, Xiaotang Chen, Kaiqi Huang |
AAMAS | 5 |
| 2024 | MemVLT: Vision-Language Tracking with Adaptive Memory-based PromptsabstractVision-language tracking (VLT) enhances traditional visual object tracking by integrating language descriptions, requiring the tracker to flexibly understand complex and diverse text in addition to visual information. However, most existing vision-language trackers still overly rely on initial fixed multimodal prompts, which struggle to provide effective guidance for dynamically changing targets. Fortunately, the Complementary Learning Systems (CLS) theory suggests that the human memory system can dynamically store and utilize multimodal perceptual information, thereby adapting to new scenarios. Inspired by this, (i) we propose a Memory-based Vision-Language Tracker (MemVLT). By incorporating memory modeling to adjust static prompts, our approach can provide adaptive prompts for tracking guidance.
(ii) Specifically, the memory storage and memory interaction modules are designed in accordance with CLS theory. These modules facilitate the storage and flexible interaction between short-term and long-term memories, generating prompts that adapt to target variations.
(iii) Finally, we conduct extensive experiments on mainstream VLT datasets (e.g., MGIT, TNL2K, LaSOT and LaSOT$_{ext}$). Experimental results show that MemVLT achieves new state-of-the-art performance. Impressively, it achieves 69.4% AUC on the MGIT and 63.3% AUC on the TNL2K, improving the existing best result by 8.4% and 4.7%, respectively. Xiaokun Feng, Xuchen Li 0001, Dailing Zhang, Meiqi Wu, Jing Zhang 0110, Xiaotang Chen, Kaiqi Huang |
NeurIPS | 7 |
| 2024 | VS-LLM: Visual-Semantic Depression Assessment Based on LLM for Drawing Projection Test
Meiqi Wu, Yaxuan Kang, Xuchen Li 0001, Xiaotang Chen, Yunfeng Kang, Weiqiang Wang 0001, Kaiqi Huang |
PRCV (9) | 5 |
| 2023 | A Hierarchical Theme Recognition Model for Sandplay Therapy
Xiaokun Feng, Xiaotang Chen, Kaiqi Huang |
PRCV (4) | 3 |
| 2023 | EKGRL: Entity-Based Knowledge Graph Representation Learning for Fact-Based Visual Question Answering
Xiaotang Chen, Kaiqi Huang |
PRCV (6) | 2 |
| 2022 | Learning Disentangled Attribute Representations for Robust Pedestrian Attribute RecognitionabstractAlthough various methods have been proposed for pedestrian attribute recognition, most studies follow the same feature learning mechanism, \ie, learning a shared pedestrian image feature to classify multiple attributes. However, this mechanism leads to low-confidence predictions and non-robustness of the model in the inference stage. In this paper, we investigate why this is the case. We mathematically discover that the central cause is that the optimal shared feature cannot maintain high similarities with multiple classifiers simultaneously in the context of minimizing classification loss. In addition, this feature learning mechanism ignores the spatial and semantic distinctions between different attributes. To address these limitations, we propose a novel disentangled attribute feature learning (DAFL) framework to learn a disentangled feature for each attribute, which exploits the semantic and spatial characteristics of attributes. The framework mainly consists of learnable semantic queries, a cascaded semantic-spatial cross-attention (SSCA) module, and a group attention merging (GAM) module. Specifically, based on learnable semantic queries, the cascaded SSCA module iteratively enhances the spatial localization of attribute-related regions and aggregates region features into multiple disentangled attribute features, used for classification and updating learnable semantic queries. The GAM module splits attributes into groups based on spatial distribution and utilizes reliable group attention to supervise query attention maps. Experiments on PETA, RAPv1, PA100k, and RAPv2 show that the proposed method performs favorably against state-of-the-art methods. Jian Jia, Naiyu Gao, Xiaotang Chen, Kaiqi Huang |
AAAI | 4 |
| 2022 | Bottom-Up Foreground-Aware Feature Fusion for Practical Person SearchabstractThe key to efficient person search is jointly localizing pedestrians and learning discriminative representation for person re-identification (re-ID). Some recently developed models are built with separate detection and re-ID branches on top of shared region feature extraction networks. There are two factors that are detrimental to re-ID feature learning. One is the background information redundancy resulting from the large receptive field of neurons. The other is the body part missing and background clutter caused by inaccurate localization. In this work, a bottom-up fusion (BUF) subnet is proposed to fuse the bounding box features pooled from multiple network stages. With a few parameters introduced, BUF leverages the multi-level features with various sizes of receptive fields to mitigate the background-bias problem. To further suppress the non-pedestrian regions, the newly introduced segmentation head generates a foreground probability map as guidance for the network to focus on the foreground regions. The resulting foreground attention module (FAM) enhances the foreground features. Moreover, for robust feature learning in practical person search, we propose to adaptively smooth the labels of the pedestrian boxes with consideration of the detection quality. Extensive experiments on PRW and CUHK-SYSU validate the effectiveness of the proposals. Our Bottom-Up Foreground-Aware Feature Fusion (BUFF) network with ALS achieves considerable gains over the state-of-the-art on PRW and competitive performance on CUHK-SYSU. Wenjie Yang 0005, Houjing Huang, Xiaotang Chen, Kaiqi Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Spatial and Semantic Consistency Regularizations for Pedestrian Attribute RecognitionabstractWhile recent studies on pedestrian attribute recognition have shown remarkable progress in leveraging complicated networks and attention mechanisms, most of them neglect the inter-image relations and an important prior: spatial consistency and semantic consistency of attributes under surveillance scenarios. The spatial locations of the same attribute should be consistent between different pedestrian images, e.g., the "hat" attribute and the "boots" attribute are always located at the top and bottom of the picture respectively. In addition, the inherent semantic feature of the "hat" attribute should be consistent, whether it is a baseball cap, beret, or helmet. To fully exploit inter-image relations and aggregate human prior in the model learning process, we construct a Spatial and Semantic Consistency (SSC) framework that consists of two complementary regularizations to achieve spatial and semantic consistency for each attribute. Specifically, we first propose a spatial consistency regularization to focus on reliable and stable attribute-related regions. Based on the precise attribute locations, we further propose a semantic consistency regularization to extract intrinsic and discriminative semantic features. We conduct extensive experiments on popular benchmarks including PA100K, RAP, and PETA. Results show that the proposed method performs favorably against state- of-the-art methods without increasing parameters. Jian Jia, Xiaotang Chen, Kaiqi Huang |
ICCV | 2 |
| 2020 | Human Parsing Based Alignment With Multi-Task Learning For Occluded Person Re-IdentificationabstractPerson re-identification (ReID) has obtained great progress in recent years. However, the problem caused by occlusion, which is frequent under surveillance camera, is not sufficiently addressed. When human body is occluded, extracted features are flooded with background noise. Moreover, without knowing location and visibility of parts, directly matching partial images with others will cause misalignment. To tackle the issue, we propose a model named HPNet to extract part-level features and predict visibility of each part, based on human parsing. By extracting features from semantic part regions and perform comparison with consideration of visibility, our method not only reduces background noise but also achieves alignment. Furthermore, ReID and human parsing are learned in a multi-task manner, without the need for an extra part model during testing. In addition to being efficient, the performance of our model surpasses previous methods by a large margin under occlusion scenarios. Houjing Huang, Xiaotang Chen, Kaiqi Huang |
ICME | 2 |
| 2020 | Proxy Task Learning For Cross-Domain Person Re-IdentificationabstractPerson re-identification (ReID) has achieved rapid improvement recently. However, exploiting the model in a new scene is always faced with huge performance drop. The cause lies in distribution discrepancy between domains, including both low-level (e.g. image quality) and high-level (e.g. pedestrian attribute) variance. To alleviate the problem of domain shift, we propose a novel framework Proxy Task Learning (PTL), which performs body perception tasks on target-domain images while training source-domain ReID, in a multi-task manner. The backbone is shared between tasks and domains, hence both low- and high-level distributions are deeply aligned. We experimentally verify two proxy tasks, i.e. human parsing and attribute recognition, that prominently enhance generalization of the model. When integrating our method into an existing cross-domain pipeline, we achieve state-of-the-art performance on large-scale benchmarks. Houjing Huang, Xiaotang Chen, Kaiqi Huang |
ICME | 2 |
| 2020 | Bottom-Up Foreground-Aware Feature Fusion for Person SearchabstractThe key to efficient person search is jointly localizing pedestrians and learning discriminative representation for person re-identification (re-ID). Some recently developed task-joint models are built with separate detection and re-ID branches on top of shared region feature extraction networks, where the large receptive field of neurons leads to background information redundancy for the following re-ID task. Our diagnostic analysis indicates the task-joint model suffers from considerable performance drop when the background is replaced or removed. In this work, we propose a subnet to fuse the bounding box features that pooled from multiple ConvNet stages in a bottom-up manner, termed bottom-up fusion (BUF) network. With a few parameters introduced, BUF leverages the multi-level features with different sizes of receptive fields to mitigate the background-bias problem. Moreover, the newly introduced segmentation head generates a foreground probability map as guidance for the network to focus on the foreground regions. The resulting foreground attention module (FAM) enhances the foreground features. Extensive experiments on PRW and CUHK-SYSU validate the effectiveness of the proposals. Our Bottom-Up Foreground-Aware Feature Fusion (BUFF) network achieves considerable gains over the state-of-the- arts on PRW and competitive performance on CUHK-SYSU. Wenjie Yang 0005, Dangwei Li, Xiaotang Chen, Kaiqi Huang |
ACM Multimedia | 3 |
| 2020 | Improve Person Re-Identification With Part Awareness LearningabstractPerson re-identification (ReID) aims to predict whether two images from different cameras belong to the same person. Due to low image quality and variance in view point and body pose, it remains a difficult task. To solve the task, a model is supposed to appropriately capture features that describe body regions for identification. With the simple intuition that explicitly incorporating ReID model with part awareness could be beneficial for learning a more discriminative feature space, we propose part segmentation as an assistant body perception task during the training of a ReID model. Specifically, we add a lightweight segmentation head to the backbone of ReID model during training, which is supervised with part labels. Note that our segmentation head is only introduced during training and that it does not change network input or the way of extracting ReID feature. Experiments show that part segmentation considerably improves the performance of ReID. Through quantitative and qualitative analyses, we further reveal that body part perception helps ReID model to capture a set of more diverse features from the body, with decreased similarity between part features and increased focus on different body regions. We experiment with various representative ReID models and achieve consistent improvement on several large-scale datasets including Market1501, CUHK03, DukeMTMC-reID and MSMT17. E.g. on MSMT17, our method increases Rank-1 Accuracy of GlobalPool-ResNet-50, PCB and MGN by 2.3%, 2.9% and 3.9%, respectively. Incorporated with MGN, our model achieves state-of-the-art performance, with Rank-1 Accuracy 95.8%, 78.8%, 90.0% and 84.0% on four datasets, respectively. Houjing Huang, Wenjie Yang 0005, Jinbin Lin, Guan Huang 0003, Jiamiao Xu, Xiaotang Chen, Kaiqi Huang |
IEEE Trans. Image Process. | 7 |
| 2019 | Towards Rich Feature Discovery With Class Activation Maps Augmentation for Person Re-IdentificationabstractThe fundamental challenge of small inter-person variation requires Person Re-Identification (Re-ID) models to capture sufficient fine-grained information. This paper proposes to discover diverse discriminative visual cues without extra assistance, e.g., pose estimation, human parsing. Specifically, a Class Activation Maps (CAM) augmentation model is proposed to expand the activation scope of baseline Re-ID model to explore rich visual cues, where the backbone network is extended by a series of ordered branches which share the same input but output complementary CAM. A novel Overlapped Activation Penalty is proposed to force the new branch to pay more attention to the image regions less activated by the old ones, such that spatial diverse visual features can be discovered. The proposed model achieves state-of-the-art results on three person Re-ID benchmarks. Moreover, a visualization approach termed ranking activation map (RAM) is proposed to explicitly interpret the ranking results in the test stage, which gives qualitative validations of the proposed method. Wenjie Yang 0005, Houjing Huang, Zhang Zhang 0001, Xiaotang Chen, Kaiqi Huang, Shu Zhang 0001 |
CVPR | 4 |
| 2019 | An Effective Adversarial Training Based Spatial-Temporal Network for Abnormal Behavior DetectionabstractUnsupervised abnormal behavior detection has attracted much attention in recent years. It is a challenging task due to the undefinition and the sparsity of abnormal behaviors, etc. Existing generative model based methods usually perform poorly due to the unknown types of abnormal behaviors and insufficient exploiting of spatial-temporal information. In this paper, we propose a novel adversarial training based spatial-temporal network to tackle these problems. Firstly, we introduce an adversarial training strategy to deal with unknown types of abnormal behaviors. Secondly, to better explore spatial-temporal information, we design an effective two-stream spatial-temporal network, which is identity mapping free and spatial-temporal complementary. Finally, we combine them together to get the final adversarial spatial-temporal network. Our method is evaluated on various challenging public datasets and achieves the state-of-the-art performance. Zhiyu Yin, Xiaotang Chen, Kaiqi Huang |
ICIP | 2 |
| 2019 | A Richly Annotated Pedestrian Dataset for Person Retrieval in Real Surveillance ScenariosabstractRetrieving specific persons with various types of queries, e.g., a set of attributes or a portrait photo has great application potential in large-scale intelligent surveillance systems. In this paper, we propose a richly annotated pedestrian (RAP) dataset which serves as a unified benchmark for both attribute-based and image-based person retrieval in real surveillance scenarios. Typically, previous datasets have three improvable aspects, including limited data scale and annotation types, heterogeneous data source, and controlled scenarios. Differently, RAP is a large-scale dataset which contains 84928 images with 72 types of attributes and additional tags of viewpoint, occlusion, body parts, and 2589 person identities. It is collected in the real uncontrolled scene and has complex visual variations in pedestrian samples due to the change of viewpoints, pedestrian postures, and cloth appearance. Towards a high-quality person retrieval benchmark, an amount of state-of-the-art algorithms on pedestrian attribute recognition and person re-identification (ReID), are performed for quantitative analysis with three evaluation tasks, i.e., attribute recognition, attribute-based and image-based person retrieval, where a new instance-based metric is proposed to measure the dependency of the prediction of multiple attributes. Finally, some interesting problems, e.g., the joint feature learning of attribute recognition and ReID, and the problem of cross-day person ReID, are explored to show the challenges and future directions in person retrieval. Dangwei Li, Zhang Zhang 0001, Xiaotang Chen, Kaiqi Huang |
IEEE Trans. Image Process. | 3 |
| 2018 | Adversarially Occluded Samples for Person Re-IdentificationabstractPerson re-identification (ReID) is the task of retrieving particular persons across different cameras. Despite its great progress in recent years, it is still confronted with challenges like pose variation, occlusion, and similar appearance among different persons. The large gap between training and testing performance with existing models implies the insufficiency of generalization. Considering this fact, we propose to augment the variation of training data by introducing Adversarially Occluded Samples. These special samples are both a) meaningful in that they resemble real-scene occlusions, and b) effective in that they are tough for the original model and thus provide the momentum to jump out of local optimum. We mine these samples based on a trained ReID model and with the help of network visualization techniques. Extensive experiments show that the proposed samples help the model discover new discriminative clues on the body and generalize much better at test time. Our strategy makes significant improvement over strong baselines on three large-scale ReID datasets, Market1501, CUHK03 and DukeMTMC-reID. Houjing Huang, Dangwei Li, Zhang Zhang 0001, Xiaotang Chen, Kaiqi Huang |
CVPR | 4 |
| 2018 | Pose Guided Deep Model for Pedestrian Attribute Recognition in Surveillance ScenariosabstractRecognizing pedestrian attributes, such as gender, backpack, and cloth types, has obtained increasing attention recently due to its great potential in intelligent video surveillance. Existing methods usually solve it with end-to-end multi-label deep neural networks, while the structure knowledge of pedestrian body has been little utilized. Considering that attributes have strong spatial correlations with human structures, e.g. glasses are around the head, in this paper, we introduce pedestrian body structure into this task and propose a Pose Guided Deep Model (PGDM) to improve attribute recognition. The PGDM consists of three main components: 1) coarse pose estimation which distillates the pose knowledge from a pre-trained pose estimation model, 2) body parts localization which adaptively locates informative image regions with only image-level supervision, 3) multiple features fusion which combines the part-based features for attribute recognition. In the inference stage, we fuse the part-based PGDM results with global body based results for final attribute prediction and the performance can be consistently improved. Compared with state-of-the-art models, the performances on three large-scale pedestrian attribute datasets, i.e., PETA, RAP, and PA-100K, demonstrate the effectiveness of the proposed method. Dangwei Li, Xiaotang Chen, Zhang Zhang 0001, Kaiqi Huang |
ICME | 2 |
| 2017 | A Multi-Task Deep Network for Person Re-IdentificationabstractPerson re-identification (ReID) focuses on identifying people across different scenes in video surveillance, which is usually formulated as a binary classification task or a ranking task in current person ReID approaches. In this paper, we take both tasks into account and propose a multi-task deep network (MTDnet) that makes use of their own advantages and jointly optimize the two tasks simultaneously for person ReID. To the best of our knowledge, we are the first to integrate both tasks in one network to solve the person ReID. We show that our proposed architecture significantly boosts the performance. Furthermore, deep architecture in general requires a sufficient dataset for training, which is usually not met in person ReID. To cope with this situation, we further extend the MTDnet and propose a cross-domain architecture that is capable of using an auxiliary set to assist training on small target sets. In the experiments, our approach outperforms most of existing person ReID algorithms on representative datasets including CUHK03, CUHK01, VIPeR, iLIDS and PRID2011, which clearly demonstrates the effectiveness of the proposed approach. Xiaotang Chen, Jianguo Zhang 0001, Kaiqi Huang |
AAAI | 2 |
| 2017 | Beyond Triplet Loss: A Deep Quadruplet Network for Person Re-identificationabstractPerson re-identification (ReID) is an important task in wide area video surveillance which focuses on identifying people across different cameras. Recently, deep learning networks with a triplet loss become a common framework for person ReID. However, the triplet loss pays main attentions on obtaining correct orders on the training set. It still suffers from a weaker generalization capability from the training set to the testing set, thus resulting in inferior performance. In this paper, we design a quadruplet loss, which can lead to the model output with a larger inter-class variation and a smaller intra-class variation compared to the triplet loss. As a result, our model has a better generalization ability and can achieve a higher performance on the testing set. In particular, a quadruplet deep network using a margin-based online hard negative mining is proposed based on the quadruplet loss for the person ReID. In extensive experiments, the proposed network outperforms most of the state-of-the-art algorithms on representative datasets which clearly demonstrates the effectiveness of our proposed method. Xiaotang Chen, Jianguo Zhang 0001, Kaiqi Huang |
CVPR | 2 |
| 2017 | Learning Deep Context-Aware Features over Body and Latent Parts for Person Re-identificationabstractPerson Re-identification (ReID) is to identify the same person across different cameras. It is a challenging task due to the large variations in person pose, occlusion, background clutter, etc. How to extract powerful features is a fundamental problem in ReID and is still an open problem today. In this paper, we design a Multi-Scale Context-Aware Network (MSCAN) to learn powerful features over full body and body parts, which can well capture the local context knowledge by stacking multi-scale convolutions in each layer. Moreover, instead of using predefined rigid parts, we propose to learn and localize deformable pedestrian parts using Spatial Transformer Networks (STN) with novel spatial constraints. The learned body parts can release some difficulties, e.g. pose variations and background clutters, in part-based representation. Finally, we integrate the representation learning processes of full body and body parts into a unified framework for person ReID through multi-class person identification tasks. Extensive evaluations on current challenging large-scale person ReID datasets, including the image-based Market1501, CUHK03 and sequence-based MARS datasets, show that the proposed method achieves the state-of-the-art results. Dangwei Li, Xiaotang Chen, Zhang Zhang 0001, Kaiqi Huang |
CVPR | 2 |
| 2017 | An Equalized Global Graph Model-Based Approach for Multicamera Object TrackingabstractNonoverlapping multicamera visual object tracking typically consists of two steps: single-camera object tracking (SCT) and inter-camera object tracking (ICT). Most of tracking methods focus on SCT, which happens in the same scene, while for real surveillance scenes, ICT is needed and single-camera tracking methods cannot work effectively. In this paper, we try to improve the overall multicamera object tracking performance by a global graph model with an improved similarity metric. Our method treats the similarities of single-camera tracking and inter-camera tracking differently and obtains the optimization in a global graph model. The results show that our method can work better even in the condition of poor SCT. Lijun Cao, Xiaotang Chen, Kaiqi Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | A novel solution for multi-camera object trackingabstractThe traditional multi-camera object tracking contains two steps: single camera object tracking (SCT) and inter-camera object tracking (ICT). The ICT performance strongly relies on the great results of SCT. In practice, most of current SCT methods are unperfect and products much more fragments. In this paper, a novel solution using a global tracklet association is proposed, which can provide a good ICT performance when the SCT results are not perfect. The proposed solution is also available in non-overlapping views through a new tracklet representation and experiments shows the effectiveness of the proposed novel solution in real scene. Lijun Cao, Xiaotang Chen, Kaiqi Huang |
ICIP | 3 |
| 2014 | Object tracking across non-overlapping views by learning inter-camera transfer models
Xiaotang Chen, Kaiqi Huang, Tieniu Tan |
Pattern Recognit. | 1 |
| 2011 | Direction-based stochastic matching for pedestrian recognition in non-overlapping camerasabstractPedestrian recognition is a challenging problem in non-overlapping multi-camera object tracking. In this paper, we present a novel approach for matching pedestrians across non-overlapping multiple cameras without the need of a training phase or spatio-temporal cues across cameras. To deal with viewpoint changes, we introduce the concept of directional angles estimated using the spatio-temporal continuity in the single camera tracking. To deal with pose changes, a stochastic matching strategy is performed, where the similarity of two blobs belonging to different viewpoints is calculated by a novel similarity measurement algorithm. The experiments are performed on different multi-view datasets. Experimental results demonstrate the effectiveness and robustness of the proposed method. Xiaotang Chen, Kaiqi Huang, Tieniu Tan |
ICIP | 1 |
| 2001 | Audio compression with desirable attributes for factory automationabstractWith increasing factory automation application requiring multimedia content, compression technology for audiovisual content is needed to minimal bandwidth requirement. An audio compression scheme is presented in here and this scheme is characterized primarily by its fast encoding and decoding speed for high fidelity audio. The bit rate for this scheme is competitive to existing international standard for high fidelity audio compression. In the context of factory automation, the compression scheme presented in here has desirable attribute including preservation of high quantitative fidelity, fast encoding speed using off-the-shelf microcontroller with integer only computation for low cost applications, minimal codec delay to satisfy real time requirement, etc. Xiaotang Chen, Y. L. Yu |
ETFA (2) | 1 |