Xiaolong Zhou 0001

dblp:19/1168-1 · also Xiao-Long Zhou 0001 · DBLP profile ↗
← Back
39ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0003-0732-5169ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 5 since 2021Systems, architecture and hardware · 7 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Prototype-based latent space distance optimization on vehicle re-identification
Sixian Chan 0001, Jiaao Cui, Zheng Wang 0059, Xiaolong Zhou 0001, Xiaoqin Zhang 0002
Expert Syst. Appl.6
2025 SGFormer: Semantic-Geometry Fusion Transformer for Multi-modal 3D Panoptic Segmentation
abstract
Modern methods for autonomous driving perception widely adopt multi-modal fusion to enhance 3D scene understanding. However, existing methods suffer from inferior semantic extraction in image encoders that treat all pixels equally, ignoring contextual differences. The generated multi-modal representations also typically lack comprehensive semantic and spatial geometry information, which is crucial for the 3D panoptic segmentation task. In this paper, we propose a novel Semantic-Geometry Fusion Transformer (SGFormer) that extracts adaptive semantic contexts, aggregates geometric information and captures the semantic-geometry fusion. First, in the Image Branch, we tailor semantic contexts for each pixel with context-guided attention and spatial context alignment to refine semantic details. Second, we transform image and voxel features into point-pixel geometry representations, simultaneously learning semantic category priors as embeddings to better represent scene geometry and semantics. Finally, to aggregate semantic information with related geometry, we design a semantic-geometry fusion that combines the transformer, effectively capturing semantic-geometry relationships into multi-modal panoptic representations. Notably, SGFormer achieves the state-of-the-art (SOTA) results on the nuScenes and SemanticPOSS, as well as yielding competitive performance on the SemanticKITTI. Moreover, SGFormer exhibits superior robustness compared to leading methods, marking an improvement of 2% to 10%.
Hongqi Yu, Sixian Chan 0001, Xiaolong Zhou 0001, Xiaoqin Zhang 0002
AAAI3
2025 Deformable Blur Sensing and Regression Analysis ReID Feature Fusion for Multitarget Multicamera Tracking Systems in Highway Scenarios
abstract
In highway scenarios, the rapid motion of vehicles can cause deformation and blur in camera footage, significantly affecting the accuracy of vehicle detection and re-identification (ReID) in multitarget multicamera tracking (MTMCT) systems. To address this issue, this article develops the deformable and blur sensing and regression analysis ReID feature fusion MTMCT system (DSRF). First, a deformable and blur sensing detection module (DFB) in DSRF is designed to overcome the limitations of cameras in capturing fast-moving objects, thereby accurately detecting vehicles moving at high speeds on highways. Then, a regression-based ReID feature fusion algorithm (RARF) in DSRF is proposed, which enhances ReID features by modeling the relationship between vehicle motion and its features, thereby better associating the detected vehicles in consecutive frames into trajectories and establishing intertrajectory relationships. Finally, extensive experiments are conducted on the highway surveillance traffic (HST) dataset developed by our team and the public dataset (CityFlow). Promising results are achieved, validating the effectiveness of our proposed method.
Sixian Chan 0001, Shenghao Ni, Jie Hu 0041, Tinglong Tang, Xiaolong Zhou 0001, Pengyi Hao
IEEE Trans. Comput. Soc. Syst.6
2025 Adaptive Target-Oriented Tracking
abstract
The current one-stream tracking pipelines are early relation modeling in feature extraction. However, insufficient discrimination may result in ambiguous relation modeling during early feature extraction. Moreover, the non-target information occupies most of the search image, rendering most relation modeling futile. To tackle the above issues, we propose tracking via learning adaptive target-oriented representation, named ATOTrack . We design an Untied positional encoding to mark the template token and the search region token separately, which reduces the confused relationship between the template and the search region. Besides, we introduce an Auto-Mask Learner to decouple the target and non-target information in the search region. Interestingly, the Auto-Mask Learner can self-learn and mask the ineffective information to interpret adaptive target-oriented representation. Extensive experiments demonstrate that ATOTrack is superior to existing methods, which achieves the state-of-the-art performance on six tracking benchmarks. In particular, ATOTrack establishes a new record on AViST with 57% AO. The code and models will be released as soon.
Sixian Chan 0001, Xianpeng Zeng, Zhoujian Wu, Yu Wang 0254, Xiaolong Zhou 0001, Tinglong Tang, Jie Hu 0041
ACM Trans. Intell. Syst. Technol.5
2025 PrimePSegter: Progressively Combined Diffusion for 3D Panoptic Segmentation With Multi-Modal BEV Refinement
abstract
Effective and robust 3D panoptic segmentation is crucial for scene perception in autonomous driving. Modern methods widely adopt multi-modal fusion based simple feature concatenation to enhance 3D scene understanding, resulting in generated multi-modal representations typically lack comprehensive semantic and geometry information. These methods focused on panoptic prediction in a single step also limit the capability to progressively refine panoptic predictions under varying noise levels, which is essential for enhancing model robustness. To address these limitations, we first utilize BEV space to unify semantic-geometry perceptual representation, allowing for a more effective integration of LiDAR and camera data. Then, we propose PrimePSegter, a progressively combined diffusion 3D panoptic segmentation model that is conditioned on BEV maps to iteratively refine predictions by denoising samples generated from Gaussian distribution. PrimePSegter adopts a conditional encoder-decoder architecture for fine-grained panoptic predictions. Specifically, a multi-modal conditional encoder is equipped with BEV fusion network to integrate semantic and geometric information from LiDAR and camera streams into unified BEV space. Additionally, a diffusion transformer decoder operates on multi-modal BEV features with varying noise levels to guide the training of diffusion model, refining the BEV panoptic representations enriched with semantics and geometry in a progressive way. PrimePSegter achieves state-of-the-art performance on the nuScenes and competitive results on the SemanticKITTI, respectively. Moreover, PrimePSegter demonstrates superior robustness towards various scenarios, outperforming leading methods.
Hongqi Yu, Sixian Chan 0001, Xiaolong Zhou 0001, Xiaoqin Zhang 0002
IEEE Trans. Multim.3
2025 SyNet: A Synergistic Network for 3D Object Detection Through Geometric-Semantic-Based Multi-Interaction Fusion
abstract
Driven by rising demands in autonomous driving, robotics,etc., 3D object detection has recently achieved great advancement by fusing optical images and LiDAR point data. On the other hand, most existing optical-LiDAR fusion methods straightly overlay RGB images and point clouds without adequately exploiting the synergy between them, leading to suboptimal fusion and 3D detection performance. Additionally, they often suffer from limited localization accuracy without proper balancing of global and local object information. To address this issue, we design a synergistic network (SyNet) that fuses geometric information, semantic information, as well as global and local information of objects for robust and accurate 3D detection. The SyNet captures synergies between optical images and LiDAR point clouds from three perspectives. The first is geometric, which derives high-quality depth by projecting point clouds onto multi-view images, enriching optical RGB images with 3D spatial information for a more accurate interpretation of image semantics. The second is semantic, which voxelizes point clouds and establishes correspondences between the derived voxels and image pixels, enriching 3D point clouds with semantic information for more accurate 3D detection. The third is balancing local and global object information, which introduces deformable self-attention and cross-attention to process the two types of complementary information in parallel for more accurate object localization. Extensive experiments show that SyNet achieves 70.7% mAP and 73.5% NDS on the nuScenes test set, demonstrating its effectiveness and superiority as compared with the state-of-the-art.
Xiaoqin Zhang 0002, Kenan Bi, Sixian Chan 0001, Shijian Lu, Xiaolong Zhou 0001
IEEE Trans. Multim.5
2024 DECNet: Dense embedding contrast for unsupervised semantic segmentation
Xiaoqin Zhang 0002, Xiaolong Zhou 0001, Sixian Chan 0001
Neural Networks3
2024 Video-Based Multi-Camera Vehicle Tracking via Appearance-Parsing Spatio-Temporal Trajectory Matching Network
abstract
Multi-camera vehicle tracking is a fundamental task for city traffic management to count traffic flow or monitor roads. This paper focuses on multi-camera tracking on the highway, which is more challenging compared with city streets in some problems such as fast-moving vehicles, tiny similar vehicles in appearance, longer tracking distance, and lighting intensity changes in the dark tunnels. In this paper, we propose a practical Appearance-Parsing Spatio-Temporal Trajectory Matching Network (ASTM-Net) based on the global appearance matching of local trajectory for addressing the cross-camera tracking tasks on the highway. Specifically, considering that the environmental disturbance and small vehicles have a similar appearance, we propose a multiple appearance-attribute parsing (MAP) module consisting of a Bi-propagation top-down (Bi-TD) block and appearance re-identification (ARe-ID) block to obtain salient global appearance-attribute features through given a video sequence. To address discrete tracking fragments caused by occlusion, we develop an appearance-joint-tracking (AJT) mechanism to merge the isolated tracklets with target interaction and occlusion handling. We then exploit an appearance-informed spatio-temporal matching (ASTM) module to achieve multi-camera tracklet-totarget assignment, which employs spatio-temporal consistency relation for intra-camera trajectory correction and coarse intercamera tracklet correlation and aggregate appearance matrix of local trajectories for assigning global trajectory ID. Finally, in order to evaluate our proposed ASTM-Net, a new dataset, named HST, collected on the highway is established.We verify the ASTM-Net on the HST and the other three public datasets,i.e., CityFlow, UA-DETRAC, and Synthehicle, whose experimental results demonstrate the effectiveness and robustness of the proposed method.
Xiaoqin Zhang 0002, Hongqi Yu, Xiaolong Zhou 0001, Sixian Chan 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 MLPT: Multilayer Perceptron based Tracking
abstract
The global receptive field plays a critical role in visual object tracking. In most popular tracking paradigms, we find that the local receptive field introduced by the convolutional neural network prevents the tracker from focusing on the long-range dependency. Although the Vision transformer brings the global receptive field in downstream tasks, its computational burden remains unaffordable. In this paper, we present a simple yet effective Multilayer Perceptron-based Tracking (MLPT), including the global receptive field. The MLPT contains three Components: Feature Correlation (FC) module, Global Information Encoder (GIE) module and Corner Head(CH). Firstly, the FC module is proposed to effectively converge the template and search region features for generating delicate features. Secondly, the GIE is designed to integrate the channel and spatial-information dealed with channel-encoding the token-encoding, separately. Specially, the same kernel is utilized for token-encoding in all channels so that our model has the global receptive field. Then, the CH is applied to establish a simple flexible way via computing the box corners coordinates for tracking. Finally, the MLPT, to our knowledge, is the first baseline of MLP-based architecture for object tracking. Extensive experiments are conducted on four challenging datasets, including GOT-10K, LaSOT, UAV123, and TrackingNet. The results show that the proposed method achieves state-of-the-art performance.
Sixian Chan 0001, Yu Wang 0254, Xiaolong Zhou 0001, Qike Shao
SMC4
2022 A wearable-based posture recognition system with AI-assisted approach for healthcare IoT
Zhen Hong, Miao Hong, Xiaolong Zhou 0001, Wei Wang 0077
Future Gener. Comput. Syst.5
2022 Online multiple object tracking using joint detection and embedding network
Sixian Chan 0001, Yangwei Jia, Xiaolong Zhou 0001, Cong Bai, Shengyong Chen, Xiaoqin Zhang 0002
Pattern Recognit.3
2022 A Multitarget Interested Region Extraction Method for Wrist X-Ray Images Based on Optimized AlexNet and Two-Class Combined Model
abstract
Bone age assessment based on X-ray Images can accurately determine the actual bone age of adolescents. Accurate extraction of the key regions of interest (ROIs) in X-ray images is required to accurately assess bone age. However, existing ROI extraction methods can only extract a few targets and have poor extraction accuracy. Thus, the strict demands for imaging in the medical field via these methods are difficult. In this article, we propose a multitarget interested region extraction method for wrist X-ray images based on optimized AlexNet and two-class combined model, named OATC, which can simultaneously extract multiple ROIs with high accuracy. Specifically, the square-wave scanning algorithm was implemented to obtain the bounding box size of each bone ROI according to the shape information of the wrist. Then, the optimized AlexNet was used to obtain the key point coordinates of each bone ROI. Bone ROIs could be extracted by combining key point coordinates with the bone bounding box size. Finally, the two-classification model was combined to improve the accuracy of ROI extraction. Experiments conducted on our wrist X-ray image dataset showed that OTAC has a fast convergence speed and small deviation. The average accuracy of extracting 14 ROIs reached 95.57%, which are 7.76% and 4.68% higher than that of VGG16 and AlexNet, respectively.
Kai Fang 0001, Xiaolong Zhou 0001, Keji Mao
IEEE Trans. Comput. Soc. Syst.3
2022 A TOPSIS-Based Relocalization Algorithm in Wireless Sensor Networks
abstract
Selecting reliable beacon nodes plays a significant role in relocalizing unknown nodes in a wireless sensor network. When the position of a beacon node is drifted or is spoofed, it becomes an unreliable beacon node, which would lead to a large relocalization deviation of unknown nodes in its neighbor. However, when selecting reliable beacon nodes, most relocalization algorithms only screen either drifting beacon nodes or malicious beacon nodes whose position is drifted or spoofed. This article proposes an algorithm that can simultaneously screen drifting beacon nodes and malicious beacon nodes. The algorithm is divided into four steps. First, three indicators are introduced, where two are for describing position drifting and one is for describing position spoofing. Second, the entropy method is used to weight the contributions of three indicators. Third, a technique for order preference by similarity to an ideal solution is used to construct a reliability evaluation model. Finally, using the reliability evaluation model select reliable beacon nodes. Experimental results illustrate that the detection accuracy of drifting beacon nodes and malicious beacon nodes of the proposed algorithm is 7.5% and 8.2% higher than that of the state-of-the-art algorithms, respectively.
Kai Fang 0001, Tingting Wang 0006, Xiaolong Zhou 0001, Yaping Ren, Hongfei Guo, Jianqing Li 0001
IEEE Trans. Ind. Informatics3
2022 Siamese Implicit Region Proposal Network With Compound Attention for Visual Tracking
abstract
Recently, siamese-based trackers have achieved significant successes. However, those trackers are restricted by the difficulty of learning consistent feature representation with the object. To address the above challenge, this paper proposes a novel siamese implicit region proposal network with compound attention for visual tracking. First, an implicit region proposal (IRP) module is designed by combining a novel pixel-wise correlation method. This module can aggregate feature information of different regions that are similar to the pre-defined anchor boxes in Region Proposal Network. To this end, the adaptive feature receptive fields then can be obtained by linear fusion of features from different regions. Second, a compound attention module including a channel and non-local attention is raised to assist the IRP module to perform a better perception of the scale and shape of the object. The channel attention is applied for mining the discriminative information of the object to handle the background clutters of the template, while non-local attention is trained to aggregate the contextual information to learn the semantic range of the object. Finally, experimental results demonstrate that the proposed tracker achieves state-of-the-art performance on six challenging benchmark tests, including VOT-2018, VOT-2019, OTB-100, GOT-10k, LaSOT, and TrackingNet. Further, our obtained results demonstrate that the proposed approach can be run at an average speed of 72 FPS in real time.
Sixian Chan 0001, Xiaolong Zhou 0001, Cong Bai, Xiaoqin Zhang 0002
IEEE Trans. Image Process.3
2021 A Multi-Level Network for Human Pose Estimation
abstract
Although multi-person human pose estimation has made great progress in recent years, the challenges such as various scales of persons, occluded keypoints, and crowded backgrounds in complex scenes are still remained to be solved. In this paper, we propose a novel multi-level pose estimation network (MLPE) to learn multi-level features that can preserve both the strong semantic clues and spatial resolution for keypoint prediction and location. More specifically, a multi-level prediction network with a feature enhancement strategy is first proposed to learn multi-level features to achieve a good trade-off between the global context information and spatial resolution. We then build a high-resolution fine network to restore high spatial resolution information based on transposed convolutions to accurately locate the keypoints. We have conducted extensive experiments on the challenging MS COCO dataset, which has proved the effectiveness of our proposed method. Code†and the experimental results are publicly online available for further research.
Zhanpeng Shao, Youfu Li 0001, Jianyu Yang 0002, Xiaolong Zhou 0001
ICRA5
2021 Box Regression-Guided Anchor-free for Robust Visual Tracking
abstract
The Siamese tracker-based approach has achieved significant success in recent years. However, these approaches do not consider the different requirements for input feature in classification and regression branches. The regression branch needs feature information slightly larger than the object region, while the classification branch needs to avoid classification failure caused by the introduction of background information. In this paper, we present a novel Box Regression-Guided Anchor-free for Robust Visual Tracking. Firstly, a scale-aware regression module is designed to satisfy the feature requirements of the regression branch, which can capture feature information of various scales. Secondly, regression-guided classification module is applied to aligning the feature between the regression result and correlation feature, thereby avoiding the introduction of background information to classification branch. In addition, the new correlation operation is introduced to gain more superb correlation feature. Comparsion experimental exhibits that the proposed tracker achieves promising results in five challenging benchmark tests, including GOT-10K, OTB-2015, VOT-2018, VOT-2019 and TrackingNet, and run at an average speed of 60 FPS in real-time.
Sixian Chan 0001, Xiaolong Zhou 0001, Cong Bai, Hua Gao, Shengyong Chen
SMC3
2020 Improved itracker combined with bidirectional long short-term memory for 3D gaze estimation using appearance cues
Xiaolong Zhou 0001, Jianing Lin, Zhuo Zhang 0012, Zhanpeng Shao, Shenyong Chen, Honghai Liu 0001
Neurocomputing1
2020 Deploy Efficiency Driven k-Barrier Construction Scheme Based on Target Circle in Directional Sensor Network
Xinggang Fan, Zhi-Cong Che, Fengdan Hu, Tao Liu 0031, Jinshan Xu, Xiaolong Zhou 0001
J. Comput. Sci. Technol.6
2020 Efficient Implementation of Truncated Reweighting Low-Rank Matrix Approximation
abstract
The weighted nuclear norm minimization and truncated nuclear norm minimization are two well-known low-rank constraint for visual applications. In this paper, by integrating their advantages into a unified formulation, we find a better weighting strategy, namely truncated reweighting norm minimization (TRNM), which provides better approximation to the target rank for some specific task. Albeit nonconvex and truncated, we prove that TRNM is equivalent to certain weighted quadratic programming problems, whose global optimum can be accessed by the newly presented reweighting singular value thresholding operator. More importantly, we design a computationally efficient optimization algorithm, namely momentum update and rank propagation (MURP), for the general TRNM regularized problems. The individual advantages of MURP include, first, reducing iterations through nonmonotonic search, and second, mitigating computational cost by reducing the size of target matrix. Furthermore, the descent property and convergence of MURP are proven. Finally, two practical models, i.e., Matrix Completion Problem via TRNM (MCTRNM) and Space Clustering Model via TRNM (SCTRNM), are presented for visual applications. Extensive experimental results show that our methods achieve better performance, both qualitatively and quantitatively, compared with several state-of-the-art algorithms.
Jianwei Zheng 0001, Xiaolong Zhou 0001, Jiafa Mao, Hongchuan Yu
IEEE Trans. Ind. Informatics3
2019 Learning A 3D Gaze Estimator with Improved Itracker Combined with Bidirectional LSTM
abstract
Free-head 3D gaze estimation which outputs gaze vector in 3D space has wide application in human-computer interaction. In this paper, we propose a novel 3D gaze estimator by improving the Itracker and employing a many-to-one bidirectional LSTM (bi-LSTM). First, we improve the conventional Itracker by removing the face-grid and reducing one network branch via concatenating the two-eye region images to predict the subject's gaze of a single frame. Then, we employ the bi-LSTM to fit the temporal information between frames to estimate gaze vector for video sequence. Experimental results show that our improved Itracker obtains 11.6% significant improvement over the state-of-the-art methods on MPIIGaze dataset (single image frame) and has robust estimation accuracy for different image resolutions. Moreover, experimental results on EyeDiap dataset (video sequence) further bring 3% accuracy improvement by employing the bi-LSTM.
Xiaolong Zhou 0001, Jianing Lin, Shengyong Chen
ICME1
2019 Dictionary Learning and Confidence Map Estimation-Based Tracker for Robot-Assisted Therapy System
Xiaolong Zhou 0001, Sixian Chan 0001, Shengyong Chen, Honghai Liu 0001
PRCV (1)1
2019 Deep learning for multiple object tracking: a survey
abstract
Deep learning has been proved effective in multiple object tracking, which confronts the difficulties of frequent occlusions, confusing appearance, in‐and‐out objects, and lack of enough labelled data. Recently, deep learning based multi‐object tracking methods make a rapid progress from representation learning to network modelling due to the development of deep learning theory and benchmark setup. In this study, the authors summarise and analyse deep learning based multi‐object tracking methods which are top‐ranked in the public benchmark test. First, they investigate functionality of deep networks in these methods, and classify the methods into three categories as description enhancement using deep features, deep network embedding, and end‐to‐end deep network construction. Second, they review deep network structures in these methods, and detail the usage and training of these networks for multi‐object tracking problem. Through experimental comparison of tracking results in the benchmarks in total and by group, they finally show the effectiveness of deep networks for tracking employed in different manners, and compare the advantages of these networks and their robustness under different tracking conditions. Moreover, they analyse the limitations of current methods, and draw some useful conclusions to facilitate the exploration of new directions for multi‐object tracking.
Yingkun Xu, Xiaolong Zhou 0001, Shengyong Chen, Fenfen Li
IET Comput. Vis.2
2019 A Hierarchical Model for Human Action Recognition From Body-Parts
abstract
As increasing attention is paid to human action recognition from skeleton data, this paper focuses on such tasks by proposing a hierarchical model to discover the structure information of body-parts involved in actions for better analysis of human actions in the skeleton data. Considering human actions as simultaneous motions of body-parts of the human skeleton, we propose a hierarchical model to simultaneously apply discriminative body-parts selection at a same scale and group coupling of bundles of body-parts at different scales, while we decompose the human skeleton into a hierarchy of body-parts of varying scales. To represent such hierarchy of body-parts, we accordingly build a hierarchical rotation and relative velocity (HRRV) descriptor. The hierarchical representations encoded by Fisher vectors of the HRRV descriptors are properly formulated into the hierarchical model via the proposed mixed norm, to apply the sparse selection of body-parts and regularize the structure of such hierarchy of body-parts. The extensive evaluations on three challenging datasets demonstrate the effectiveness of our proposed approach, which achieves superior performance compared to the state-of-the-art algorithms on datasets with various sizes, showing it is more widely applicable than existing approaches.
Zhanpeng Shao, Youfu Li 0001, Yao Guo 0002, Xiaolong Zhou 0001, Shengyong Chen
IEEE Trans. Circuits Syst. Video Technol.4
2018 Oscillation Detection and Parameter-Adaptive Hedge Algorithm for Real-Time Visual Tracking
Bolin Lv, Xiaolong Zhou 0001, Shengyong Chen
PRCV (4)2
2018 Online classification for object tracking based on superpixel
Sixian Chan 0001, Xiaolong Zhou 0001, Shengyong Chen
Neurocomputing2
2017 Compressive tracking with locality sensitive histograms features
abstract
Currently, Compressive Tracking (CT) method has drawn great attention because of its high efficiency. However, it cannot well deal with some appearance variations due to its limitations of feature expression and it only uses a fixed parameter to update the appearance model. In order to handle such matters, we propose an adaptive CT method that combines the predicted target position with CT based on Locality Sensitive Histograms (LSH) features. Our method significantly improves CT in four aspects. First, the efficient illumination invariant features extracted based on LSH are used to represent an effective appearance model that is robust to illumination changes. Second, the color attributes tracker is adopted to predict the target position for re-building the new weighted discriminant function which brings in the color information to make up for the inadequacy of Haar-like characteristics. Third, a new model update mechanism is proposed to preserve the stable features while avoid the noisy appearance variations during tracking. Fourth, a trajectory rectification method is employed to refine the tracking location when possible inaccurate tracking occurs. Finally, we show that our tracker achieves state-of-the-art performance in a comprehensive evaluation over 47 challenging color sequences.
Sixian Chan 0001, Xiaolong Zhou 0001, Zhuo Zhang 0012, Shengyong Chen
ICRA2
2017 Two-eye model-based gaze estimation from a Kinect sensor
abstract
In this paper, we present an effective and accurate gaze estimation method based on two-eye model of a subject with the tolerance of free head movement from a Kinect sensor. To accurately and efficiently determine the point of gaze, i) we employ two-eye model to improve the estimation accuracy; ii) we propose an improved convolution-based means of gradients method to localize the iris center in 3D space; iii) we present a new personal calibration method that only needs one calibration point. The method approximates the visual axis as a line from the iris center to the gaze point to determine the eyeball centers and the Kappa angles. The final point of gaze can be calculated by using the calibrated personal eye parameters. We experimentally evaluate the proposed gaze estimation method on eleven subjects. Experimental results demonstrate that our gaze estimation method has an average estimation accuracy around 1.99°, which outperforms many leading methods in the state-of-the-art.
Xiaolong Zhou 0001, Haibin Cai, Youfu Li 0001, Honghai Liu 0001
ICRA1
2017 Object tracking using a convolutional network and a structured output SVM
abstract
Object tracking has been a challenge in computer vision. In this paper, we present a novel method to model target appearance and combine it with structured output learning for robust online tracking within a tracking-by-detection framework. We take both convolutional features and handcrafted features into account to robustly encode the target appearance. First, we extract convolutional features of the target by kernels generated from the initial annotated frame. To capture appearance variation during tracking, we propose a new strategy to update the target and background kernel pool. Secondly, we employ a structured output SVM for refining the target’s location to mitigate uncertainty in labeling samples as positive or negative. Compared with existing state-of-the-art trackers, our tracking method not only enhances the robustness of the feature representation, but also uses structured output prediction to avoid relying on heuristic intermediate steps to produce labelled binary samples. Extensive experimental evaluation on the challenging OTB-50 video sequences shows competitive results in terms of both success and precision rate, demonstrating the merits of the proposed tracking method.
Xiaolong Zhou 0001, Sixian Chan 0001, Shengyong Chen
Comput. Vis. Media2
2017 Adaptive Compressive Tracking based on Locality Sensitive Histograms
Sixian Chan 0001, Xiaolong Zhou 0001, Shengyong Chen
Pattern Recognit.2
2016 A pipeline using multi-layer Tumors Automata for interactive multi-label image segmentation
abstract
In this paper, we investigate a novel algorithm to the problem of interactive image segmentation. We propose an extension of the Growcut framework using the Tumors Automata (TA) formed from the superpixel. The proposed TA is similar to Cellular Automata but can directly deal with superpixel. The superpixels (image segments) can provide powerful boundary cues to guide segmentation, where superpixels can be collected easily by over-segmenting the image using any reasonable existing segmentation algorithms. Given a small number of user-labelled superpixels, the rest of the image is segmented automatically by a TA. When the automaton labels the image, the segmentation evolution is faster than Growcut because of the iterative process. Moreover, a level set method and multi-layer TA are employed to further improve the performance. Experiments conducted on the Berkeley Segmentation Database demonstrate the superior performance of our method over the state-of-the-art methods.
Sixian Chan 0001, Xiaolong Zhou 0001, Zhuo Zhang 0012, Shengyong Chen
HSI2
2016 Combining 3D joints Moving Trend and Geometry property for human action recognition
abstract
Depth image based human action recognition has attracted many attentions due to the popularity of the depth sensors. However, accurate recognition still remains a challenge because of various object appearances, poses and video sequences. In this paper, a novel skeleton joints descriptor based on 3D Moving Trend and Geometry (3DMTG) property is proposed for human action recognition. Specifically, a histogram of 3D moving directions between consecutive frames for each joint is constructed to represent the 3D moving trend feature in spatial domain. The geometry information of joints in each frame is modelled by the relative motion with the initial status. The proposed feature descriptor is evaluated on two popular datasets. The experimental results demonstrate the superior performance of our method over the state-of-the-art methods, especially the higher recognition rates for complex actions.
Bangli Liu, Hui Yu 0001, Xiaolong Zhou 0001, Honghai Liu 0001
SMC3
2015 Learning Local Appearances With Sparse Representation for Robust and Fast Visual Tracking
abstract
In this paper, we present a novel appearance model using sparse representation and online dictionary learning techniques for visual tracking. In our approach, the visual appearance is represented by sparse representation, and the online dictionary learning strategy is used to adapt the appearance variations during tracking. We unify the sparse representation and online dictionary learning by defining a sparsity consistency constraint that facilitates the generative and discriminative capabilities of the appearance model. An elastic-net constraint is enforced during the dictionary learning stage to capture the characteristics of the local appearances that are insensitive to partial occlusions. Hence, the target appearance is effectively recovered from the corruptions using the sparse coefficients with respect to the learned sparse bases containing local appearances. In the proposed method, the dictionary is undercomplete and can thus be efficiently implemented for tracking. Moreover, we employ a median absolute deviation based robust similarity metric to eliminate the outliers and evaluate the likelihood between the observations and the model. Finally, we integrate the proposed appearance model with the particle filter framework to form a robust visual tracking algorithm. Experiments on benchmark video sequences show that the proposed appearance model outperforms the other state-of-the-art approaches in tracking performance.
Tianxiang Bai, Youfu Li 0001, Xiaolong Zhou 0001
IEEE Trans. Cybern.3
2014 Entropy distribution and coverage rate-based birth intensity estimation in GM-PHD filter for multi-target visual tracking
Xiaolong Zhou 0001, Youfu Li 0001, Bingwei He
Signal Process.1
2014 GM-PHD-Based Multi-Target Visual Tracking Using Entropy Distribution and Game Theory
abstract
Tracking multiple moving targets in a video is a challenge because of several factors, including noisy video data, varying number of targets, and mutual occlusion problems. The Gaussian mixture probability hypothesis density (GM-PHD) filter, which aims to recursively propagate the intensity associated with the multi-target posterior density, can overcome the difficulty caused by the data association. This paper develops a multi-target visual tracking system that combines the GM-PHD filter with object detection. First, a new birth intensity estimation algorithm based on entropy distribution and coverage rate is proposed to automatically and accurately track the newborn targets in a noisy video. Then, a robust game-theoretical mutual occlusion handling algorithm with an improved spatial color appearance model is proposed to effectively track the targets in mutual occlusion. The spatial color appearance model is improved by incorporating interferences of other targets within the occlusion region. Finally, the experiments conducted on publicly available videos demonstrate the good performance of the proposed visual tracking system.
Xiaolong Zhou 0001, Youfu Li 0001, Bingwei He, Tianxiang Bai
IEEE Trans. Ind. Informatics1
2013 Multi-target visual tracking with game theory-based mutual occlusion handling
abstract
Tracking multiple moving targets in video is still a challenge because of mutual occlusion problem. This paper presents a Gaussian mixture probability hypothesis density-based visual tracking system with game theory-based mutual occlusion handling. First, a two-step occlusion reasoning algorithm is proposed to determine the occlusion region. Then, the spatial constraint-based appearance model with other interacting targets¶ interferences is modeled. Finally, an n-person, non-zero-sum, non-cooperative game is constructed to handle the mutual occlusion problem. The individual measurements within the occlusion region are regarded as the players in the constructed game competing for the maximum utilities by using the certain strategies. The Nash Equilibrium of the game is the optimal estimation of the locations of the players within the occlusion region. Experiments conducted on publicly available videos demonstrate the good performance of the proposed occlusion handling algorithm.
Xiaolong Zhou 0001, Youfu Li 0001, Bingwei He, Tianxiang Bai
IROS1
2013 Game-theoretical occlusion handling for multi-target visual tracking
Xiaolong Zhou 0001, Youfu Li 0001, Bingwei He
Pattern Recognit.1
2012 Robust and fast visual tracking using constrained sparse coding and dictionary learning
abstract
We present a novel appearance model using sparse coding with online sparse dictionary learning techniques for robust visual tracking. In the proposed appearance model, the target appearance is modeled via online sparse dictionary learning technique with an “elastic-net constraint”. This scheme allows us to capture the characteristics of the target local appearance, and promotes the robustness against partial occlusions during tracking. Additionally, we unify the sparse coding and online dictionary learning by defining a “sparsity consistency constraint” that facilitates the generative and discriminative capabilities of the appearance model. Moreover, we propose a robust similarity metric that can eliminate the outliers from the corrupted observations. We then integrate the proposed appearance model with the particle filter framework to form a robust visual tracking algorithm. Experiments on publicly available benchmark video sequences demonstrate that the proposed appearance model improves the tracking performance compared with other state-of-the-art approaches.
Tianxiang Bai, Youfu Li 0001, Xiaolong Zhou 0001
IROS3
2012 Birth intensity online estimation in GM-PHD filter for multi-target visual tracking
abstract
Multi-target tracking in video is a challenge due to noisy video data, varying number of targets, and the data association problems. In this paper, a multi-target visual tracking system that incorporates object detection with the Gaussian mixture PHD filter is developed. The main contribution of this paper is to propose a new birth intensity online estimation method that based on the entropy distribution and the coverage rate. First, the birth intensity is initialized by using the previously obtained targets' states and measurements. The measurements are obtained by object detection and classified into the birth measurements and the survival measurements. Then it is updated according to the currently obtained birth measurements. In the update stage, the instability of the entropy distribution is applied to remove components like noises within the birth intensity which are irrelevant with the currently obtained birth measurements. And the coverage rate between each birth intensity component and corresponding birth measurement is computed to further eliminate the noises. Finally, experiments are implemented to show the performance of the proposed visual tracking system, especially to show the good performance for tracking the newborn targets.
Xiaolong Zhou 0001, Youfu Li 0001, Bingwei He, Tianxiang Bai, Yazhe Tang
IROS1
2009 A Novel View Planning Method for Automatic Reconstruction of Unknown 3-D Objects Based on the Limit Visual Surface
abstract
Automatic reconstruction of unknown 3-D objects has been of great importance in the areas of machine vision, object recognition, and automatic modeling. In this paper, a new planning approach of generating 3-D models automatically is proposed. The new algorithm incorporates the limit visual surfaces of unknown model which are obtained according to both of the known object boundary knowledge and the visual region of the vision system and selects the suitability of viewpoints as the next best view on scanning coverage. The limit visual surfaces are used to predict the maximal information of unknown model and then the visibility criterion of next viewpoint is determined. And the position which can obtain the maximal visual surface area is defined as the next best view position. The reconstruction result of real model with proposed method show the efficiency in practical implementation.
Xiaolong Zhou 0001, Bingwei He, Youfu Li 0001
ICIG1