Yan Zhou 0003

dblp:60/5157-3 · DBLP profile ↗
← Back
36ranked-venue papers
14as first author
21since 2021 · last 2026
0000-0002-2372-4947ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-authorSystems, architecture and hardware · 3Human-computer interaction and ubiquitous computing · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorComputer networks · 1
YearPublicationVenuePosition
2026 Deep learning-based group activity recognition in videos: A survey
Xiaolin Zhu 0004, Dongli Wang, Yan Zhou 0003, Yongcan Weng
Neurocomputing3
2026 Enhanced skeleton-based Group Activity Recognition through spatio-temporal graph convolution with cross-dimensional attention
Dongli Wang, Yongcan Weng, Xiaolin Zhu 0004, Yan Zhou 0003, Richard Irampaye
Image Vis. Comput.4
2026 ExGes: Expressive Human Motion Retrieval and Modulation for Audio-Driven Gesture Synthesis
abstract
Audio-driven human gesture synthesis is a crucial task with broad applications in virtual avatars, human-computer interaction, and creative content generation. Despite notable progress, existing methods often produce coarse gestures, lack expressiveness, and fail to fully align with audio semantics. To address these challenges, we propose ExGes, a novel retrieval-enhanced diffusion framework with three key designs: (1) a Motion Base Construction, which builds a gesture library from the training dataset; (2) a Motion Retrieval Module, employing contrastive learning and momentum distillation for retrieving fine-grained reference poses; and (3) a Precise Control Module, integrating partial masking and stochastic masking to enable flexible and fine-grained control. Experimental evaluations on BEAT2 demonstrate that ExGes reduces Fréchet Gesture Distance by 4.55%and improves motion diversity by 5.3% over EMAGE, with user studies revealing a 71.3% preference for its naturalness and semantic relevance.
Xukun Zhou, Fengxin Li, Yan Zhou 0003, Pengfei Wan 0001, Yeying Jin, Hongyuan Zhang 0001, Hongyan Liu 0002, Zhaoxin Fan, Jun He 0008, Xuelong Li 0001
IEEE Trans. Vis. Comput. Graph.4
2025 Multi-dimensional convolution transformer for group activity recognition
Dongli Wang, Xiaolin Zhu 0004, Yan Zhou 0003
Multim. Tools Appl.5
2025 Voxel completion and 3D asymmetrical convolution networks for Lidar semantic segmentation
Yan Zhou 0003, Haibin Zhou
Multim. Tools Appl.1
2025 Dynamical Attention Hypergraph Convolutional Network for Group Activity Recognition
abstract
Recently, group activity recognition (GAR) has drawn growing interests in video analysis and computer vision communities. The current models of GAR tasks are often impractical in that they suppose that all interactions between actors are pairwise, which only models and leverages part of the information in real entire interactions. Motivated by this, we design a distinct dynamical attention hypergraph convolutional network framework, referred to as DAHGCN, for precise GAR, modeling the entire interactions and capturing the high-order relationships among involved actors in a real-life scenario. Specifically, to learn complementary feature representations for fine-grained GAR, a multilevel feature descriptor (MLFD) module is proposed. Furthermore, for learning higher order interaction relationships, we construct a DAHGCN to accommodate complex group interactions, which can dynamically change the topology of the hypergraph and learn these key representations by virtue of the "similarity-based shared nearest-neighbor (SSNN) clustering" and "attention mechanisms" on hypergraph. Finally, a multiscale temporal convolution (MSTC) module is utilized to explore various long-range temporal dynamic correlations across different frames. In addition, comprehensive experiments on three commonly used GAR datasets clearly demonstrate that, when compared with the state-of-the-art methods, our proposed method can achieve the most optimal performance.
Xiaolin Zhu 0004, Dongli Wang, Qin Wan 0001, Yan Zhou 0003
IEEE Trans. Neural Networks Learn. Syst.6
2025 DiffCap: Diffusion-Based Real-Time Human Motion Capture Using Sparse IMUs and a Monocular Camera
abstract
Combining sparse IMUs and a monocular camera is a new promising setting to perform real-time human motion capture. This paper proposes a diffusion-based solution to learn human motion priors and fuse the two modalities of signals together seamlessly in a unified framework. By delicately considering the characteristics of the two signals, the sequential visual information is considered as a whole and transformed into a condition embedding, while the inertial measurement is concatenated with the noisy body pose frame by frame to construct a sequential input for the diffusion model. Firstly, we observe that the visual information may be unavailable in some frames due to occlusions or subjects moving out of the camera view. Thus incorporating the sequential visual features as a whole to get a single feature embedding is robust to the occasional degenerations of visual information in those frames. On the other hand, the IMU measurements are robust to occlusions and always stable when signal transmission has no problem. So incorporating them frame-wisely could better explore the temporal information for the system. Experiments have demonstrated the effectiveness of the system design and its state-of-the-art performance in pose estimation compared with the previous works. The code will be released.
Shaohua Pan 0002, Xinyu Yi, Yan Zhou 0003, Weihua Jian, Yuan Zhang 0020, Pengfei Wan 0001, Feng Xu 0005
IEEE Trans. Vis. Comput. Graph.3
2025 Lunet: an enhanced upsampling fusion network with efficient self-attention for semantic segmentation
Yan Zhou 0003, Haibin Zhou, Yin Yang 0003, Richard Irampaye, Dongli Wang, Zhengpeng Zhang
Vis. Comput.1
2024 Multi-object tracking using context-sensitive enhancement via feature fusion
Yan Zhou 0003, Dongli Wang, Xiaolin Zhu 0004
Multim. Tools Appl.1
2024 DCMA-Net: dual cross-modal attention for fine-grained few-shot recognition
Yan Zhou 0003, Xiao Ren, Yin Yang 0003, Haibin Zhou
Multim. Tools Appl.1
2023 Hierarchical Spatial-Temporal Transformer with Motion Trajectory for Individual Action and Group Activity Recognition
abstract
Group activity recognition, which aims to simultaneously understand individual action and group activity in video clips, plays a fundamental role in video analysis. In this paper, we propose a novel reasoning network, Hierarchical Spatial-Temporal Transformer termed HSTT, for individual action and group activity recognition, which focuses on capturing the various degrees of spatial-temporal dynamic interactions adaptively and jointly among actors. Specifically, we first design a hierarchical spatial-temporal Transformer by capturing different levels of relationships to deal with unequal interaction relationships among actors. Furthermore, our proposed spatial-temporal Transformer (STT) block is capable of fully mining long-range spatial-temporal interactions with the virtue of the merge function and cross attention mechanism. Besides, we adopt the motion trajectory branch to provide complementary dynamic features for improving recognition performance. Extensive experiments on the two public GAR datasets clearly show that our approach can achieve very competitive performance by comparing them with state-of-the-art works.
Xiaolin Zhu 0004, Dongli Wang, Yan Zhou 0003
ICASSP3
2023 Automatic Human Scene Interaction through Contact Estimation and Motion Adaptation
abstract
Human scene interaction (HSI) aims to understand and accommodate the various ways humans interact with their environment. However, existing works typically struggle to produce high-precision contact estimation for understanding these interactions and lack sufficient fine-grained semantics to adapt to the complexities of diverse characters and scenes, resulting in unnatural and inaccurate interacting motions. In this paper, we present a novel approach to automatic human scene interaction that effectively recovers the human body mesh and high-precision contact information, subsequently enabling adaptation to different environments. Our main contributions include the proposal of a contact estimation framework that leverages semantic features from 2D images and 3D model recovered with inverse kinematics to guide the learning of vertex-level human scene contact (HSC) estimation. For motion adaptation, we propose an enhanced Laplacian semantics descriptor combined with kinematic constraints, enabling precise retargeting between variously sized human models and distinct 3D scenes. Through extensive experiments, we demonstrate our method's superiority against state-of-the-art approaches and showcase the results of the automatic human scene interaction process.
Yan Zhou 0003, Li Chen 0031, Weihua Jian, Pengfei Wan 0001
ACM Multimedia3
2023 Adversarial learning based intermediate feature refinement for semantic segmentation
Dongli Wang, Zhitian Yuan, Wanli Ouyang, Baopu Li, Yan Zhou 0003
Appl. Intell.5
2023 Multi-directional feature refinement network for real-time semantic segmentation in urban street scenes
abstract
Abstract Efficient and accurate semantic segmentation is crucial for autonomous driving scene parsing. Capturing detailed information and semantic information efficiently through two‐branch networks has been widely utilised in real‐time semantic segmentation. This study proposes a network named MRFNet based on two‐branch strategy to solve the problem of accuracy and speed of segmentation in urban scenes. Many real‐time networks do not comprehensively consider contextual information from sub‐regions in different directions and at different scales. To handle this problem, a Multi‐directional Feature Refinement Module (MFRM) which has three sub‐paths to capture information at different scales and directions is proposed. And MFRM reduces computation by using strip pooling and dilated convolution operations. In particular, the authors propose a Feature Cross‐guide Aggregation Module to aggregate detailed information and contextual information through the mutual guidance of detailed information and semantic information. This module guides the extraction of feature maps in a more precise direction. Experiments on Cityscapes and CamVid datasets demonstrate the effectiveness of our method by achieving a balance between accuracy and inference speed. Specially, on single 1080Ti GPU, our method yields 78.9% mean intersection over union (mIoU) and 77.4% mIoU at speed of 144.5 frames per second (FPS) and 120.8 FPS on Cityscapes and CamVid datasets respectively.
Yan Zhou 0003, Xihong Zheng, Yin Yang 0003, Jinzhen Mu, Richard Irampaye
IET Comput. Vis.1
2023 A Strip Dilated Convolutional Network for Semantic Segmentation
Yan Zhou 0003, Xihong Zheng, Wanli Ouyang, Baopu Li
Neural Process. Lett.1
2023 MLST-Former: Multi-Level Spatial-Temporal Transformer for Group Activity Recognition
abstract
Group activity recognition, which aims to simultaneously understand individual action and group activity in video clips, plays a fundamental role in computer vision and video analysis. In this paper, we propose a novel relational inference framework, termed MLST-Former, for individual action and group activity recognition, capturing the various degrees of spatial-temporal dynamic interactions adaptively and jointly among actors and generating reasonable group representations. Specifically, we first design a multi-level spatial-temporal Transformer to capture the miscellaneous actors’ spatial-temporal contextual information to deal with unbalanced interactions between actors. Furthermore, our proposed network is capable of fully mining long-range spatial-temporal dependencies with the virtue of the merge function and cross attention mechanism. We then propose an inter-frame gating fusion mechanism (IGFM) to selectively aggregate the temporal and structural features of the interacting actors. A new multi-task learning strategy, consisting of the classification cost of individual actions and group activities, is also developed. Moreover, we adopt the motion trajectory branch to provide complementary dynamic features for improving recognition performance. A series of ablation studies demonstrate the effectiveness and respective contributions of the different components within the proposed method. Extensive experiments on four public GAR datasets clearly show that our approach can achieve very competitive performance by comparing them with state-of-the-art methods.
Xiaolin Zhu 0004, Yan Zhou 0003, Dongli Wang, Wanli Ouyang
IEEE Trans. Circuits Syst. Video Technol.2
2022 Lightweight Self-Attention Network for Semantic Segmentation
abstract
The deep neural network model based on self-attention (SA) for obtaining rich contextual information has been widely adopted in semantic segmentation. However, the computational complexity of the standard self-attentive module is high, which partly limits the use of this module. In this work, we propose the lightweight self-attention network (LSANet) for semantic segmentation. Specifically, the Lightweight Self-Attentive Module (LSAM) captures information using a hand-designed compact feature representation, and weighted fusion of position information. In the decoder structure, an improved up-sampling module is proposed. Compared with the bilinear upsampling, this method achieves better results in restoring image details. The experimental results on PASCAL VOC 2012, and Cityscapes datasets show the effectiveness of our method, which simplifies operations and improves performance.
Yan Zhou 0003, Haibin Zhou, Nanjun Li, Dongli Wang
IJCNN1
2022 Taming mode collapse in generative adversarial networks using cooperative realness discriminators
abstract
Abstract Generative adversarial networks (GANs) are able to produce realistic images. However, GANs may suffer mode collapse in their output data distribution. Here, we theoretically and empirically justify generalizing the GAN framework to multiple discriminators with one generator for improving generative performance. First, a comprehensive perspective is adopted to understand why mode collapse occurs. Second, an array of cooperative realness discriminators is introduced into the GAN framework to combat mode collapse and explore discriminator roles ranging from a formidable adversary to a forgiving teacher. Third, two types of simple yet effective regularization are proposed for generating realistic and diverse images. Experiments on various datasets show the effectiveness of the GAN compared to previous methods in alleviating mode collapse and improving the quality of the generated samples.
Jinzhen Mu, Chunyan Chen, Wenshan Zhu, Shuang Li 0004, Yan Zhou 0003
IET Image Process.5
2021 Integration of gradient guidance and edge enhancement into super-resolution for small object detection in aerial images
abstract
Abstract Detecting small objects are difficult because of their poor‐quality appearance and small size, and such issues are especially pronounced for aerial images of great importance. To address the small object detection (SOD) problem, a united architecture that tries to upsample small objects into super‐resolved versions, achieving characteristics similar to those large objects and thus resulting in more discriminative detection is used. For this purpose, a new end‐to‐end multi‐task generative adversarial network (GAN) is proposed. In the architecture, the generator is a super‐resolution (SR) network, and the discriminator is a multi‐task network. In the generator, a gradient guide and an edge‐enhancement strategy are introduced to alleviate structural distortions. In the discriminator, a faster region‐based convolutional neural network (FRCNN) is incorporated for the task of object detection. Specifically, the discriminator outputs a distribution scalar to measure the realness. Then, each super‐resolved image passes through the discriminator with a realness distribution, classification scores, and bounding box regression offsets. Furthermore, the losses of the detection task are backpropagated into the generator during training rather than being optimized independently. Extensive experiments on the challenging cars overhead with context dataset (COWC), detectIon in optical remote sensing images (DIOR), vision meets drones (VisDrone), and dataset for object detection in aerial images (DOTA) demonstrate the effectiveness of the proposed method in reconstructing structures while generating natural super‐resolved images and show the superiority of the proposed method in detecting small objects over state‐of‐the‐art detectors.
Jinzhen Mu, Shuang Li 0004, Zongming Liu, Yan Zhou 0003
IET Image Process.4
2021 Bilateral attention network for semantic segmentation
abstract
Abstract Enhancing network feature representation capabilities and reducing the loss of image details have become the focus of semantic segmentation task. This work proposes the bilateral attention network for semantic segmentation. The authors embed two attention modules in the encoder and decoder structures . Specifically, high‐level features of the encoder structure integrate all channel maps through dense channel relationships learned by the channel correlation coefficient attention module. The positively correlated channels promote each other, and the negatively correlated channels suppress each other. In the decoder structure, low‐level features selectively emphasize the edge detail information in the feature map through the position attention module. The feature expression of semantic segmentation is improved by feature fusion of the two attention modules to obtain more accurate segmentation results . Finally, to verify the effectiveness of the model, the authors conduct experiments on the PASCAL VOC 2012 and Cityscapes scene analysis benchmark data sets and achieve a mean intersection‐over‐union of 74.92% and 66.63%, respectively.
Dongli Wang, Nanjun Li, Yan Zhou 0003, Jinzhen Mu
IET Image Process.3
2021 Self-attention feature fusion network for semantic segmentation
Yan Zhou 0003, Dongli Wang, Jinzhen Mu, Haibin Zhou
Neurocomputing2
2019 Human action recognition based on multi-mode spatial-temporal feature fusion
Dongli Wang, Yan Zhou 0003
FUSION3
2019 LANet: A Ladder Attention Network for Image Semantic Segmentation
abstract
A Ladder Attention Network (LANet) is proposed to recalibrate feature maps by using channel information in image semantic segmentation. Unlike previous works which use global average pooling to obtain global information of channel map, we use sub-region average pooling (SAP) to extract more abundant channel information. Meanwhile, the attention of low-stage channel is related to that of high-stage channel. In order to make the attention information of low-stage channel better transmit to high-stage features, we add the attention expansion mechanism of low-stage channel. On the other hand, noting the influence of channel number on multi-scale features fusion of Atrous Spatial Pyramid Pooling (ASPP) model in DeepLabv3 network, we allocate different channels to different convolution or pooling operations without increasing the total number of channels. Finally, experimental results on MIT Scene Parsing Benchmark dataset (SceneParse150) are included to demonstrate the effectiveness of the proposed model.
Dongli Wang, Yan Zhou 0003
IECON3
2019 Moving Human Focus Inference Model for Action Recognition
abstract
Spatio-temporal information is crucial to human action recognition. Inspired by the primary visual cortex of the human brain, a new end-to-end action recognition method based on deep learning was proposed to extract pure spatiotemporal information in videos. It is named moving human focus inference model that can capture long-term temporal information with low time costs and eliminates the interference of dynamic background. It obtains the temporal and spatial information in the video from the temporal pathway with the focus block and spatial pathway, respectively. Then the temporal and spatial information are merged as spatio-temporal information by fusion region. Finally, the softmax layer is used to recognize the actions, followed by the bayes inference block that is used to correct the recognition results through the methods of maximum posterior probability inference. The model is evaluated on two challenge datasets: UCF101 and HMDB51. The experimental results show that the approach has better accuracy and generalization ability compared to the state-of-the-art methods.
Dongli Wang, Hexue Xiao, Fang Ou, Yan Zhou 0003
IECON4
2019 A Hybrid Statistics-based Channel Pruning Method for Deep CNNs
abstract
In recent years, depth and width Convolution Neural Networks (CNNs) have achieved significant results in various fields of machine vision. To be honest, deploying the depth models to devices directly is Incredible. Reducing the width of the framework is an effective strategy to make the model more slender. In this paper, we propose a channel pruning method to simultaneously accelerate and compress deep CNNs while maintaining their accuracy. First, a pre-trained CNN model is evaluated layer-by-layer according to the hybrid statistics-based criterion. The negative-score channel and the corresponding filters are discarded. For the performance damaged by channel pruning, knowledge distillation is adopted to fine-tune to achieve stronger generalization ability. Experiments on the ILSVRC-12 benchmark demonstrate the effectiveness of our method when applying to the popular CNN framework. We achieve $3.79\times$ speed-up and $19.1\times$ compression baselines on VGG-16, $3.6\times$ acceleration and $5.5\times$ compression on Darknet (fully convolutional network), both with a little accuracy decrease. Even for modern networks like ResNet, we achieve $2.1\times$ speed-up and $2.2\times$ compression, suffering only 0.4% accuracy loss.
Yan Zhou 0003, Dongli Wang
SMC1
2018 Action Recognition Based on Multi-feature Depth Motion Maps
abstract
Depth motion maps (DMM), containing abundant information on appearance and motion, are captured from the absolute difference between two consecutive depth video sequences. In this paper, each depth frame is first projected onto three orthogonal planes (front, side, top). Then the DMMf, DMMs and DMMt are generated under the three projection view respectively. In order to describe DMM in local and global, histogram of oriented gradient (HOG), local binary patterns (LBP), a local Gist feature description based on a dense grid are computed respectively. Considering the advantages of features fusion and information entropy quantitative evaluation of the Principal Component Analysis (PCA), three descriptors are weighted and fused based on information entropy improved PCA to represent the depth video. A reconstruction error adaptively weighted combination collaborative classifier based on l1-norm and l2-norm is employed for action recognition, the adaptively weights are determined by Entropy Method. Experimental results on MSR Action3D dataset show that the present approach has strong robustness, discriminability and stability.
Dongli Wang, Fang Ou, Yan Zhou 0003
IECON3
2015 Distributed multi-object localisation by consensus on compressive sampling received signal strength fingerprints
abstract
Recent growing interest for location‐based services has created a demand on object localisation approaches with low cost and high accuracy. In this study, the problem of distributed multi‐object localisation using fingerprints of received signal strength (RSS) is addressed by combining average consensus and compressed sensing. First, Bayesian compressed sensing is employed at each agent to recover the sparse index vector from RSS measurements, which are corrupted by noises. It relaxes the requirement on accurate prior position knowledge of beacon nodes and is applicable in non‐line‐of‐sight conditions. Then, average consensus is adopted to compel all agents to reach an agreement on the index vector, and in turn, on the location of objects. By using only one‐hop neighbours’ information, the proposed distributed localisation method is applicable to large‐scale networks. Moreover, the final location of each object is obtainable from each individual agent, which makes the proposed method flexible to the network administration. Experimental results are included to demonstrate the effectiveness of the proposed method.
Dongli Wang, Yan Zhou 0003, Yanhua Wei, Tingrui Pei
IET Commun.2
2014 An indirect Lyapunov approach to the observer-based robust control for fractional-order complex dynamic networks
Yong-Hong Lan, Hai-Bo Gu, Cai-Xue Chen, Yan Zhou 0003
Neurocomputing4
2014 Consensus 3-D bearings-only tracking in switching senor networks
Yan Zhou 0003, Dongli Wang
Signal Process.1
2013 Cross-layer design of quantized-innovation-based target tracking in wireless sensor networks
Yan Zhou 0003, Dongli Wang, Tingrui Pei, Shujuan Tian
FUSION1
2012 Target tracking in wireless sensor networks using adaptive measurement quantization
Yan Zhou 0003, Dongli Wang
Sci. China Inf. Sci.1
2011 Binary tree of posterior probability support vector machines
abstract
Posterior probability support vector machines (PPSVMs) prove robust against noises and outliers and need fewer storage support vectors (SVs). Gonen et al. (2008) extended PPSVMs to a multiclass case by both single-machine and multimachine approaches. However, these extensions suffer from low classification efficiency, high computational burden, and more importantly, unclassifiable regions. To achieve higher classification efficiency and accuracy with fewer SVs, a binary tree of PPSVMs for the multiclass classification problem is proposed in this letter. Moreover, a Fisher ratio separability measure is adopted to determine the tree structure. Several experiments on handwritten recognition datasets are included to illustrate the proposed approach. Specifically, the Fisher ratio separability accelerated binary tree of PPSVMs obtains overall test accuracy, if not higher than, at least comparable to those of other multiclass algorithms, while using significantly fewer SVs and much less test time.
Dongli Wang, Jianguo Zheng, Yan Zhou 0003
J. Zhejiang Univ. Sci. C3
2010 A scalable support vector machine for distributed classification in ad hoc sensor networks
Dongli Wang, Jianguo Zheng, Yan Zhou 0003
Neurocomputing3
2010 Posterior CramÉr-Rao Lower Bounds for Target Tracking in Sensor Networks With Quantized Range-Only Measurements
abstract
We consider the problem of target tracking in a wireless sensor network (WSN) that consists of randomly distributed range-only sensors. Quantized measurements are usually adopted in such a network to attack the problem of limited power supply and communication bandwidth. Assuming that local sensor noises are mutually independent, we derive the posterior Cramer-Rao lower bound (CRLB) on the mean squared error (MSE) of target tracking in WSNs with quantized range-only measurements. Recursion of posterior CRLB on tracking based on both constant velocity (CV) and constant acceleration (CA) model for target dynamics and a general range-only measuring model for local sensors are obtained. Due to the analytical difficulties, particle filter is applied to approximate the theoretical bounds. To illustrate the posterior CRLB, an example on tracking a target with noisy circular trajectories is given.
Yan Zhou 0003, Dongli Wang
IEEE Signal Process. Lett.1
2009 Collaborative Maneuvering Target Tracking in Wireless Sensor Network with Quantized Measurements
abstract
Maneuvering target tracking in wireless sensor network (WSN) with quantized measurements is investigated. The measurement in each local sensor is quantized by uniform quantization scheme and then transmitted to a fusion center (FC). To estimate the state of the target in the FC, the quantized messages are first fused in a weighted average way. Then interactive multiple-model (IMM) scheme using sigma-point Kalman filtering (SPKF) is employed. Focuses are on tradeoff between bandwidth of each sensor and the global tracking accuracy. By performing a change of variable and Lagrange technique, the closed-form solution to the optimization problem for bandwidth scheduling is given, where the mean square error (MSE) incurred by weighted average fusion is minimized subject to a constraint on the total energy consumption. Simulation results reveal that the proposed scheme performs very closely to the clairvoyant IMM-SPKF that based on the analog-amplitude measurements, while obtaining average communication energy saving up to 51.2% and computational burden reduction 31%.
Yan Zhou 0003
SMC1
2009 Distributed Sigma-Point Kalman Filtering for Sensor Networks: Dynamic Consensus Approach
abstract
A scalable Sigma-Point Kalman filter (DSPKF) is proposed for distributed target tracking in a sensor network in this paper. The main idea is to use dynamic consensus strategy to the information form sigma-point Kalman filter (ISPKF) that derived from weighted statistical linearization perspective. Each node estimates the global average information contribution by using local and neighbors' information rather than by the information from all nodes in the network. Therefore, the proposed DSPKF algorithm is completely distributed and applicable to large-scale sensor network. A novel dynamic consensus filter is proposed, and its asymptotical convergence performance and stability are discussed. Finally, a numerical example is given to illustrate the proposed scheme.
Yan Zhou 0003
SMC1