Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xiang Zhang 0006

dblp:91/4353-6 · DBLP profile ↗
← Back
26ranked-venue papers
10as first author
8since 2021 · last 2026
0000-0002-9726-0858ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Segmentation and scene understanding · 54% Video understanding and tracking · 36% Probabilistic and Bayesian machine learning · 11%
Network and information security
1 paper
Privacy and data protection · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
change detection
0.412020
Extended Motion Diffusion-Based Change Detection for Airport Ground Surveillance · IEEE Trans. Image Process. 2020
Computer vision › Video understanding and tracking
background subtraction
0.312017
An Imbalance Compensation Framework for Background Subtraction · IEEE Trans. Multim. 2017
Privacy and data protection
surveillance
0.112020
Extended Motion Diffusion-Based Change Detection for Airport Ground Surveillance · IEEE Trans. Image Process. 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian classification
0.112017
An Imbalance Compensation Framework for Background Subtraction · IEEE Trans. Multim. 2017

Methods — techniques the papers use, named apart from their topics

motion diffusion · 0.9foreground modeling · 0.9background modeling · 0.9spatio-temporal oversampling · 0.3selective downsampling · 0.3bayesian classification · 0.3
YearPublicationVenuePosition
2026 A Multiple Aircraft Tracking Dataset for Airport Traffic Surveillance
abstract
Multiple Aircraft Tracking (MAT) is the foundation of many traffic safety applications in airports. However, experiments show that the state-of-the-art algorithms in Multiple Object Tracking (MOT) deteriorate significantly in the airport scene, and the performance drop can even reach 40%. This is because aircraft and airport scenes possess unique characteristics. For instance, aircraft’s low-compact design causes significant changes in appearance from different angles, while the expansive nature of airports presents challenges in multi-scale issues, particularly at smaller scales. In this paper, we introduce a new dataset Airport Ground Video Surveillance-Tracking24 (AGVS-T24), which could servers as a benchmark to study the challenges in MAT. AGVS-T24 includes 53 airport videos, totaling more than 150,000 frames, and precise manual annotations. AGVS-T24 comprehensively presents various motion patterns of the aircraft, such as takeoff, landing, docking, undocking, taxiing, turning, acceleration, deceleration, etc. AGVS-T24 also contains a variety of MAT challenges, such as appearance change, simultaneous multi-scales, weather and illumination changes, similar appearance between aircraft, as well as tracking in infrared and panoramic modes, and so on. Furthermore, we conduct a simple review of current MOT algorithms, and 20 classic algorithms are selected and tested on AGVS-T24. Finally, we also summarize some principles for designing MAT algorithms. This dataset can be downloaded fromwww.agvs-caac.com/AGVS-T/agvst24.html
Xiang Zhang 0006, Tinyu Li, Runhe Huang
IEEE Trans. Intell. Transp. Syst.1
2024 MBSNet: To Distinguish Motion From Stillness for Airport Traffic Safety
abstract
Background subtraction forms the basis of many safety applications in airport traffic management, such as the visual conflict warning system. However, deep learning methods often mistakenly identify stationary aircraft as foreground, mainly because they prioritize learning appearance over motion features. This means that stationary aircraft with a similar appearance to moving ones are often incorrectly classified as foreground. To address this issue, a Motion-enhanced Background Subtraction Network (MBSNet) is proposed in this paper. MBSNet is designed to focus more on motion information within an encoder-decoder framework. Firstly, a Motion Augmentation Encoder Module (MAEM) is introduced, which generates a clean background frame without foreground from previous frames. This module compares the background frame with the current frame containing moving objects, indirectly enhancing the motion component in the encoded features. Because targets on the airport ground are relatively sparse, MAEM ensures a clean background image. Secondly, a Motion Accumulation Decoder Module (MADM) is designed, which accumulates motion-augmented features from the current frame and past frames based on feature dissimilarity measurement. Since aircraft exhibit consistent motion patterns, such as continuous straight travel with occasional turns, MADM further enhances the motion component in the accumulated feature vector. Finally, MBSNet is evaluated on the AGVS dataset, and our experiments demonstrate the effectiveness of the proposed method for airport background subtraction.
Xiang Zhang 0006, Yingqi Tang, Maozhang Zhou, Celimuge Wu, Zhi Liu 0002
IEEE Trans. Intell. Transp. Syst.1
2022 AGVS: A New Change Detection Dataset for Airport Ground Video Surveillance
abstract
Change detection is the foundation of intelligent video surveillance of the airport ground. However, experiments have shown that change detection algorithms with good performance on traditional datasets (e.g., CDnet2014) perform poorly in airport ground surveillance. The reason is that traditional datasets focus on the diversity of scenarios, while the practical application requires robustness against various changes in a single scene. We posit that the solution to this problem is to establish a unique dataset for airport ground surveillance and develop specific algorithms for this scenario. In this paper, we present an Airport Ground Video Surveillance benchmark (AGVS) for change detection of the airport ground. AGVS includes 25 long videos, amounting to about 100000 frames and accurate ground truth for all frames. Each video contains multiple challenges specific to the airport ground (e.g., haze, camouflage, strip shape, shadow and illumination change, simultaneous multi-scale objects) and various appearance changes of the aircraft). Change detection ground truth is generated by manual annotation. The AGVS benchmark can be downloaded fromhttps://www.agvs-caac.com. Furthermore, we conduct a simple review of current change detection algorithms, both unsupervised or supervised, and then 21 state-of-the-art algorithms are tested and analyzed on the AGVS benchmark. Finally, we conclude with algorithm design principles of change detection for airport ground surveillance.
Xiang Zhang 0006, Shuai Li 0005, Celimuge Wu, Zhi Liu 0002
IEEE Trans. Intell. Transp. Syst.1
2022 ADS-B-Based Spatiotemporal Alignment Network for Airport Video Object Segmentation
abstract
Video object segmentation (VOS) is the fundamental problem of vision-based intelligent transportation, and many VOS algorithms relying on inference from reference masks have been proposed. Due to the inherent defects of the inference strategy and the complex changes of targets, VOS methods that perform well on public datasets are usually ineffective in airport scenarios. We propose a spatiotemporal alignment network (STA-Net) that makes use of Automatic Dependent Surveillance-Broadcast (ADS-B) data as prior information to guide the long-term segmentation of aircraft. ADS-B is an airport-specific signal, which indicates the location of aircraft in real time. Based on ADS-B, we continuously generate new reference masks instead of using previous masks for inference, which greatly reduces the accumulation of inference errors. To achieve this, previous masks of each aircraft are aligned on the temporal domain based on the position information in ADS-B. All temporally-aligned masks are compared, and the one most similar to the current instant is reserved. This mask is both temporally and spatially aligned; hence it is a better reference mask for inference. Aligned masks are updated every time new ADS-B data arrive, so that they can support long-term inference. With the selected mask as a reference, aircraft of interest are segmented within a unified encoder-decoder framework over the long term. Experiments on a benchmark dataset and in a real airport scenario verify the effectiveness of the presented method.
Xiang Zhang 0006, Honggang Wu, Zhi Liu 0002, Celimuge Wu
IEEE Trans. Intell. Transp. Syst.1
2021 Multi-object Tracking Based on Nearest Optimal Template Library
Xiang Zhang 0006, Donghang Chen
ICANN (1)2
2021 Online Multi-Object Tracking with United Siamese Network and Candidate-Refreshing Model
abstract
Current mainstream multi-object tracking (MOT) algorithms aim to maintain the identities of targets by data-association. However, the accuracy of multi-object tracking would be affected by the unreliability results of detectors. Simultaneously, the separation of motion module and appearance module cannot fully exploit affinity features that will increase computational complexity. To address these issues, we propose a multi-task tracking framework including the United Siamese Network (USN) and Candidate-Refreshing (CR) model. The USN integrates motion affinity and appearance affinity into an end-to-end network. Such design can pay attention to joint learning of similarities between motion patterns and appearance features and enhance the ability to distinguish similar targets in complex environments. The CR model aims to combine detection candidates with tracking candidates, and compensate for unreliable detection results by scoring and regressing the bounding boxes. In addition, we equip our framework with a tracklet confidence function, which is based on the spatialtemporal information. This function can determine the vitality of unmatched tracklets and further alleviate the influence of detectors. Experiments on the MOT16 and MOT17 benchmark datasets show that our framework has achieved outstanding performance on online MOT algorithms.
Donghang Chen, Xiang Zhang 0006, Yingqi Tang, Shaozhi Wu
IJCNN2
2021 Motion-augmented Change Detection for Video Surveillance
abstract
The main task of change detection is to segment moving objects from the background. Recently, the deep learning-based change detection method has attracted much attention, but it tends to classify stationary objects as foreground. The reason for this phenomenon is that it identifies objects mainly based on appearance features while the same object has a similar appearance both in motion and stationary state. To address this problem, we propose a novel network called Motion-augmented Change Detection Network (MCDNet) to distinguish moving objects using both motion and appearance information. To achieve this purpose, firstly, we introduce a Motion-augmented Background Model (MBM) to simulate scene background without any foreground objects which can be dynamically updated by the predicted mask. In this way, the motion information can be implicitly highlighted by comparing the current frame and the background. Secondly, we design an Attention Memory Module (AMM) to store past features and use it to guide the segmentation of the current frame, facilitating the extraction of motion and appearance features. Experiments on two challenging public benchmarks (i.e. AGVS1and CDnet2014) demonstrate that our proposed method achieves compelling performance against state-of-the-art methods.
Yingqi Tang, Xiang Zhang 0006, Donghang Chen, Zhizhuo Zhang, Haifei Yu
MMSP2
2021 A pruning method based on the measurement of feature extraction ability
Honggang Wu, Xiang Zhang 0006
Mach. Vis. Appl.3
2020 Mask-Ranking Network for Semi-supervised Video Object Segmentation
Xiang Zhang 0006, Yingqi Tang
ACCV (1)2
2020 Online Multiple Object Tracking Using Single Object Tracker and Maximum Weight Clique Graph
abstract
Tracking multiple objects is a challenging task in time-critical video analysis systems. In the popular tracking-by-detection framework, the core problems of a tracker are the quality of the employed input detections and the effectiveness of the data association. Towards this end, we propose a multiple object tracking method which employs a single object tracker to improve the results of unreliable detection and data association simultaneously. Besides, we utilize maximum weight clique graph algorithm to handle the optimal assignment in an online mode. In our method, a robust single object tracker is used to connect previous tracked objects to tackle the current noise detection and improve the data association as a motion cue. Furthermore, we use person re-identification network to learn the historical appearances of the tracklets in order to promote the tracker's identification ability. We conduct extensive experiments on the MOT benchmark to demonstrate the effectiveness of our tracker.
Xiang Zhang 0006, Yexin Li
MMSP2
2020 Robust block tensor principal component analysis
Lanlan Feng, Yipeng Liu 0001, Longxi Chen, Xiang Zhang 0006, Ce Zhu
Signal Process.4
2020 Extended Motion Diffusion-Based Change Detection for Airport Ground Surveillance
abstract
Change detection in airport ground is important for airport security. Due to the particularity of ground environment, e.g. haze and camouflage, airport ground change detection is generally incomplete. If an incomplete detection is used as reference for the detection in subsequent frames, it may result in noticeable detection defects across the frames. In this paper, extended motion diffusion (EMD) is proposed to address the problems. The core idea of the EMD is to design a novel model insensitive to incomplete detection. Firstly the one-to-many correspondence in traditional motion diffusion is extended in the prediction step of EMD to build up correspondence from incomplete detection to intact objects. Prior information, e.g. aircraft motion prior and ground structure prior, is employed in the development of the correspondence. Then based on the correspondence a number of new samples are synthesized and filtered in the identification step of the EMD to compensate possible detection defects. Finally, the reserved samples are collected to train a foreground model, which is used in conjunction with another background model for classification. The proposed method is verified based on the Airport Ground Video Surveillance (AGVS) benchmark. Experimental results show effectiveness of the proposed algorithm in dealing with haze and camouflage.
Xiang Zhang 0006, Honggang Wu, Celimuge Wu
IEEE Trans. Image Process.1
2019 A Pruning Method Based on Feature Abstraction Capability of Filters
Xiang Zhang 0006, Ce Zhu
ICIG (2)2
2018 Video analytical coding: When video coding meets video analysis
Ce Zhu, Min Mao, Fangliang Song, Frédéric Dufaux, Xiang Zhang 0006
Signal Process. Image Commun.6
2017 Contouring error vector and cross-coupled control of multi-axis servo system
abstract
The contouring error and cross-coupled gains calculation have always been the critical issues in the application of cross-coupled control. Traditionally, the linear approximation and circular approximation are widely used to determine the contouring error and cross-coupled gains. However, for linear approximation and circular approximation, the contouring error and cross-coupled gains are calculated sophisticatedly, especially in three-dimensional applications. In this paper, a contouring error vector is established under task coordinate frame, then the contouring error and cross-coupled gains can be easily obtained based on the magnitude and orientation of the contouring error vector. The experimental results on a three-axis CNC machine indicate the proposed approach simplifies the calculation of contouring error and cross-coupled gains.
Xiang Zhang 0006, Yunjiang Lou
IROS2
2017 Analytical distortion aware video coding for computer based video analysis
abstract
With the development of artificial intelligence, more and more multimedia applications for various tasks have emerged in our daily life. Meanwhile, as one of the main information sources of the applications, a huge amount of video data has been being generated by portable or mounted cameras in daily basis for varying purposes including surveillance, in which case we may need computers to "watch" videos to save labor cost. However, most video coding standards are designed for the highest human perceptual quality given a bit rate by minimizing a fidelity cost function (e.g., mean squared error, MSE), assuming the content will be consumed by human beings. In view of the above considerations, this paper proposes a new rate-analytical-distortion optimization method (RADO) for video analysis. Specifically, we consider moving object detection as the analysis task. Accordingly, we develop a novel rate analytical distortion (RAD) model for video coding, where the analytical distortion is related to the object detection performance expressed in terms of F-measure. As shown in the experimental results, the performance of the video analysis task can be significantly improved (up to 40% reduction of analytical distortion) with a slight bit rate increase.
Ce Zhu, Min Mao, Fangliang Song, Frédéric Dufaux, Xiang Zhang 0006
MMSP6
2017 A Bayesian Approach to Camouflaged Moving Object Detection
abstract
Moving object detection is about foreground and background separation based on motion detection. Detecting moving objects from similarly colored background (known as camouflage problem) has been a long-standing open question in this field. Discriminative modeling (DM), which focuses on enhancing the performance to distinguish foreground from background with discriminative features and well-designed classifiers, has been widely used for moving object detection. However, DM may tend to fail when encountering the camouflage problem, as the class separability in camouflaged areas is generally poor. In this paper, we propose a new strategy, camouflage modeling (CM), to identify camouflaged foreground pixels. In view of the fact that camouflage involves both foreground and background, we need to model both the background and the foreground, and compare them in a well-designed way in camouflage detection. Specifically, we develop a global model for the background, and an integration of global and local models for the foreground, respectively. Based on both background and foreground models, we introduce a factor to measure the degree of camouflage, and further identify truly camouflaged areas. In view of the fact that a moving object is usually composed of both camouflaged and noncamouflaged areas, CM and DM are fused in a Bayesian framework to perform complete object detection. Experiments are conducted on testing sequences to demonstrate the effectiveness of the proposed algorithm.
Xiang Zhang 0006, Ce Zhu, Yipeng Liu 0001, Mao Ye 0001
IEEE Trans. Circuits Syst. Video Technol.1
2017 An Imbalance Compensation Framework for Background Subtraction
abstract
Class imbalance refers to the instance where the number of training samples for the majority classes is far more than that of the minority classes (relative imbalance), and the quality of training samples for the minority classes is inferior to that of the majority classes (absolute imbalance), which are further complicated by other imbalance factors, e.g., data overlapping. Video background subtraction aims to classify each pixel into two classes: foreground and background. This paper first reveals that background subtraction is a class imbalance problem, where the foreground and background are the minority and majority classes, respectively. By exploring spatial and temporal correlation inherent in video data, we present an imbalance compensation framework for background subtraction, which consists of two sequential modules, imbalance-compensated bilayer modeling, and imbalance-compensated Bayesian classification. In the first module, spatio-temporal oversampling (SOS) and selective downsampling (SDS) are proposed to compensate the imbalance at data level. SOS attempts to synthesize representative samples appended to the minority sample set, while SDS selectively deletes a number of majority samples in data overlapping areas. The rebalanced samples are then used to learn a bilayer model. In the second module, novel cost functions are proposed to compensate the effect of class imbalance at algorithm level. The cost functions are based on imbalance measurement, and used to construct the prior term in the Bayesian classification scheme. Experiments are conducted on public databases to demonstrate the effectiveness of the proposed method.
Xiang Zhang 0006, Ce Zhu, Honggang Wu, Zhi Liu 0003, Yuanyuan Xu 0001
IEEE Trans. Multim.1
2015 Spatiotemporal saliency detection based on superpixel-level trajectory
Zhi Liu 0003, Xiang Zhang 0006, Olivier Le Meur, Liquan Shen
Signal Process. Image Commun.3
2014 Co-saliency detection based on region-level fusion and pixel-level refinement
abstract
This paper addresses the problem of co-saliency detection, which aims to identify the common salient objects in a set of images and is important for many applications such as object co-segmentation and co-recognition. First, the segmentation driven low-rank matrix recovery model is used for intra saliency detection in each individual image of the image set, to highlight the regions whose features are sparse in each image. Then, a region-level fusion method, which exploits inter-region dissimilarities on color histograms and global consistency of regions over the image set, adjusts the intra saliency maps to obtain the region-level co-saliency maps, which can highlight co-salient object regions and suppress irrelevant regions. Finally, a pixel-level refinement method, which integrates color-spatial similarity between pixel and region with image border connectivity based object prior, generates the pixel-level co-saliency maps with better quality. Extensive experiments on two benchmark datasets demonstrate that the proposed co-saliency model consistently outperforms the state-of-the-art co-saliency models in both subjective and objective evaluation.
Zhi Liu 0003, Wenbin Zou, Xiang Zhang 0006, Olivier Le Meur
ICME4
2014 Statistical background subtraction based on imbalanced learning
abstract
In this paper, we study the class imbalance problem in statistical background subtraction. Firstly, we discuss the imbalance essence in background subtraction, and conclude that foreground and background are inherently imbalanced. Secondly, following the imbalanced learning strategy in machine learning, we present a spatio-temporal over-sampling method to resolve the class imbalance in background subtraction. Our method densely generate synthesized foreground samples in compact 3D spatio-temporal domain. Those generated samples could reduce the imbalance level between foreground and background from both quantity and quality, and therefore contribute to improvement of detection performance. We also define a new index to measure the change of imbalance level during over-sampling. Experiments are conducted on public datasets to demonstrate the effectiveness of our method.
Xiang Zhang 0006, Xu Zhao 0001
ICME1
2014 Superpixel-Based Spatiotemporal Saliency Detection
abstract
This paper proposes a superpixel-based spatiotemporal saliency model for saliency detection in videos. Based on the superpixel representation of video frames, motion histograms and color histograms are extracted at the superpixel level as local features and frame level as global features. Then, superpixel-level temporal saliency is measured by integrating motion distinctiveness of superpixels with a scheme of temporal saliency prediction and adjustment, and superpixel-level spatial saliency is measured by evaluating global contrast and spatial sparsity of superpixels. Finally, a pixel-level saliency derivation method is used to generate pixel-level temporal and spatial saliency maps, and an adaptive fusion method is exploited to integrate them into the spatiotemporal saliency map. Experimental results on two public datasets demonstrate that the proposed model outperforms six state-of-the-art spatiotemporal saliency models in terms of both saliency detection and human fixation prediction.
Zhi Liu 0003, Xiang Zhang 0006, Shuhua Luo, Olivier Le Meur
IEEE Trans. Circuits Syst. Video Technol.2
2013 Cost-sensitive background subtraction
abstract
Foreground and background are treated without distinction at classification stage in most background subtraction algorithms. However, correct classification of foreground is the primary requirement, and thus misclassification costs of the two classes should be different. Based on this fact, we present a new method to introduce cost sensitivity into background subtraction, where a cost matrix is created to represent the costs of misclassification. Some items in the cost matrix are not constants, but functions of foreground occurence at each pixel location. By the use of such non-constant costs, detection rate of foreground is improved while increase of false alarms is prevented at the same time. Experiments demonstrate the effectiveness of the proposed algorithm.
Xiang Zhang 0006, Jian Cheng 0001, Zhi Liu 0003, Jie Yang 0002
ICIP1
2012 Region Diversity Maximization for Salient Object Detection
abstract
Salient object detection is an important technique for many content-based applications, but it becomes a challenging work when handling the cluttered saliency maps, which cannot completely highlight salient object regions and cannot suppress background regions. In this letter, we propose a novel approach to detect salient object from saliency map without manually setting any parameters. Region diversity maximization is used as the objective function to direct the object detection, and the optimal window for locating the salient object is obtained using an efficient iterative search scheme. Experimental results on different saliency maps demonstrate the overall better detection performance and computational efficiency of our approach.
Zhi Liu 0003, Huan Du, Xiang Zhang 0006, Liquan Shen
IEEE Signal Process. Lett.4
2011 Interactive object segmentation using iterative adjustable graph cut
abstract
Interactive object segmentation is widely used for extracting any user-interested objects from natural images. A common problem with many interactive segmentation approaches is that the object segmentation quality is degraded due to inaccurate object/background seeds provided by the user. This paper proposes an iterative adjustable graph cut to efficiently solve this problem. First, object/background seeds are initialized based on the object segmentation result obtained with the user-specified scribbles as the interactive input. Then, an iterative seed adjustment scheme is exploited to correct inaccurate seeds and extract new suitable seeds via graph cut, in which the balancing weight between energy terms are adaptively updated to protect stable seeds and speedup the iteration process. Finally, suitable seeds are obtained and graph cut is used to segment the objects. Experimental results demonstrate the better segmentation performance of our approach even if user provides rather rough seeds.
Zhi Liu 0003, Yinzhu Xue, Xiang Zhang 0006
VCIP4
2009 The analysis of the color similarity problem in moving object detection
Xiang Zhang 0006, Jie Yang 0002
Signal Process.1