EDBT 2026 Demo / reviewers in the wild / expert
Zeyu Fu
dblp:195/9236
· DBLP profile ↗
27ranked-venue papers
7as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MultiHateLoc: Towards Temporal Localisation of Multimodal Hate Content in Online Videos
Qiyue Sun, Tailin Chen, Jiangbei Yue, Jianbo Jiao, Zeyu Fu |
WWW | 7 |
| 2026 | Cross-modal progressive modeling for neuro-visual representation learningabstractNeural decoding from scalp signals requires models that respect spatial, temporal, and spectral structure while leveraging strong visual priors. In this paper, we introduce CFT-NET for disentangled neural visual representation together with a progressive visual–semantic adaptation (PVSA) framework that aligns EEG embeddings to pretrained visual backbones under a contrastive objective followed by pairwise matching. CFT-NET integrates Frequency-Separated Weights (FSW), Spatial-Context Aggregation (SCA), and Adaptive Temporal Filtering (ATF) to explicitly extract spectral, spatial, and temporal factors. PVSA consists of an instance-guided visual encoder and a visual-guided semantic decoder linked by cross attention, enabling fine-grained neuro–image interaction. On THINGS-EEG and THINGS-MEG dataset, the approach consistently outperforms state-of-the-art baselines in both subject-dependent and subject-independent zero-shot classification. By aligning model architecture with visual cognition principles and coupling it to strong visual priors, our methods narrows the gap between neural activity and visual cognition, provides novel cross-modal neural decoding method that achieves competitive performance against recent state-of-the-art baselines. Jiyao Pu, Kaili Sun, Zeyu Fu, Haoran Duan 0001, Yang Long 0001 |
Neurocomputing | 4 |
| 2026 | Probing 3D anomalies via multi-view registration and dual-residual analysis
Yuxing Yang, Zeyu Fu, Liewei Wang, Siyue Yu, Jimin Xiao |
Neurocomputing | 2 |
| 2025 | A Black-Box Evaluation Framework for Semantic Robustness in Bird's Eye View DetectionabstractCamera-based Bird's Eye View (BEV) perception models receive increasing attention for their crucial role in autonomous driving, a domain where concerns about the robustness and reliability of deep learning have been raised. While only a few works have investigated the effects of randomly generated semantic perturbations, aka natural corruptions, on the multi-view BEV detection task, we develop a black-box robustness evaluation framework that adversarially optimises three common semantic perturbations: geometric transformation, colour shifting, and motion blur, to deceive BEV models, serving as the first approach in this emerging field. To address the challenge posed by optimising the semantic perturbation, we design a smoothed, distance-based surrogate function to replace the mAP metric and introduce SimpleDIRECT, a deterministic optimisation algorithm that utilises observed slopes to guide the optimisation process. By comparing with randomised perturbation and two optimisation baselines, we demonstrate the effectiveness of the proposed framework. Additionally, we provide a benchmark on the semantic robustness of ten recent BEV models. The results reveal that PolarFormer, which emphasises geometric information from multi-view images, exhibits the highest robustness, whereas BEVDet is fully compromised, with its precision reduced to zero. Yanghao Zhang, Xiangyu Yin 0001, Zeyu Fu, Xiaowei Huang 0001, Wenjie Ruan |
AAAI | 5 |
| 2025 | Edge-Intelligent Unmanned Aerial Vehicle Oil Tank Inspection Method Based on GWO-PSO SchedulingabstractThis paper proposes a hierarchical task scheduling and adaptive resource management framework for intelligent unmanned aerial vehicle (UAV) systems to address the complex demands of oil tank monitoring tasks. A hybrid path planning strategy is adopted: at the global level, an improved Grey Wolf Optimization algorithm is used for task allocation and initial path generation; in local emergency scenarios, an enhanced Particle Swarm Optimization algorithm is employed to generate refined schedules. The system architecture supports dynamic task distribution and efficient resource utilization within UAV swarms, while the embedded implementation enables real-time execution in edge environments. Experimental evaluations demonstrate the effectiveness of the proposed method in enhancing scheduling flexibility, response efficiency, and robustness in complex inspection scenarios. Xiaokang Yin 0001, Cai Luo, Chunbo Luo, Lei Liu 0003, Zeyu Fu |
HPCC | 6 |
| 2025 | Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation
Xujiong Ye, Yanda Meng, Zeyu Fu |
MICCAI (10) | 5 |
| 2025 | PySimPace v2.0: An Easy-to-Use Simulation Tool with Machine Learning Pipelines for Realistic MRI Motion Artifact GenerationabstractMotion artifacts in structural and functional magnetic resonance imaging (MRI) pose a significant challenge for both clinical use and machine learning (ML)-based image analysis. Existing ML approaches for artifact correction require paired clean and corrupted datasets, which are difficult to acquire. We present py-simpace, an open-source, pip-installable MRI motion artifact simulation toolkit with native ML integration. py-simpace supports structural MRI and functional MRI (fMRI) simulation, offering configurable k-space and image-space motion, ghosting, Gibbs ringing, and physiological noise. It provides an end-to-end pipeline with a ready-to-use PyTorch Dataset interface for ML training. We describe the design of py-simpace v2.0, compare it with existing tools, and demonstrate its utility for robust artifact correction model development. Snehil Kumar, Neil Vaughan, Zeyu Fu, Heather Wilson |
ACM Multimedia | 3 |
| 2025 | DeHate: A Holistic Hateful Video Dataset for Explicit and Implicit Hate DetectionabstractHate speech poses a persistent threat to society, causing profound harm to both individuals and communities. Detecting such content is essential for promoting safer and more inclusive environments. While previous research has primarily focused on text-based or image-based hate speech detection, video-based hate detection remains relatively underexplored. A key barrier is the limited availability of high-quality video datasets. Existing hateful video datasets are typically limited in scale, diversity, and annotation depth, often labeling hateful content without further distinguishing between explicit and implicit forms. In this work, we present DeHate, which, to the best of our knowledge, is the largest hateful video dataset to date. DeHate comprises 6689 videos collected from two platforms and spanning six social groups. Each video is annotated with fine-grained labels that differentiate explicit, implicit, and non-hateful content, along with segment-level localization of hate, identification of contributing modalities, and specification of the targeted groups. Through detailed analysis of annotated videos across platforms, we reveal distinct patterns in how hateful content is conveyed, offering a comprehensive comparison between explicit and implicit hate in terms of their prevalence and characteristics. Furthermore, we benchmark state-of-the-art models, including both uni-modal and multi-modal architectures, and identify persistent challenges in detecting subtle and context-dependent forms of hate. Our findings highlight the importance of holistic and fine-grained hateful video datasets for advancing research in hate speech detection. Disclaimer: This paper contains sensitive content that may be disturbing to some readers. Tailin Chen, Jiangbei Yue, Jianbo Jiao, Zeyu Fu |
ACM Multimedia | 7 |
| 2025 | Pose-oriented scene-adaptive matching for abnormal event detectionabstractFor intelligent surveillance systems, abnormal event detection automatically analyses surveillance video sequences and detects abnormal objects or unusual human actions at the frame level. Due to the lack of labelled data, most approaches are semi-supervised based on reconstruction or prediction methods. However, these methods may not generalize well to unseen scene contexts. To address this issue, we present a novel self and mutual scene-adaptive matching method for abnormal event detection. In the framework, we propose synergistic pose estimation and object detection, which effectively integrates human pose and object detection information to improve pose estimation accuracy. Then, the poses are resized to reduce the spatial distance between the source and target domains. The improved pose sequences are further fed into a spatio-temporal graph convolutional network to extract the geometric features. Finally, the features are embedded in a clustering layer to classify action types and compute normality scores. The training data is taken from the training part of common video anomaly detection datasets: UCSD PED1 & PED2, CHUK Avenue, and ShanghaiTech Campus. The proposed framework is evaluated on video sequences with unseen scene contexts in the UCSD PED2 and ShanghaiTech Campus datasets. The detection accuracy and efficiency are also evaluated in detail, and the proposed method for abnormal event detection achieves the highest AUC performance, 84.6%, on the ShanghaiTech Campus dataset and relatively high AUC performance, 96.9% and 74.8%, on UCSD PED2 & PED1 datasets. Compared with other state-of-the-art works, the performance analysis and results confirm the robustness and effectiveness of our proposed framework for cross-scene abnormal event detection. Yuxing Yang, Leiyu Xie, Zeyu Fu, Jiawei Yan, Syed M. Naqvi |
Neurocomputing | 3 |
| 2025 | OSDMamba: Enhancing Oil Spill Detection From Remote Sensing Images Using Selective State-Space ModelabstractSemantic segmentation is commonly used for Oil Spill Detection (OSD) in remote sensing images. However, the limited availability of labelled oil spill samples and class imbalance present significant challenges that can reduce detection accuracy. Furthermore, most existing methods, which rely on convolutional neural networks (CNNs), struggle to detect small oil spill areas due to their limited receptive fields and inability to effectively capture global contextual information. This study explores the potential of State-Space Models (SSMs), particularly Mamba, to overcome these limitations, building on their recent success in vision applications. We propose OSDMamba, the first Mambabased architecture specifically designed for oil spill detection. OSDMamba leverages Mamba’s selective scanning mechanism to effectively expand the model’s receptive field while preserving critical details. Moreover, we designed an asymmetric decoder incorporating ConvSSM and deep supervision to strengthen multiscale feature fusion, thereby enhancing the model’s sensitivity to minority class samples. Experimental results show that the proposed OSDMamba achieves state-of-the-art performance, yielding improvements of 8.9% and 11.8% in OSD across two publicly available datasets. The source codes will be made publicly available at https://github.com/Chenshuaiyu1120/Oil-Spill-detection. Shuaiyu Chen, Peng Ren 0001, Chunbo Luo, Zeyu Fu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Position and Orientation Aware One-Shot Learning for Medical Action Recognition From Signal DataabstractIn this article, we propose a position and orientation-aware one-shot learning framework for medical action recognition from signal data. The proposed framework comprises two stages and each stage includes signal-level image generation (SIG), cross-attention (CsA), and dynamic time warping (DTW) modules and the information fusion between the proposed privacy-preserved position and orientation features. The proposed SIG method aims to transform the raw skeleton data into privacy-preserved features for training. The CsA module is developed to guide the network in reducing medical action recognition bias and more focusing on important human body parts for each specific action, aimed at addressing similar medical action related issues. Moreover, the DTW module is employed to minimize temporal mismatching between instances and further improve model performance. Furthermore, the proposed privacy-preserved orientation-level features are utilized to assist the position-level features in both of the two stages for enhancing medical action recognition performance. Extensive experimental results on the widely-used and well-known NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD datasets all demonstrate the effectiveness of the proposed method, which outperforms the other state-of-the-art methods with general dataset partitioning by 2.7%, 6.2% and 4.1%, respectively. Leiyu Xie, Yuxing Yang, Zeyu Fu, Syed M. Naqvi |
IEEE Trans. Multim. | 3 |
| 2024 | Robust implementation of foreground extraction and vessel segmentation for X-ray coronary angiography image sequence
Zeyu Fu, Zhuang Fu, Chenzhuo Lu, Jian Fei |
Pattern Recognit. | 1 |
| 2023 | One-Shot Medical Action Recognition With A Cross-Attention Mechanism And Dynamic Time WarpingabstractIn this paper, we address the classification of medical actions with only one single sample by developing a novel one-shot learning framework which contains both cross-attention and dynamic time warping (DTW) modules. To be concrete, we firstly transform the raw skeleton sequence into the signal-level image representation. We exploit a metric learning approach, which is the prototypical network for the proposed one-shot learning framework and choose the residual network (ResNet18) as the backbone which is widely used in recent years. Cross-attention is applied for guiding the network to focus on the more important joints from each specific action. The cross-attention mechanism that applies between the support and query set will be adapted for mining and matching the relationships with the human body. Furthermore, a DTW module is introduced to mitigate the temporal information mismatching issue between the actions from the support and query sets. The experimental results on the NTU RGB+D 120 dataset demonstrate the effectiveness of our proposed approach and the improved performance compared to the baseline approach. The code of this work is available at1. Leiyu Xie, Yuxing Yang, Zeyu Fu, Syed M. Naqvi |
ICASSP | 3 |
| 2023 | Self-adaptive Adversarial Training for Robust Medical Segmentation
Zeyu Fu, Yanghao Zhang, Wenjie Ruan |
MICCAI (3) | 2 |
| 2023 | Abnormal event detection for video surveillance using an enhanced two-stream fusion methodabstractAbnormal event detection is a critical component of intelligent surveillance systems, focusing on identifying abnormal objects or unusual human behaviours in video sequences. However, conventional methods struggle due to the scarcity of labelled data. Existing solutions typically train on normal data, establish boundaries for regular events, and identify outliers during testing. These approaches are often inadequate as they do not efficiently leverage the geometry and image texture information, and they lack a specific focus on different types of abnormal events. This paper introduces a novel two-stream fusion algorithm for abnormal event detection to address these diverse abnormal events better. We first extract the object, pose, and optical flow features. Then, the object and pose information is combined early on to eliminate occluded pose graphs. The trusted pose graphs are fed into a Spatio-Temporal Graph Convolutional Network (ST-GCN) to detect abnormal behaviours. Simultaneously, we propose a video prediction framework that identifies abnormal frames by measuring the difference between predicted and ground truth frames. Lastly, we execute a decision-level fusion between the classification and prediction streams to achieve the final results. Our results on the UCSD PED1 dataset indicate the enhanced performance of the fusion model for various abnormal events. Furthermore, experimental results on the UCSD PED2 dataset and the ShanghaiTech campus dataset underscore our approach’s effectiveness compared to other related works. Yuxing Yang, Zeyu Fu, Syed M. Naqvi |
Neurocomputing | 2 |
| 2023 | A Machine Learning Method for Automated Description and Workflow Analysis of First Trimester Ultrasound ScansabstractObstetric ultrasound assessment of fetal anatomy in the first trimester of pregnancy is one of the less explored fields in obstetric sonography because of the paucity of guidelines on anatomical screening and availability of data. This paper, for the first time, examines imaging proficiency and practices of first trimester ultrasound scanning through analysis of full-length ultrasound video scans. Findings from this study provide insights to inform the development of more effective user-machine interfaces, of targeted assistive technologies, as well as improvements in workflow protocols for first trimester scanning. Specifically, this paper presents an automated framework to model operator clinical workflow from full-length routine first-trimester fetal ultrasound scan videos. The 2D+t convolutional neural network-based architecture proposed for video annotation incorporates transfer learning and spatio-temporal (2D+t) modelling to automatically partition an ultrasound video into semantically meaningful temporal segments based on the fetal anatomy detected in the video. The model results in a cross-validation A1 accuracy of 96.10% , F1=0.95 , precision =0.94 and recall =0.95 . Automated semantic partitioning of unlabelled video scans (n=250) achieves a high correlation with expert annotations ( ρ = 0.95, p=0.06 ). Clinical workflow patterns, operator skill and its variability can be derived from the resulting representation using the detected anatomy labels, order, and distribution. It is shown that nuchal translucency (NT) is the toughest standard plane to acquire and most operators struggle to localize high-quality frames. Furthermore, it is found that newly qualified operators spend 25.56% more time on key biometry tasks than experienced operators. Robail Yasrab, Zeyu Fu, He Zhao 0002, Lok Hin Lee, Harshita Sharma, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
IEEE Trans. Medical Imaging | 2 |
| 2022 | A Two-Stream Information Fusion Approach to Abnormal Event Detection in VideoabstractHuman abnormal activity detection for automatic surveillance systems is to detect abnormal objects and human behaviours in videos. In this paper, we propose to explicitly address different kinds of abnormal events by developing a two-stream fusion approach that integrates both geometry and image texture information. To be concrete, we firstly propose to utilize an object detector to divide the abnormal events into two catalogues: abnormal human behaviors and abnormal objects. For the detection of abnormal human behaviours, we exploit a spatial-temporal graph convolutional network (ST-GCN) which considers both spatial and temporal domains to capture the geometrical features from human pose graphs. The extracted geometric feature embeddings are further adapted with a clustering step to cluster the temporal graphs and output normality scores. For the detection of abnormal objects, the obtained from the object detector are reused to assist with generating normality scores of possible anomalies. Finally, a late fusion is performed to integrate normality scores from both screams for final decision. The experimental results on the datasets of UCSD PED2 and ShanghaiTech Campus demonstrate the effectiveness of our proposed approach and the improved performance compared to other state-of-the-art approaches. Yuxing Yang, Zeyu Fu, Syed M. Naqvi |
ICASSP | 2 |
| 2022 | Facial Anatomical Landmark Detection Using Regularized Transfer Learning With Application to Fetal Alcohol Syndrome RecognitionabstractFetal alcohol syndrome (FAS) caused by prenatal alcohol exposure can result in a series of cranio-facial anomalies, and behavioral and neurocognitive problems. Current diagnosis of FAS is typically done by identifying a set of facial characteristics, which are often obtained by manual examination. Anatomical landmark detection, which provides rich geometric information, is important to detect the presence of FAS associated facial anomalies. This imaging application is characterized by large variations in data appearance and limited availability of labeled data. Current deep learning-based heatmap regression methods designed for facial landmark detection in natural images assume availability of large datasets and are therefore not well-suited for this application. To address this restriction, we develop a new regularized transfer learning approach that exploits the knowledge of a network learned on large facial recognition datasets. In contrast to standard transfer learning which focuses on adjusting the pre-trained weights, the proposed learning approach regularizes the model behavior. It explicitly reuses the rich visual semantics of a domain-similar source model on the target task data as an additional supervisory signal for regularizing landmark detection optimization. Specifically, we develop four regularization constraints for the proposed transfer learning, including constraining the feature outputs from classification and intermediate layers, as well as matching activation attention maps in both spatial and channel levels. Experimental evaluation on a collected clinical imaging dataset demonstrate that the proposed approach can effectively improve model generalizability under limited training samples, and is advantageous to other approaches in the literature. Zeyu Fu, Jianbo Jiao, Michael Suttie, J. Alison Noble |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Video Anomaly Detection for Surveillance Based on Effective Frame Area
Yuxing Yang, Yang Xian, Zeyu Fu, Syed M. Naqvi |
FUSION | 3 |
| 2020 | 2D Pose-Based Real-Time Human Action Recognition With Occlusion-HandlingabstractHuman Action Recognition (HAR) for CCTV-oriented applications is still a challenging problem. Real-world scenarios HAR implementations is difficult because of the gap between Deep Learning data requirements and what the CCTV-based frameworks can offer in terms of data recording equipments. We propose to reduce this gap by exploiting human poses provided by the OpenPose, which has been already proven to be an effective detector in CCTV-like recordings for tracking applications. Therefore, in this work, we first propose ActionXPose: a novel 2D pose-based approach for pose-level HAR. ActionXPose extracts low- and high-level features from body poses which are provided to a Long Short-Term Memory Neural Network and a 1D Convolutional Neural Network for the classification. We also provide a new dataset, named ISLD, for realistic pose-level HAR in a CCTV-like environment, recorded in the Intelligent Sensing Lab. ActionXPose is extensively tested on ISLD under multiple experimental settings, e.g. Dataset Augmentation and Cross-Dataset setting, as well as revising other existing datasets for HAR. ActionXPose achieves state-of-the-art performance in terms of accuracy, very high robustness to occlusions and missing data, and promising results for practical implementation in real-world applications. Federico Angelini, Zeyu Fu, Yang Long 0001, Ling Shao 0001, Syed M. Naqvi |
IEEE Trans. Multim. | 2 |
| 2019 | Enhanced Detection Reliability for Human Tracking Based Video Analytics
Zeyu Fu, Syed M. Naqvi |
FUSION | 1 |
| 2019 | Enhanced pooling method for convolutional neural networks based on optimal search theoryabstractTo obtain the best pooling effect and higher accuracy in image recognition, an improved method based on optimal search theory for the pooling layer of convolutional neural networks (CNNs) is proposed. The purpose is to solve the problems of the traditional pooling method, namely that it is too simplistic and it is difficult to extract effective features. The basic principle and network structure of CNN are introduced in the study. A new optimum‐pooling method is proposed, and the authors study how to obtain the maximum probability to detect the target function under the constrained condition. Comparison experiments of different pooling methods are performed on three widely used datasets: LFW, CIFAR‐10, and ImageNet. The experimental results show that the proposed method has the characteristics of more effective feature extraction and wide adaptability, and leads to higher accuracy and lower error rate in image recognition. Zeyu Fu, Syed M. Naqvi, Jonathon A. Chambers |
IET Image Process. | 3 |
| 2019 | Multi-Level Cooperative Fusion of GM-PHD Filters for Online Multiple Human TrackingabstractIn this paper, we propose a multi-level cooperative fusion approach to address the online multiple human tracking problem in a Gaussian mixture probability hypothesis density (GM-PHD) filter framework. The proposed fusion approach consists essentially of three steps. First, we integrate two human detectors with different characteristics (full-body and body-parts), and investigate their complementary benefits for tracking multiple targets. For each detector domain, we then propose a novel discriminative correlation matching model, and fuse it with spatio-temporal information to address ambiguous identity association in the GM-PHD filter. Finally, we develop a robust fusion center with virtual and real zones to make a global decision based on preliminary candidate targets generated by each detector. This center also mitigates the sensitivity of missed detections in the generalized covariance intersection fusion process, thereby improving the fusion performance and tracking consistency. Experiments on the MOTChallenge Benchmark demonstrate that the proposed method achieves improved performance over other state-of-the-art RFS-based tracking methods. Zeyu Fu, Federico Angelini, Jonathon A. Chambers, Syed M. Naqvi |
IEEE Trans. Multim. | 1 |
| 2018 | Collaborative Detector Fusion of Data-Driven PHD Filter for Online Multiple Human TrackingabstractThe use of multiple data sources (measurements) has been recently demonstrated to improve the accuracy and reliability of a tracking system as it is capable of providing redundancy in different aspects, and also eliminating interferences of individual sources. This paper focuses on addressing the multiple human tracking problem from a multi-detector approach. This approach integrates two detectors with different characteristics (full-body and body-parts) to perform robust collaborative fusion based on data-driven Gaussian Mixture Probability Hypothesis Density (GM-PHD) filters. To leverage the maximum strengths from multiple detectors, we propose a robust fusion center at the track level, which manages to perform Generalized Intersection Covariance (GCI) fusions for survival and birth tracks independently, and also eliminates false tracks caused by a cluttered environment. Moreover, an identity reassignment mechanism is also developed to address the identity mismatching problem in the target birth process, so as to enhance the fusion performance and track consistency. Experimental results on two challenging benchmark video sequences confirm the effectiveness of the proposed approach. Zeyu Fu, Syed M. Naqvi, Jonathon A. Chambers |
FUSION | 1 |
| 2018 | 3D-Hog Embedding Frameworks for Single and Multi-Viewpoints Action Recognition Based on Human SilhouettesabstractGiven the high demand for automated systems for human action recognition, great efforts have been undertaken in recent decades to progress the field. In this paper, we present frameworks for single and multi-viewpoints action recognition based on Space-Time Volume (STV) of human silhouettes and 3D-Histogram of Oriented Gradient (3D-HOG) embedding. We exploit fast-computational approaches involving Principal Component Analysis (PCA) over the local feature spaces for compactly describing actions as combinations of local gestures and L2-Regularized Logistic Regression (L2-RLR) for learning the action model from local features. Outperforming results on Weizmann and i3DPost datasets confirm efficacy of the proposed approaches as compared to the baseline method and other works, in terms of accuracy and robustness to appearance changes. Federico Angelini, Zeyu Fu, Sergio A. Velastin, Jonathon A. Chambers, Syed M. Naqvi |
ICASSP | 2 |
| 2018 | GM-PHD Filter Based Online Multiple Human Tracking Using Deep Discriminative Correlation MatchingabstractIn this paper, we propose deep discriminative correlation matching within the Gaussian Mixture Probability Hypothesis Density (GM-PHD) filter for online multiple human tracking. In this matching scheme, we mainly exploit the Convolutional Neural Network (CNN) based Discriminative Correlation Filter (DCF) as a target-specific classifier to discriminate the desired target from background and remaining targets. DCFs are learned through the extracted features obtained from the outputs of the last convolutional layers which are capable to encode the target appearances with better discriminativity and robustness to appearance changes. Moreover, we present a hybrid likelihood function that fuses the spatio-temporal relation and correlation matching score to collaboratively enhance the PHD association step. Experimental results on the MOT17 Challenge benchmark [1] confirm the improved performance of our proposed method as compared with other state-of-the-art techniques. Zeyu Fu, Federico Angelini, Syed M. Naqvi, Jonathon A. Chambers |
ICASSP | 1 |
| 2017 | Particle PHD filter based multi-target tracking using discriminative group-structured dictionary learningabstractStructured sparse representation has been recently found to achieve better efficiency and robustness in exploiting the target appearance model in tracking systems with both holistic and local information. Therefore, to better simultaneously discriminate multi-targets from their background, we propose a novel video-based multi-target tracking system that combines the particle probability hypothesis density (PHD) filter with discriminative group-structured dictionary learning. The discriminative dictionary with group structure learned by the hierarchical K-means clustering algorithm implicitly associates the dictionary atoms with the group labels, simultaneously enforcing the target candidates from the same group (class) to share the same structured sparsity pattern. Furthermore, we propose a new joint likelihood calculation by relating the discriminative sparse codes with the maximum voting technique to enhance the particle PHD updating step. Experimental results on two publicly available benchmark video sequences confirm the improved performance of our proposed method over other state-of-the-art techniques in video-based multi-target tracking. Zeyu Fu, Pengming Feng, Syed M. Naqvi, Jonathon A. Chambers |
ICASSP | 1 |