VLDB 2026 Research / reviewers in the wild / expert
Chao Fan 0001
dblp:56/5961-1 · also Fan Chao 0001
· DBLP profile ↗
24ranked-venue papers
9as first author
22since 2021 · last 2026
0000-0002-3605-2705ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain generalization for lesion classification via frequency swapping and soft-mask disentanglement
Xibin Jia, Shaowu Xu, Chao Fan 0001, Zhenghan Yang |
Expert Syst. Appl. | 5 |
| 2026 | VLAlignSeg: Vision-language alignment for few-shot medical image segmentation
Haipeng Qiao, Xibin Jia, Chao Fan 0001 |
Expert Syst. Appl. | 4 |
| 2026 | SAFE: A Semantic-Appearance Full-body Editor for pedestrian video anonymization
Jingzhe Ma, Chao Fan 0001, Dingqiang Ye, Jinfeng Yang, Dongyang Jin, Fuad Mire Hassan, Shiqi Yu 0001 |
Pattern Recognit. | 2 |
| 2025 | Exploring More from Multiple Gait Modalities for Human IdentificationabstractThe gait, as a kind of soft biometric characteristic, can reflect the distinct walking patterns of individuals at a distance, exhibiting a promising technique for unrestrained human identification. With largely excluding gait-unrelated cues hidden in RGB videos, the silhouette and skeleton, though visually compact, have acted as two of the most prevailing gait modalities for a long time. Recently, several attempts have been made to introduce more informative data forms like human parsing and optical flow images to capture gait characteristics, along with multi-branch architectures. However, due to the inconsistency within model designs and experiment settings, we argue that a comprehensive and fair comparative study among these popular gait modalities, involving the representational capacity and fusion strategy exploration, is still lacking. From the perspectives of fine vs. coarse-grained shape and whole vs. pixel-wise motion modeling, this work presents an in-depth investigation of three popular gait representations, i.e., silhouette, human parsing, and optical flow, with various fusion evaluations, and experimentally exposes their similarities and differences. Based on the obtained insights, we further develop a C²Fusion strategy, consequently building our new framework MultiGait++. C²Fusion preserves commonalities while highlighting differences to enrich the learning of gait features. To verify our findings and conclusions, extensive experiments on Gait3D, GREW, CCPG, and SUSTech1K are conducted. Dongyang Jin, Chao Fan 0001, Shiqi Yu 0001 |
AAAI | 2 |
| 2025 | On Denoising Walking Videos for Gait RecognitionabstractTo capture individual gait patterns, excluding identity-irrelevant cues in walking videos, such as clothing texture and color, remains a persistent challenge for vision-based gait recognition. Traditional silhouette- and pose-based methods, though theoretically effective at removing such distractions, often fall short of high accuracy due to their sparse and less informative inputs. Emerging end-to-end methods address this by directly denoising RGB videos using human priors. Building on this trend, we propose DenoisingGait, a novel gait denoising method. Inspired by the philosophy that "what I cannot create, I do not understand", we turn to generative diffusion models, uncovering how they partially filter out irrelevant factors for gait understanding. Additionally, we introduce a geometry-driven Feature Matching module, which, combined with background removal via human silhouettes, condenses the multi-channel diffusion features at each foreground pixel into a two-channel direction vector. Specifically, the proposed within- and cross-frame matching respectively capture the local vectorized structures of gait appearance and motion, producing a novel flow-like gait representation termed Gait Feature Field, which further reduces residual noise in diffusion features. Experiments on the CCPG, CASIA-B*, and SUSTech1K datasets demonstrate that DenoisingGait achieves a new SoTA performance in most cases for both within- and cross-domain evaluations. Code is available at https://github.com/ShiqiYu/OpenGait. Dongyang Jin, Chao Fan 0001, Jingzhe Ma, Jingkai Zhou, Shiqi Yu 0001 |
CVPR | 2 |
| 2025 | SPENet: Self-guided Prototype Enhancement Network for Few-Shot Medical Image Segmentation
Chao Fan 0001, Xibin Jia, Anqi Xiao, Hongyuan Yu, Zhenghan Yang, Yan Huang 0008, Liang Wang 0001 |
MICCAI (5) | 1 |
| 2025 | Pose as Clinical Prior: Learning Dual Representations for Scoliosis Screening
Zirui Zhou, Zizhao Peng, Dongyang Jin, Chao Fan 0001, Fengwei An, Shiqi Yu 0001 |
MICCAI (13) | 4 |
| 2025 | Cross-Modal Dual-Causal Learning for Long-Term Action RecognitionabstractLong-term action recognition (LTAR) is challenging due to extended temporal spans with complex atomic action correlations and visual confounders. Although vision-language models (VLMs) have shown promise, they often rely on statistical correlations instead of causal mechanisms. Moreover, existing causality-based methods address modal-specific biases but lack cross-modal causal modeling, limiting their utility in VLM-based LTAR. This paper proposes Cross-Modal Dual-Causal Learning (CMDCL), which introduces a structural causal model to uncover causal relationships between videos and label texts. CMDCL addresses cross-modal biases in text embeddings via textual causal intervention and removes confounders inherent in the visual modality through visual causal intervention guided by the debiased text. These dual-causal interventions enable robust action representations to address LTAR challenges. Experimental results on three benchmarks including Charades, Breakfast and COIN, demonstrate the effectiveness of the proposed model. Our code is available at https://github.com/xushaowu/CMDCL. Shaowu Xu, Xibin Jia, Junyu Gao 0002, Qianmei Sun, Jing Chang 0006, Chao Fan 0001 |
ACM Multimedia | 6 |
| 2025 | BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision ModelsabstractLarge vision models (LVM) based gait recognition has achieved impressive performance.
However, existing LVM-based approaches may overemphasize gait priors while neglecting the intrinsic value of LVM itself, particularly the rich, distinct representations across its multi-layers.
To adequately unlock LVM's potential, this work investigates the impact of layer-wise representations on downstream recognition tasks.
Our analysis reveals that LVM's intermediate layers offer complementary properties across tasks, integrating them yields an impressive improvement even without rich well-designed gait priors.
Building on this insight, we propose a simple and universal baseline for LVM-based gait recognition, termed BiggerGait.
Comprehensive evaluations on CCPG, CAISA-B*, SUSTech1K, and CCGR_MINI validate the superiority of BiggerGait across both within- and cross-domain tasks, establishing it as a simple yet practical baseline for gait representation learning.
All the models and code are available at https://github.com/ShiqiYu/OpenGait/. Dingqiang Ye, Chao Fan 0001, Zhanbo Huang, Chengwen Luo 0001, Jianqiang Li 0001, Shiqi Yu 0001, Xiaoming Liu 0002 |
NeurIPS | 2 |
| 2025 | OpenGait: A Comprehensive Benchmark Study for Gait Recognition Toward Better PracticalityabstractGait recognition, a rapidly advancing vision technology for person identification from a distance, has made significant strides in indoor settings. However, evidence suggests that existing methods often yield unsatisfactory results when applied to newly released real-world gait datasets. Furthermore, conclusions drawn from indoor gait datasets may not easily generalize to outdoor ones. Therefore, the primary goal of this paper is to present a comprehensive benchmark study aimed at improving practicality rather than solely focusing on enhancing performance. To this end, we developed OpenGait, a flexible and efficient gait recognition platform. Using OpenGait, we conducted in-depth ablation experiments to revisit recent developments in gait recognition. Surprisingly, we detected some imperfect parts of some prior methods and thereby uncovered several critical yet previously neglected insights. These findings led us to develop three structurally simple yet empirically powerful and practically robust baseline models: DeepGaitV2, SkeletonGait, and SkeletonGait++, which represent the appearance-based, model-based, and multi-modal methodologies for gait pattern description, respectively. In addition to achieving state-of-the-art performance, our careful exploration provides new perspectives on the modeling experience of deep gait models and the representational capacity of typical gait modalities. In the end, we discuss the key trends and challenges in current gait recognition, aiming to inspire further advancements towards better practicality. Chao Fan 0001, Saihui Hou, Chuanfu Shen, Jingzhe Ma, Dongyang Jin, Yongzhen Huang, Shiqi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | SliceMamba With Neural Architecture Search for Medical Image SegmentationabstractDespite the progress made in Mamba-based medical image segmentation models, existing methods utilizing unidirectional or multi-directional feature scanning mechanisms struggle to effectively capture dependencies between neighboring positions, limiting the discriminant representation learning of local features. These local features are crucial for medical image segmentation as they provide critical structural information about lesions and organs. To address this limitation, we propose SliceMamba, a simple yet effective locally sensitive Mamba-based medical image segmentation model. SliceMamba features an efficient Bidirectional Slicing and Scanning (BSS) module, which performs bidirectional feature slicing and employs varied scanning mechanisms for sliced features with distinct shapes. This design keeps spatially adjacent features close in the scan sequence, preserving the local structure of the image and enhancing segmentation performance. Additionally, to fit the varying sizes and shapes of lesions and organs, we introduce an Adaptive Slicing Search method that automatically identifies the optimal feature slicing method based on the characteristics of the target data. Extensive experiments on two skin lesion datasets (ISIC2017 and ISIC2018), two polyp segmentation datasets (Kvasir and ClinicDB), one ultra-wide field retinal hemorrhage segmentation dataset (UWF-RHS), and one multi-organ segmentation dataset (Synapse) demonstrate the effectiveness of our method. Chao Fan 0001, Hongyuan Yu, Yan Huang 0008, Liang Wang 0001, Zhenghan Yang, Xibin Jia |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | SkeletonGait: Gait Recognition Using Skeleton MapsabstractThe choice of the representations is essential for deep gait recognition methods. The binary silhouettes and skeletal coordinates are two dominant representations in recent literature, achieving remarkable advances in many scenarios. However, inherent challenges remain, in which silhouettes are not always guaranteed in unconstrained scenes, and structural cues have not been fully utilized from skeletons. In this paper, we introduce a novel skeletal gait representation named skeleton map, together with SkeletonGait, a skeleton-based method to exploit structural information from human skeleton maps. Specifically, the skeleton map represents the coordinates of human joints as a heatmap with Gaussian approximation, exhibiting a silhouette-like image devoid of exact body structure. Beyond achieving state-of-the-art performances over five popular gait datasets, more importantly, SkeletonGait uncovers novel insights about how important structural features are in describing gait and when they play a role. Furthermore, we propose a multi-branch architecture, named SkeletonGait++, to make use of complementary features from both skeletons and silhouettes. Experiments indicate that SkeletonGait++ outperforms existing state-of-the-art methods by a significant margin in various scenarios. For instance, it achieves an impressive rank-1 accuracy of over 85% on the challenging GREW dataset. The source code is available at https://github.com/ShiqiYu/OpenGait. Chao Fan 0001, Jingzhe Ma, Dongyang Jin, Chuanfu Shen, Shiqi Yu 0001 |
AAAI | 1 |
| 2024 | Cross-Covariate Gait Recognition: A BenchmarkabstractGait datasets are essential for gait research. However, this paper observes that present benchmarks, whether conventional constrained or emerging real-world datasets, fall short regarding covariate diversity. To bridge this gap, we undertake an arduous 20-month effort to collect a cross-covariate gait recognition (CCGR) dataset. The CCGR dataset has 970 subjects and about 1.6 million sequences; almost every subject has 33 views and 53 different covariates. Compared to existing datasets, CCGR has both population and individual-level diversity. In addition, the views and covariates are well labeled, enabling the analysis of the effects of different factors. CCGR provides multiple types of gait data, including RGB, parsing, silhouette, and pose, offering researchers a comprehensive resource for exploration. In order to delve deeper into addressing cross-covariate gait recognition, we propose parsing-based gait recognition (ParsingGait) by utilizing the newly proposed parsing data. We have conducted extensive experiments. Our main results show: 1) Cross-covariate emerges as a pivotal challenge for practical applications of gait recognition. 2) ParsingGait demonstrates remarkable potential for further advancement. 3) Alarmingly, existing SOTA methods achieve less than 43% accuracy on the CCGR, highlighting the urgency of exploring cross-covariate gait recognition. Link: https://github.com/ShinanZou/CCGR. Shinan Zou, Chao Fan 0001, Jianbo Xiong, Chuanfu Shen, Shiqi Yu 0001, Jin Tang 0001 |
AAAI | 2 |
| 2024 | BigGait: Learning Gait Representation You Want by Large Vision ModelsabstractGait recognition stands as one of the most pivotal remote identification technologies and progressively expands across research and industry communities. However, existing gait recognition methods heavily rely on task-specific upstream driven by supervised learning to provide explicit gait representations like silhouette sequences, which in-evitably introduce expensive annotation costs and poten-tial error accumulation. Escaping from this trend, this work explores effective gait representations based on the all-purpose knowledge produced by task-agnostic Large Vision Models (LVMs) and proposes a simple yet efficient gait framework, termed B igGait. Specifically, the Gait Repre-sentation Extractor (GRE) within BigGait draws upon design principles from established gait representations, effectively transforming all-purpose knowledge into implicit gait representations without requiring third-party supervision signals. Experiments on CCPG, CAISA-B* and SUSTechlK indicate that BigGait significantly outperforms the previous methods in both within-domain and cross-domain tasks in most cases, and provides a more practical paradigm for learning the next-generation gait representation. Fi-nally, we delve into prospective challenges and promising directions in LVMs-based gait recognition, aiming to in-spire future work in this emerging topic. The source code is available at https://github.com/ShiqiYu/OpenGait. Dingqiang Ye, Chao Fan 0001, Jingzhe Ma, Xiaoming Liu 0002, Shiqi Yu 0001 |
CVPR | 2 |
| 2024 | Gait Patterns as Biomarkers: A Video-Based Approach for Classifying Scoliosis
Zirui Zhou, Zizhao Peng, Chao Fan 0001, Fengwei An, Shiqi Yu 0001 |
MICCAI (5) | 4 |
| 2023 | OpenGait: Revisiting Gait Recognition Toward Better PracticalityabstractGait recognition is one of the most critical long-distance identification technologies and increasingly gains popularity in both research and industry communities. Despite the significant progress made in indoor datasets, much evidence shows that gait recognition techniques perform poorly in the wild. More importantly, we also find that some conclusions drawn from indoor datasets cannot be generalized to real applications. Therefore, the primary goal of this paper is to present a comprehensive benchmark study for better practicality rather than only a particular model for better performance. To this end, we first develop a flexible and efficient gait recognition codebase named OpenGait. Based on OpenGait, we deeply revisit the recent development of gait recognition by re-conducting the ablative experiments. Encouragingly, we detect some unperfect parts of certain prior woks, as well as new insights. Inspired by these discoveries, we develop a structurally simple, empirically powerful, and practically robust baseline model, Gait-Base. Experimentally, we comprehensively compare Gait-Base with many current gait recognition methods on multiple public datasets, and the results reflect that GaitBase achieves significantly strong performance in most cases regardless of indoor or outdoor situations. Code is available at https://github.com/ShiqiYu/OpenGait. Chao Fan 0001, Chuanfu Shen, Saihui Hou, Yongzhen Huang, Shiqi Yu 0001 |
CVPR | 1 |
| 2023 | LidarGait: Benchmarking 3D Gait Recognition with Point CloudsabstractVideo-based gait recognition has achieved impressive results in constrained scenarios. However, visual cameras neglect human 3D structure information, which limits the feasibility of gait recognition in the 3D wild world. Instead of extracting gait features from images, this work explores precise 3D gait features from point clouds and proposes a simple yet efficient 3D gait recognition framework, termed LidarGait. Our proposed approach projects sparse point clouds into depth maps to learn the representations with 3D geometry information, which outperforms existing point-wise and camera-based methods by a significant margin. Due to the lack of point cloud datasets, we build the first large-scale LiDAR-based gait recognition dataset, SUSTech1K, collected by a LiDAR sensor and an RGB camera. The dataset contains 25,239 sequences from 1,050 subjects and covers many variations, including visibility, views, occlusions, clothing, carrying, and scenes. Extensive experiments show that (1) 3D structure information serves as a significant feature for gait recognition. (2) LidarGait outperforms existing point-based and silhouette-based methods by a significant margin, while it also offers stable cross-view results. (3) The LiDAR sensor is superior to the RGB camera for gait recognition in the outdoor environment. The source code and dataset have been made available at https://lidargait.github.io. Chuanfu Shen, Chao Fan 0001, Wei Wu 0041, George Q. Huang, Shiqi Yu 0001 |
CVPR | 2 |
| 2023 | PointGait: Boosting End-to-End 3D Gait Recognition with Point Clouds via Spatiotemporal ModelingabstractLiDAR is a new type of sensor used for gait recognition. Previous LiDAR-based state-of-the-art methods mostly exploit gait features from the depth maps generated by projecting point clouds in a 3D-to-2D manner, rather than directly using the raw 3D point data. However, these projection-based methods require an additional preprocessing step, which obstructs the universality of the method among different types of LiDARs. On the other hand, while existing point-based methods have achieved promising results in 3D object recognition, they have underperformed in 3D gait recognition, indicating the presence of a domain gap between coarse-grained 3D object classification and fine-grained 3D pedestrians recognition. By analyzing the success achieved by camera-based methods, we perceive that point-based gait recognition fails mainly because of neglecting to capture local representation. To address this issue, we propose an end-to-end 3D gait recognition framework named PointGait, which can directly capture informative gait features from point cloud data. Specifically, PointGait is a multi-stream model consisting of a Global and Local Gait Feature Extractor to extract holistic and fine-grained spatial features. Besides, a Personalized Motion Extractor is introduced to capture inter-frame motion features. Our experimental results on a LiDAR gait dataset, SUSTech1K, outperform all popular point-based methods, demonstrating the effectiveness and potential of our approach. In conclusion, the proposed PointGait promotes the development of point-based gait recognition by highlighting the importance of incorporating fine-grained spatiotemporal information. Chuanfu Shen, Chao Fan 0001, George Q. Huang, Shiqi Yu 0001 |
IJCB | 3 |
| 2023 | A Multi-Stage Adaptive Feature Fusion Neural Network for Multimodal Gait RecognitionabstractGait recognition is a biometric technology that has received extensive attention. Most existing gait recognition algorithms are unimodal, and a few multimodal gait recognition algorithms perform multimodal fusion only once. None of these algorithms may fully exploit the complementary advantages of the multiple modalities. In this paper, by considering the temporal and spatial characteristics of gait data, we propose a multi-stage feature fusion strategy (MSFFS), which performs multimodal fusions at different stages in the feature extraction process. Also, we propose an adaptive feature fusion module (AFFM) that considers the semantic association between silhouettes and skeletons. The fusion process fuses different silhouette areas with their more related skeleton joints. Since visual appearance changes and time passage co-occur in a gait period, we propose a multiscale spatial-temporal feature extractor (MSSTFE) to learn the spatial-temporal linkage features thoroughly. Specifically, MSSTFE extracts and aggregates spatial-temporal linkages information at different spatial scales. Combining the strategy and modules mentioned above, we propose a multi-stage adaptive feature fusion (MSAFF) neural network, which shows state-of-the-art performance in many experiments on three datasets. Besides, MSAFF is equipped with feature dimensional pooling (FD Pooling), which can significantly reduce the dimension of the gait representations without hindering the accuracy. Shinan Zou, Jianbo Xiong, Chao Fan 0001, Shiqi Yu 0001, Jin Tang 0001 |
IJCB | 3 |
| 2023 | Learning Gait Representation From Massive Unlabelled Walking Videos: A BenchmarkabstractGait depicts individuals' unique and distinguishing walking patterns and has become one of the most promising biometric features for human identification. As a fine-grained recognition task, gait recognition is easily affected by many factors and usually requires a large amount of completely annotated data that is costly and insatiable. This paper proposes a large-scale self-supervised benchmark for gait recognition with contrastive learning, aiming to learn the general gait representation from massive unlabelled walking videos for practical applications via offering informative walking priors and diverse real-world variations. Specifically, we collect a large-scale unlabelled gait dataset GaitLU-1M consisting of 1.02M walking sequences and propose a conceptually simple yet empirically powerful baseline model GaitSSB. Experimentally, we evaluate the pre-trained model on four widely-used gait benchmarks, CASIA-B, OU-MVLP, GREW and Gait3D with or without transfer learning. The unsupervised results are comparable to or even better than the early model-based and GEI-based methods. After transfer learning, GaitSSB outperforms existing methods by a large margin in most cases, and also showcases the superior generalization capacity. Further experiments indicate that the pre-training can save about 50% and 80% annotation costs of GREW and Gait3D. Theoretically, we discuss the critical issues for gait-specific contrastive framework and present some insights for further study. As far as we know, GaitLU-1M is the first large-scale unlabelled gait dataset, and GaitSSB is the first method that achieves remarkable unsupervised results on the aforementioned benchmarks. Chao Fan 0001, Saihui Hou, Jilong Wang 0010, Yongzhen Huang, Shiqi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | SiamON: Siamese Occlusion-Aware Network for Visual TrackingabstractOcclusion has been proven to be one of the most challenging factors faced by most visual trackers. There are mainly two difficulties, the first one is that the number of occlusion samples are very limited even though collecting a large-scale training data set, and another one is how to correctly learn the features of the target when comes to occlusion situations. In this paper, we tried to solve these two problems together in our proposed model. To this end, we propose a novel Siamese Occlusion-aware Network (SiamON) for high-performance visual tracking. In particular, we predefine some soft-masks to solve the problem of fewer occlusion samples, which perceive patterns of occlusion contents at different locations and take these masks as the conditions to guide occlusion-aware feature learning. Meanwhile, we propose a target-aware attention mechanism allows the model to pay more attention to the target and further weaken the impact of occlusion. Extensive experiments on several popular benchmarks show that our tracking method exceeds many state-of-the-art trackers especially in the presence of occlusion and meets the requirements of real-time. Chao Fan 0001, Hongyuan Yu, Yan Huang 0008, Caifeng Shan, Liang Wang 0001, Chenglong Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | GaitEdge: Beyond Plain End-to-End Gait Recognition for Better Practicality
Chao Fan 0001, Saihui Hou, Chuanfu Shen, Yongzhen Huang, Shiqi Yu 0001 |
ECCV (5) | 2 |
| 2020 | GaitPart: Temporal Part-Based Model for Gait RecognitionabstractGait recognition, applied to identify individual walking patterns in a long-distance, is one of the most promising video-based biometric technologies. At present, most gait recognition methods take the whole human body as a unit to establish the spatio-temporal representations. However, we have observed that different parts of human body possess evidently various visual appearances and movement patterns during walking. In the latest literature, employing partial features for human body description has been verified being beneficial to individual recognition. Taken above insights together, we assume that each part of human body needs its own spatio-temporal expression. Then, we propose a novel part-based model GaitPart and get two aspects effect of boosting the performance: On the one hand, Focal Convolution Layer, a new applying of convolution, is presented to enhance the fine-grained learning of the part-level spatial features. On the other hand, the Micro-motion Capture Module (MCM) is proposed and there are several parallel MCMs in the GaitPart corresponding to the pre-defined parts of the human body, respectively. It is worth mentioning that the MCM is a novel way of temporal modeling for gait task, which focuses on the short-range temporal features rather than the redundant long-range features for cycle gait. Experiments on two of the most popular public datasets, CASIA-B and OU-MVLP, richly exemplified that our method meets a new state-of-the-art on multiple standard benchmarks. The source code will be available on https://github.com/ChaoFan96/GaitPart. Chao Fan 0001, Yunjie Peng, Chunshui Cao, Xu Liu 0008, Saihui Hou, Jiannan Chi, Yongzhen Huang, Qing Li 0015, Zhiqiang He 0002 |
CVPR | 1 |
| 2019 | Visual Tracking Via Siamese Network With Global SimilarityabstractVisual tracking is a very important and challenging problem in the field of computer vision. In recent years, Siamese networks have been widely used for visual tracking due to their fast tracking speed, but many trackers based on Siamese network train their networks by utilizing either pairwise loss or triplet loss, which easily leads to over-fitting. In addition, it is difficult to distinguish some hard samples in the training samples. In this paper, we propose a novel global similarity loss to train the network. Specifically, we utilize two Gaussian distributions to simulate and optimize the distribution of positive and negative samples in the train set and add constraint on the hard samples. In experiments, without any other modification, we apply the proposed method to the Siamese network. And the results on several popular tracking benchmarks show our method achieves superior tracking performance than the baseline. Chao Fan 0001, Chenglong Li 0002, Jin Tang 0001 |
ICIP | 1 |