VLDB 2026 Research / reviewers in the wild / expert
Bo Li 0114
dblp:50/3402-114
· DBLP profile ↗
8ranked-venue papers
2as first author
5since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | SiamSampler: Video-Guided Sampling for Siamese Visual TrackingabstractWe present SiamSampler, the first to our knowledge investigating video sampling in visual object tracking. We observe that the random sampling applied in Siamese-based trackers cannot focus on important data or ensure data diversity, hindering the effective training of networks. This paper proposes the Video-Guided Sampling Strategy to solve the problems in random sampling from both inter and intra-video levels. At the inter-video level, we propose Modified Gaussian Sampling Strategy (MGSS) to automatically assign higher sampling probabilities to longer and more difficult videos and reduce the sampling probabilities of shorter and easier videos. At the intra-video level, the Farthest Image Pair Sampling Strategy (FPSS) is proposed to increase the diversity of training data. Extensive experiments on general benchmarks demonstrate the effectiveness of our method. Compared with the baseline model, our method improves tracking performance on five datasets, without affecting the testing speed. Peixia Li, Lei Bai 0001, Lei Qiao 0004, Bo Li 0114, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Cross Domain Object Detection by Target-Perceived Dual Branch DistillationabstractCross domain object detection is a realistic and challenging task in the wild. It suffers from performance degradation due to large shift of data distributions and lack of instance-level annotations in the target domain. Existing approaches mainly focus on either of these two difficulties, even though they are closely coupled in cross domain object detection. To solve this problem, we propose a novel Target-perceived Dual-branch Distillation (TDD) framework. By integrating detection branches of both source and target domains in a unified teacher-student learning scheme, it can reduce domain shift and generate reliable supervision effectively. In particular, we first introduce a distinct Target Proposal Perceiver between two domains. It can adaptively enhance source detector to perceive objects in a target image, by leveraging target proposal contexts from iterative cross-attention. Afterwards, we design a concise Dual Branch Self Distillation strategy for model training, which can progressively integrate complementary object knowledge from different domains via self-distillation in two branches. Finally, we conduct extensive experiments on a number of widely-used scenarios in cross domain object detection. The results show that our TDD significantly outperforms the state-of-the-art methods on all the benchmarks. The codes and models will be released afterwards. Mengzhe He, Yali Wang 0001, Yiru Wang 0003, Hanqing Li, Bo Li 0114, Weihao Gan, Wei Wu 0021, Yu Qiao 0001 |
CVPR | 6 |
| 2022 | Unsupervised Learning of Accurate Siamese TrackingabstractUnsupervised learning has been popular in various computer vision tasks, including visual object tracking. However, prior unsupervised tracking approaches rely heavily on spatial supervision from templatesearch pairs and are still unable to track objects with strong variation over a long time span. As unlimited self-supervision signals can be obtained by tracking a video along a cycle in time, we investigate evolving a Siamese tracker by tracking videos forward-backward. We present a novel unsupervised tracking framework, in which we can learn temporal correspondence both on the classification branch and regression branch. Specifically, to propagate reliable template feature in the forward propagation process so that the tracker can be trained in the cycle, we first propose a consistency propagation transformation. We then identify an ill-posed penalty problem in conventional cycle training in backward propagation process. Thus, a differentiable region mask is proposed to select features as well as to implicitly penalize tracking errors on intermediate frames. Moreover, since noisy labels may degrade training, we propose a mask-guided loss reweighting strategy to assign dynamic weights based on the quality of pseudo labels. In extensive experiments, our tracker outperforms preceding unsupervised methods by a substantial margin, performing on par with supervised methods on large-scale datasets such as TrackingNet and LaSOT. Code is available at https://github.com/FlorinShum/ULAST. Qiuhong Shen, Lei Qiao 0004, Jinyang Guo 0002, Peixia Li, Xin Li 0034, Bo Li 0114, Weihao Gan, Wei Wu 0021, Wanli Ouyang |
CVPR | 6 |
| 2022 | Target-Relevant Knowledge Preservation for Multi-Source Domain Adaptive Object DetectionabstractDomain adaptive object detection (DAOD) is a promising way to alleviate performance drop of detectors in new scenes. Albeit great effort made in single source domain adaptation, a more generalized task with multiple source domains remains not being well explored, due to knowledge degradation during their combination. To address this issue, we propose a novel approach, namely target-relevant knowledge preservation (TRKP), to unsupervised multi-source DAOD. Specifically, TRKP adopts the teacher-student framework, where the multi-head teacher network is built to extract knowledge from labeled source domains and guide the student network to learn detectors in unlabeled target domain. The teacher network is further equipped with an adversarial multi-source disentanglement (AMSD) module to preserve source domain-specific knowledge and simultaneously perform cross-domain alignment. Besides, a holistic target-relevant mining (HTRM) scheme is developed to re-weight the source images according to the source-target relevance. By this means, the teacher network is enforced to capture target-relevant knowledge, thus benefiting decreasing domain shift when mentoring object detection in the target domain. Extensive experiments are conducted on various widely used benchmarks with new state-of-the-art scores reported, highlighting the effectiveness. Jiaxin Chen 0002, Mengzhe He, Yiru Wang 0003, Bo Li 0114, Bingqi Ma, Weihao Gan, Wei Wu 0021, Yali Wang 0001, Di Huang 0001 |
CVPR | 5 |
| 2022 | Backbone is All Your Need: A Simplified Architecture for Visual Object Tracking
Peixia Li, Lei Bai 0001, Lei Qiao 0004, Qiuhong Shen, Bo Li 0114, Weihao Gan, Wei Wu 0021, Wanli Ouyang |
ECCV (22) | 6 |
| 2019 | SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep NetworksabstractSiamese network based trackers formulate tracking as convolutional feature cross-correlation between target template and searching region. However, Siamese trackers still have accuracy gap compared with state-of-the-art algorithms and they cannot take advantage of feature from deep networks, such as ResNet-50 or deeper. In this work we prove the core reason comes from the lack of strict translation invariance. By comprehensive theoretical analysis and experimental validations, we break this restriction through a simple yet effective spatial aware sampling strategy and successfully train a ResNet-driven Siamese tracker with significant performance gain. Moreover, we propose a new model architecture to perform depth-wise and layer-wise aggregations, which not only further improves the accuracy but also reduces the model size. We conduct extensive ablation studies to demonstrate the effectiveness of the proposed tracker, which obtains currently the best results on four large tracking benchmarks, including OTB2015, VOT2018, UAV123, and LaSOT. Our model will be released to facilitate further studies based on this problem. Bo Li 0114, Wei Wu 0021, Qiang Wang 0051, Fangyi Zhang, Junliang Xing |
CVPR | 1 |
| 2018 | High Performance Visual Tracking With Siamese Region Proposal NetworkabstractVisual object tracking has been a fundamental topic in recent years and many deep learning based trackers have achieved state-of-the-art performance on multiple benchmarks. However, most of these trackers can hardly get top performance with real-time speed. In this paper, we propose the Siamese region proposal network (Siamese-RPN) which is end-to-end trained off-line with large-scale image pairs. Specifically, it consists of Siamese subnetwork for feature extraction and region proposal subnetwork including the classification branch and regression branch. In the inference phase, the proposed framework is formulated as a local one-shot detection task. We can pre-compute the template branch of the Siamese subnetwork and formulate the correlation layers as trivial convolution layers to perform online tracking. Benefit from the proposal refinement, traditional multi-scale test and online fine-tuning can be discarded. The Siamese-RPN runs at 160 FPS while achieving leading performance in VOT2015, VOT2016 and VOT2017 real-time challenges. Bo Li 0114, Wei Wu 0021 |
CVPR | 1 |
| 2018 | Distractor-Aware Siamese Networks for Visual Object Tracking
Qiang Wang 0051, Bo Li 0114, Wei Wu 0021, Weiming Hu 0004 |
ECCV (9) | 3 |