VLDB 2026 Research / reviewers in the wild / expert
Peixia Li
dblp:213/5896
· DBLP profile ↗
12ranked-venue papers
5as first author
7since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Semi-Supervised Semantic Segmentation under Label Noise via Diverse Learning GroupsabstractSemi-supervised semantic segmentation methods use a small amount of clean pixel-level annotations to guide the interpretation of a larger quantity of unlabelled image data. The challenges of providing pixel-accurate annotations at scale mean that the labels are typically noisy, and this contaminates the final results. In this work, we propose an approach that is robust to label noise in the annotated data. The method uses two diverse learning groups with different network architectures to effectively handle both label noise and unlabelled images. Each learning group consists of a teacher network, a student network and a novel filter module. The filter module of each learning group utilizes pixel-level features from the teacher network to detect incorrectly labelled pixels. To reduce confirmation bias, we employ the labels cleaned by the filter module from one learning group to train the other learning group. Experimental results on two different benchmarks and settings demonstrate the superiority of our method over state-of-the-art approaches. Peixia Li, Pulak Purkait, Thalaiyasingam Ajanthan, Majid Abdolshah, Ravi Garg, Hisham Husain, Stephen Gould, Wanli Ouyang, Anton van den Hengel |
ICCV | 1 |
| 2023 | SiamSampler: Video-Guided Sampling for Siamese Visual TrackingabstractWe present SiamSampler, the first to our knowledge investigating video sampling in visual object tracking. We observe that the random sampling applied in Siamese-based trackers cannot focus on important data or ensure data diversity, hindering the effective training of networks. This paper proposes the Video-Guided Sampling Strategy to solve the problems in random sampling from both inter and intra-video levels. At the inter-video level, we propose Modified Gaussian Sampling Strategy (MGSS) to automatically assign higher sampling probabilities to longer and more difficult videos and reduce the sampling probabilities of shorter and easier videos. At the intra-video level, the Farthest Image Pair Sampling Strategy (FPSS) is proposed to increase the diversity of training data. Extensive experiments on general benchmarks demonstrate the effectiveness of our method. Compared with the baseline model, our method improves tracking performance on five datasets, without affecting the testing speed. Peixia Li, Lei Bai 0001, Lei Qiao 0004, Bo Li 0114, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Unsupervised Learning of Accurate Siamese TrackingabstractUnsupervised learning has been popular in various computer vision tasks, including visual object tracking. However, prior unsupervised tracking approaches rely heavily on spatial supervision from templatesearch pairs and are still unable to track objects with strong variation over a long time span. As unlimited self-supervision signals can be obtained by tracking a video along a cycle in time, we investigate evolving a Siamese tracker by tracking videos forward-backward. We present a novel unsupervised tracking framework, in which we can learn temporal correspondence both on the classification branch and regression branch. Specifically, to propagate reliable template feature in the forward propagation process so that the tracker can be trained in the cycle, we first propose a consistency propagation transformation. We then identify an ill-posed penalty problem in conventional cycle training in backward propagation process. Thus, a differentiable region mask is proposed to select features as well as to implicitly penalize tracking errors on intermediate frames. Moreover, since noisy labels may degrade training, we propose a mask-guided loss reweighting strategy to assign dynamic weights based on the quality of pseudo labels. In extensive experiments, our tracker outperforms preceding unsupervised methods by a substantial margin, performing on par with supervised methods on large-scale datasets such as TrackingNet and LaSOT. Code is available at https://github.com/FlorinShum/ULAST. Qiuhong Shen, Lei Qiao 0004, Jinyang Guo 0002, Peixia Li, Xin Li 0034, Bo Li 0114, Weihao Gan, Wei Wu 0021, Wanli Ouyang |
CVPR | 4 |
| 2022 | Backbone is All Your Need: A Simplified Architecture for Visual Object Tracking
Peixia Li, Lei Bai 0001, Lei Qiao 0004, Qiuhong Shen, Bo Li 0114, Weihao Gan, Wei Wu 0021, Wanli Ouyang |
ECCV (22) | 2 |
| 2021 | BN-NAS: Neural Architecture Search with Batch NormalizationabstractWe present BN-NAS, neural architecture search with Batch Normalization (BN-NAS), to accelerate neural architecture search (NAS). BN-NAS can significantly reduce the time required by model training and evaluation in NAS. Specifically, for fast evaluation, we propose a BN-based indicator for predicting subnet performance at a very early training stage. The BN-based indicator further facilitates us to improve the training efficiency by only training the BN parameters during the supernet training. This is based on our observation that training the whole supernet is not necessary while training only BN parameters accelerates network convergence for network architecture search. Extensive experiments show that our method can significantly shorten the time of training supernet by more than 10 times and shorten the time of evaluating subnets by more than 600,000 times without losing accuracy. The source codes are available at https://github.com/bychen515/BNNAS. Peixia Li, Baopu Li, Chen Lin 0003, Chuming Li, Ming Sun 0008, Wanli Ouyang |
ICCV | 2 |
| 2021 | GLiT: Neural Architecture Search for Global and Local Image TransformerabstractWe introduce the first Neural Architecture Search (NAS) method to find a better transformer architecture for image recognition. Recently, transformers without CNN-based backbones are found to achieve impressive performance for image recognition. However, the transformer is designed for NLP tasks and thus could be sub-optimal when directly used for image recognition. In order to improve the visual representation ability for transformers, we propose a new search space and searching algorithm. Specifically, we introduce a locality module that models the local correlations in images explicitly with fewer computational cost. With the locality module, our search space is defined to let the search algorithm freely trade off between global and local information as well as optimizing the low-level design choice in each module. To tackle the problem caused by huge search space, a hierarchical neural architecture search method is proposed to search the optimal vision transformer from two levels separately with the evolutionary algorithm. Extensive experiments on the ImageNet dataset demonstrate that our method can find more discriminative and efficient trans-former variants than the ResNet family (e.g., ResNet101) and the baseline ViT for image classification. The source codes are available at https://github.com/bychen515/GLiT. Peixia Li, Chuming Li, Baopu Li, Lei Bai 0001, Chen Lin 0003, Ming Sun 0008, Wanli Ouyang |
ICCV | 2 |
| 2021 | Residual multi-task learning for facial landmark localization and expression recognition
Wenlong Guan, Peixia Li, Naoki Ikeda, Kosuke Hirasawa, Huchuan Lu |
Pattern Recognit. | 3 |
| 2020 | Visual tracking by dynamic matching-classification network switching
Peixia Li, Dong Wang 0004, Huchuan Lu |
Pattern Recognit. | 1 |
| 2019 | GradNet: Gradient-Guided Network for Visual Object TrackingabstractThe fully-convolutional siamese network based on template matching has shown great potentials in visual tracking. During testing, the template is fixed with the initial target feature and the performance totally relies on the general matching ability of the siamese network. However, this manner cannot capture the temporal variations of targets or background clutter. In this work, we propose a novel gradient-guided network to exploit the discriminative information in gradients and update the template in the siamese network through feed-forward and backward operations. To be specific, the algorithm can utilize the information from the gradient to update the template in the current frame. In addition, a template generalization training method is proposed to better use gradient information and avoid overfitting. To our knowledge, this work is the first attempt to exploit the information in the gradient for template update in siamese-based trackers. Extensive experiments on recent benchmarks demonstrate that our method achieves better performance than other state-of-the-art trackers. Peixia Li, Wanli Ouyang, Dong Wang 0004, Xiaoyun Yang, Huchuan Lu |
ICCV | 1 |
| 2019 | Multi attention module for visual tracking
Peixia Li, Dong Wang 0004, Gang Yang 0002, Huchuan Lu |
Pattern Recognit. | 2 |
| 2018 | Real-Time 'Actor-Critic' Tracking
Dong Wang 0004, Peixia Li, Huchuan Lu |
ECCV (7) | 3 |
| 2018 | Deep visual tracking: Review and experimental comparison
Peixia Li, Dong Wang 0004, Lijun Wang 0001, Huchuan Lu |
Pattern Recognit. | 1 |