Dashan Guo

dblp:203/6579 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
6since 2021 · last 2023
0000-0001-6833-8882ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2023 Few-Shot Class-Incremental Learning via Class-Aware Bilateral Distillation
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims to continually learn novel classes based on only few training samples, which poses a more challenging task than the well-studied Class-Incremental Learning (CIL) due to data scarcity. While knowledge distillation, a prevailing technique in CIL, can alleviate the catastrophic forgetting of older classes by regularizing outputs between current and previous model, it fails to consider the overfitting risk of novel classes in FSCIL. To adapt the powerful distillation technique for FSCIL, we propose a novel distillation structure, by taking the unique challenge of overfitting into account. Concretely, we draw knowledge from two complementary teachers. One is the model trained on abundant data from base classes that carries rich general knowledge, which can be leveraged for easing the overfitting of current novel classes. The other is the updated model from last incremental session that contains the adapted knowledge of previous novel classes, which is used for alleviating their forgetting. To combine the guidances, an adaptive strategy conditioned on the class-wise semantic similarities is introduced. Besides, for better preserving base class knowledge when accommodating novel concepts, we adopt a two-branch network with an attention-based aggregation module to dynamically merge predictions from two complementary branches. Extensive experiments on 3 popular FSCIL datasets: mini-ImageNet, CIFAR100 and CUB200 validate the effectiveness of our method by surpassing existing works by a significant margin. Code is available at https://github.com/LinglanZhao/BiDistFSCIL.
Linglan Zhao, Jing Lu 0004, Yunlu Xu, Zhanzhan Cheng, Dashan Guo, Xiangzhong Fang
CVPR5
2022 DavarOCR: A Toolbox for OCR and Multi-Modal Document Understanding
abstract
This paper presents DavarOCR, an open-source toolbox for OCR and document understanding tasks. DavarOCR currently implements 19 advanced algorithms, covering 9 different task forms. DavarOCR provides detailed usage instructions and the trained models for each algorithm. Compared with the previous open-source OCR toolbox, DavarOCR has relatively more complete support for the sub-tasks of the cutting-edge technology of document understanding. In order to promote the development and application of OCR technology in academia and industry, we pay more attention to the use of modules that different sub-domains of technology can share. DavarOCR is publicly released at https://github.com/hikopensource/Davar-Lab-OCR.
Liang Qiao 0001, Zaisheng Li, Baorui Zou, Dashan Guo, Yingda Xu, Yunlu Xu, Zhanzhan Cheng
ACM Multimedia8
2022 Boosting Few-shot visual recognition via saliency-guided complementary attention
Linglan Zhao, Dashan Guo, Wei Li 0084, Xiangzhong Fang
Neurocomputing3
2021 A Strong Baseline for Semi-Supervised Incremental Few-Shot Learning
Linglan Zhao, Dashan Guo, Yunlu Xu, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Xiangzhong Fang
BMVC2
2021 Saliency-Guided Complementary Attention for Improved Few-Shot Learning
abstract
Despite significant progress in recent deep neural networks, most deep learning algorithms rely heavily on abundant training samples. To address this problem, we propose an effective and interpretable few-shot classification model using Saliency-Guided Complementary Attention (SGCA), which aims to learn transferable representations and to build a robust classification module simultaneously. Concretely, we propose to train our feature extractor using an auxiliary task to separate object regions from background clutter guided by saliency detection signals. In addition, to make the separation beneficial to the downstream tasks, we introduce a complementary attention mechanism to force the classification module to focus on various informative parts of the image. Extensive experiments on few-shot learning tasks demonstrate the effectiveness of our proposed method, e.g., we achieve 68.81% and 84.60% for 5-way 1-shot and 5-shot settings on mini-ImageNet, respectively.
Linglan Zhao, Dashan Guo, Wei Li 0084, Xiangzhong Fang
ICME3
2021 Class-wise Metric Scaling for Improved Few-Shot Classification
abstract
Few-shot classification aims to generalize basic knowledge to recognize novel categories from a few samples. Recent centroid-based methods achieve promising classification performance with the nearest neighbor rule. However, we consider that those methods intrinsically ignore per-class distribution, as the decision boundaries are biased due to the diversity of intra-class variances. Hence, we propose a class-wise metric scaling (CMS) mechanism, which can be applied to both training and testing stages. Concretely, metric scalars are set as learnable parameters in the training stage, helping to learn a more discriminative and transferable feature representation. As for testing, we construct a convex optimization problem to generate an optimal scalar vector for refining the nearest neighbor decisions. Besides, we also involve a low-ranking bilinear pooling layer for improved representation capacity, which further provides significant performance gains. Extensive experiments are conducted on a series of feature extractor backbones, datasets, and testing modes, which have shown consistent improvements compared to prior SOTA methods, e.g., we achieve accuracies of 66.64 % and 83.63 % for 5-way 1-shot and 5-shot settings on the mini-ImageNet, respectively. Under the semi-supervised inductive mode, results are further up to 78.34 % and 87.53 %, respectively.
Linglan Zhao, Wei Li 0084, Dashan Guo, Xiangzhong Fang
WACV4
2019 Refining Proposals with Neighboring Contexts for Temporal Action Detection
abstract
While many methods have been proposed for generating temporal proposals, the performance of existing temporal detection pipelines is still limited by the quality of proposals. In this paper, we introduce a new refining model for temporal action detection, which incorporates the evaluation of Intersection-over-Union (IoU) value into the action classification, and then regresses a better segment using the neighboring contexts of one candidate proposal. To refine one candidate proposal, we augment it with two neighboring proposals of equal length, which capture the contextual information from the past and future segments. After extracting regional features for the augmented proposals, we utilize the dilated convolutions for contextual modeling to regress the offsets between the candidate proposal and the target segment within this augmented area. Extensive experiments on THUMOS14 demonstrate that our method successfully refines the generated proposals and achieves superior detection performance over other methods.
Dashan Guo, Wei Li 0084, Jianhui Sun, Xiangzhong Fang
ICME1
2018 Multimodal architecture for video captioning with memory networks and an attention mechanism
Wei Li 0084, Dashan Guo, Xiangzhong Fang
Pattern Recognit. Lett.2
2018 Fully Convolutional Network for Multiscale Temporal Action Proposals
abstract
Similar to the function of object proposals in localizing objects within images, temporal action proposals can facilitate the extraction of semantic segments and simplify the computations required for temporal action localization in untrimmed videos. In this paper, we propose a fully convolutional network to identify multiscale temporal action proposals (FCN-TAP) that utilizes only the temporal convolutions to retrieve accurate action proposals for video sequences. Using gated linear units, our network enables simple but powerful inferences, and by parallelizing the computations, it significantly improves performances compared with previous recurrent models. To capture more temporal contexts with fewer parameters, we apply dilated convolutions to expand the receptive fields of our network. Moreover, we divide the receptive fields into multiple scale ranges and then refine the corresponding temporal boundaries using duration regression at each scale. To generate suitable segments with arbitrary durations for training, we design a new strategy to select sampled candidates within the corresponding scale range. The power of our method is demonstrated through experiments on the THUMOS'14 and ActivityNet datasets, where FCN-TAP performs better and achieves a remarkable speedup compared to other state-of-the-art methods. Additional experiments show that our method generates high-quality proposals and improves the localization stage of existing action detection pipelines.
Dashan Guo, Wei Li 0084, Xiangzhong Fang
IEEE Trans. Multim.1
2017 Capturing Temporal Structures for Video Captioning by Spatio-temporal Contexts and Channel Attention Mechanism
Dashan Guo, Wei Li 0084, Xiangzhong Fang
Neural Process. Lett.1