Baopu Li

dblp:44/3801 · DBLP profile ↗
← Back
83ranked-venue papers
11as first author
58since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 46 · 3 first-author · 37 since 2021Artificial intelligence and machine learning · 44 · 7 first-author · 31 since 2021Systems, architecture and hardware · 9 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 $\beta $-DARTS++: Bi-Level Regularization for Proxy-Robust Differentiable Architecture Search
abstract
Neural Architecture Search (NAS) has attracted increasing attention in recent years because of its capability to design neural networks automatically. Among them, differential NAS approaches such as DARTS, have gained popularity for search efficiency. However, they still suffer from three main issues, that are, the weak stability due to the performance collapse, the poor generalization ability of the searched architectures, and the inferior robustness to different kinds of proxies (i.e., computationally reduced search configurations). To solve the search stability and searched architecture's generalization problems, a simple-but-effective regularization method, termed as Beta-Decay, is proposed to regularize the DARTS-based NAS searching process (referred as $\beta$β-DARTS). Specifically, Beta-Decay regularization can impose constraints to keep the value and variance of activated architecture parameters from being too large, thereby ensuring fair competition among architecture parameters and making the supernet less sensitive to the impact of input on the operation set. In-depth theoretical analyses on how it works and why it works are provided, and comprehensive experiments on a variety of search spaces and datasets validate that Beta-Decay regularization can help to stabilize the searching process and make the searched network more transferable across different datasets. To address the proxy robustness problem, we first benchmark differentiable NAS methods under a wide range of proxy data, proxy channels, proxy layers, and proxy epochs, since the robustness of NAS under different kinds of proxies has not been explored before. We then conclude some interesting findings and find that $\beta$β-DARTS always achieves the best result among all compared NAS methods under almost all proxy settings. We further introduce the novel flooding regularization to the weight optimization of $\beta$β-DARTS (termed as Bi-level regularization), and experimentally and theoretically verify its effectiveness for improving the proxy robustness of differentiable NAS.
Peng Ye 0006, Tong He 0001, Baopu Li, Tao Chen 0003, Lei Bai 0001, Wanli Ouyang
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models
abstract
Upcycled Mixture-of-Experts (MoE) models have shown great potential in various tasks by converting the original Feed-Forward Network (FFN) layers in pre-trained dense models into MoE layers. However, these models still suffer from significant parameter inefficiency due to the introduction of multiple experts. In this work, we propose a novel DeRS (Decompose, Replace, and Synthesis) paradigm to overcome this shortcoming, which is motivated by our observations about the unique redundancy mechanisms of upcycled MoE experts. Specifically, DeRS decomposes the experts into one expert-shared base weight and multiple expert-specific delta weights, and subsequently represents these delta weights in lightweight forms. Our proposed DeRS paradigm can be applied to enhance parameter efficiency in two different scenarios, including: 1) DeRS Compression for inference stage, using sparsification or quantization to compress vanilla upcycled MoE models; and 2) DeRS Up-cycling for training stage, employing lightweight sparse or low-rank matrixes to efficiently upcycle dense models into MoE models. Extensive experiments across three different tasks show that the proposed methods can achieve extreme parameter efficiency while maintaining the performance for both training and compression of upcycled MoE models.
Yongqi Huang, Peng Ye 0006, Chenyu Huang 0001, Jianjian Cao, Lin Zhang 0055, Baopu Li, Gang Yu 0002, Tao Chen 0003
CVPR6
2025 NAS-PED: Neural Architecture Search for Pedestrian Detection
abstract
Pedestrian detection currently suffers from two issues in crowded scenes: occlusion and dense boundary prediction, making it still challenging in complex real-world scenarios. In recent years, Convolutional Neural Networks (CNN) and Vision Transformers (ViT) have shown their superiorities in addressing these issues, where ViTs capture global feature dependency to infer occlusion parts and CNNs make accurate dense predictions by local detailed features. Nevertheless, limited by the narrow receptive field, CNNs fail to infer occlusion parts, while ViTs tend to ignore local features that are vital to distinguish different pedestrians in the crowd. Therefore, it is essential to combine the advantages of CNN and ViT for pedestrian detection. However, manually designing a specific CNN and ViT hybrid network requires enormous time and resources for trial and error. To address this issue, we propose the first Neural Architecture Search (NAS) framework specifically designed for pedestrian detection named NAS-PED, which automatically designs an appropriate CNNs and ViTs hybrid backbone for the crowded pedestrian detection task. Specifically, we formulate transformers and convolutions with various kernel sizes in the same format, which provides an unconstrained space for diverse hybrid network search. Furthermore, to search for a suitable backbone, we propose an information bottleneck based NAS objective function, which treats the process of NAS as an information extraction process, preserving relevant information and suppressing redundant information from the dense pedestrians in crowd scenes Extensive experiments on CrowdHuman, CityPersons and EuroCity Persons datasets demonstrate the effectiveness of the proposed method. Our NAS-PED obtains absolute gains of 4.0% MR and 1.9% AP over the state-of-the-art (SOTA) pedestrian detection framework on CrowdHuman datasets. For the CityPersons and EuroCity Persons datasets, the searched backbone achieves stable improvement across all three subsets, outperforming some large language-image pre-trained models. Code will be released after acceptance.
Min Liu 0008, Baopu Li, Yaonan Wang 0001, Wanli Ouyang
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Stimulative Training++: Go Beyond the Performance Limits of Residual Networks
abstract
Residual networks have shown great success and become indispensable in recent deep neural network models. In this work, we aim to re-investigate the training process of residual networks from a novel perspective of loafing, and further propose a new training scheme as well as three improved strategies for boosting residual networks beyond their performance limits. Previous research has suggested that residual networks can be considered as ensembles of shallow networks, which implies that the final performance of a residual network is influenced by a group of subnetworks. Furthermore, we identify a previously overlooked problem, where subnetworks within a residual network are prone to exert less effort when working as part of a group compared to working alone. We define this problem as network loafing. Since network loafing may inevitably cause the sub-par performance of the residual network, we propose a novel training scheme called stimulative training, which randomly samples a residual subnetwork and calculates the KL divergence loss between the sampled subnetwork and the given residual network for extra supervision. In order to unleash the potential of stimulative training, we further propose three simple-yet-effective strategies, including a novel KL- loss that only aligns the network logits direction, random smaller inputs for subnetworks, and inter-stage sampling rules. Comprehensive experiments and analysis verify the effectiveness of stimulative training as well as its three improved strategies. For example, the proposed method can boost the performance of ResNet50 on ImageNet to 80.5% Top1 accuracy without using any extra data, model, trick, or changing the structure. With only uniform augment, the performance can be further improved to 81.0% Top1 accuracy, better than the best training recipes provided by Timm library and PyTorch official version. We also verify its superiority on various typical models, datasets, and tasks and give some theoretical analysis. As such, we advocate utilizing the proposed method as a general and next-generation technology to train residual networks.
Peng Ye 0006, Tong He 0001, Shengji Tang, Baopu Li, Tao Chen 0003, Lei Bai 0001, Wanli Ouyang
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 BridgeNet: Comprehensive and Effective Feature Interactions via Bridge Feature for Multi-Task Dense Predictions
abstract
Multi-task dense prediction aims at handling multiple pixel-wise prediction tasks within a unified network simultaneously for visual scene understanding. However, cross-task feature interactions of current methods are still suffering from incomplete levels of representations, less discriminative semantics in feature participants, and inefficient pair-wise task interaction processes. To tackle these under-explored issues, we propose a novel BridgeNet framework, which extracts comprehensive and discriminative intermediate Bridge Features, and conducts interactions based on them. Specifically, a Task Pattern Propagation (TPP) module is first applied to ensure highly semantic task-specific feature participants are prepared for subsequent interactions, and a Bridge Feature Extractor (BFE) is specially designed to selectively integrate both high-level and low-level representations to generate the comprehensive bridge features. Then, instead of conducting heavy pair-wise cross-task interactions, a Task-Feature Refiner (TFR) is developed to efficiently take guidance from bridge features and form final task predictions. To the best of our knowledge, this is the first work considering the completeness and quality of feature participants in cross-task interactions. Extensive experiments are conducted on NYUD-v2, Cityscapes and PASCAL Context benchmarks, and the superior performance shows the proposed architecture is effective and powerful in promoting different dense prediction tasks simultaneously.
Jingdong Zhang 0003, Jiayuan Fan 0001, Peng Ye 0006, Bo Zhang 0069, Hancheng Ye, Baopu Li, Yancheng Cai, Tao Chen 0003
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Efficient Architecture Search via Bi-Level Data Pruning
abstract
Improving the efficiency of Neural Architecture Search (NAS) is a challenging but significant task that has received much attention. Previous studies mainly adopt the Differentiable Architecture Search (DARTS) and improve its search strategies or modules to enhance search efficiency. Recently, some methods have started considering data reduction for speedup, but they are not tightly coupled with the architecture search process and cannot capture the training dynamics of DARTS well, resulting in sub-optimal performances. To this end, this work pioneers an exploration into the critical role of dataset characteristics in the bi-level optimization of DARTS, and then proposes a novel Bi-level Data Pruning (BDP) paradigm that targets the weights and architecture levels of DARTS to enhance efficiency from a data perspective. Specifically, we introduce a progressive bi-level data pruning strategy that utilizes supernet prediction dynamics as the metric to gradually prune unsuitable samples for DARTS during the search. An effective automatic class balance constraint is also integrated into BDP, to suppress potential class imbalances resulting from data-efficient algorithms. Comprehensive evaluations on the NAS-Bench-201 search space, DARTS search space, and MobileNet-like search space validate that BDP reduces search costs by over 50% while achieving superior performance when applied to the baseline DARTS. Besides, we demonstrate that BDP can harmoniously integrate with advanced DARTS variants, like P-DARTS, PC-DARTS, EG-NAS, and$\beta $-DARTS, offering an approximately$2\times $speedup with minimal performance compromise.
Chongjun Tu, Peng Ye 0006, Weihao Lin 0002, Hancheng Ye, Chong Yu 0001, Tao Chen 0003, Baopu Li, Wanli Ouyang
IEEE Trans. Circuits Syst. Video Technol.7
2024 Boosting Residual Networks with Group Knowledge
abstract
Recent research understands the residual networks from a new perspective of the implicit ensemble model. From this view, previous methods such as stochastic depth and stimulative training have further improved the performance of the residual network by sampling and training of its subnets. However, they both use the same supervision for all subnets of different capacities and neglect the valuable knowledge generated by subnets during training. In this manuscript, we mitigate the significant knowledge distillation gap caused by using the same kind of supervision and advocate leveraging the subnets to provide diverse knowledge. Based on this motivation, we propose a group knowledge based training framework for boosting the performance of residual networks. Specifically, we implicitly divide all subnets into hierarchical groups by subnet-in-subnet sampling, aggregate the knowledge of different subnets in each group during training, and exploit upper-level group knowledge to supervise lower-level subnet group. Meanwhile, we also develop a subnet sampling strategy that naturally samples larger subnets, which are found to be more helpful than smaller subnets in boosting performance for hierarchical groups. Compared with typical subnet training and other methods, our method achieves the best efficiency and performance trade-offs on multiple datasets and network structures. The code is at https://github.com/tsj-001/AAAI24-GKT.
Shengji Tang, Peng Ye 0006, Baopu Li, Weihao Lin 0002, Tao Chen 0003, Tong He 0001, Chong Yu 0001, Wanli Ouyang
AAAI3
2024 3D Multi-frame Fusion for Video Stabilization
abstract
In this paper, we present RStab, a novel framework for video stabilization that integrates 3D multi-frame fusion through volume rendering. Departing from conventional methods, we introduce a 3D multi-frame perspective to generate stabilized images, addressing the challenge of full-frame generation while preserving structure. The core of our RStab framework lies in Stabilized Rendering (SR), a volume rendering module, fusing multi-frame information in 3D space. Specifically, SR involves warping features and colors from multiple frames by projection, fusing them into descriptors to render the stabilized image. However, the precision of warped information depends on the projection accuracy, a factor significantly influenced by dynamic regions. In response, we introduce the Adaptive Ray Range (ARR) module to integrate depth priors, adaptively defining the sampling range for the projection process. Additionally, we propose Color Correction (CC) assisting geometric constraints with optical flow for accurate color aggregation. Thanks to the three modules, our RStab demonstrates superior performance compared with previous stabilizers in the field of view (FOV), image quality, and video stability across various datasets.
Weiyue Zhao, Tianqi Liu 0003, Huiqiang Sun, Baopu Li, Zhiguo Cao 0001
CVPR6
2024 DreamMover: Leveraging the Prior of Diffusion Models for Image Interpolation with Large Motion
Liao Shen, Tianqi Liu 0003, Huiqiang Sun, Baopu Li, Jianming Zhang 0001, Zhiguo Cao 0001
ECCV (15)5
2024 Enhanced Sparsification via Stimulative Training
Shengji Tang, Weihao Lin 0002, Hancheng Ye, Peng Ye 0006, Chong Yu 0001, Baopu Li, Tao Chen 0003
ECCV (51)6
2024 Robust 3D Face Alignment with Multi-Path Neural Architecture Search
abstract
3D face alignment is a very challenging and fundamental problem in computer vision. Existing deep learning-based methods manually design different networks to regress either parameters of a 3D face model or 3D positions of face vertices. However, designing such networks relies on expert knowledge, and these methods often struggle to produce consistent results across various face poses. To address this limitation, we employ Neural Architecture Search (NAS) to automatically discover the optimal architecture for 3D face alignment. We propose a novel Multi-path One-shot Neural Architecture Search (MONAS) framework that leverages multi-scale features and contextual information to enhance face alignment across various poses. The MONAS comprises two key algorithms: Multi-path Networks Unbiased Sampling Based Training and Simulated Annealing based Multi-path One-shot Search. Experimental results on three popular benchmarks demonstrate the superior performance of the MONAS for both sparse alignment and dense alignment.
Zhichao Jiang, Hongsong Wang 0001, Xi Teng, Baopu Li
ICME4
2024 DeNKD: Decoupled Non-Target Knowledge Distillation for Complementing Transformer-Based Unsupervised Domain Adaptation
abstract
There is a growing need to explore the potential of transformers in Unsupervised Domain Adaptation (UDA) due to their increasing success in various vision tasks. However, the application of transformers in UDA has yet to be thoroughly investigated and requires further research. In this study, our primary focus is to design a novel pipeline specifically tailored for transformer-based UDA, to address a crucial challenge: the overemphasis on the transfer of target-oriented information, mainly caused by the self-attention blocks in transformers and the cross-domain adversarial learning scheme. First, we show that non-target information, including semantic contextual information such as background features and non-target classes, must be addressed in the domain adaptation process. Recognizing the importance of incorporating non-target knowledge, we propose a decoupled non-target knowledge distillation method called DeNKD. DeNKD decouples non-target information across domains at both feature and logit levels. This decoupling is achieved through a bi-directional knowledge distillation approach that facilitates the interaction and exchange of non-target knowledge to facilitate an effective transformer-based cross-domain knowledge transfer. We perform extensive evaluations on several well-established UDA benchmark datasets. The results consistently show that DeNKD outperforms other methods, achieving the best performance across the board. For example, on the Office-Home dataset, DeNKD achieves an accuracy of 85.54%, while on the VisDA-2017 dataset, it achieves an accuracy of 89.95%. These results highlight the effectiveness of DeNKD in transformer-based UDA and its potential for improving cross-domain adaptation performance.
Peng Ye 0006, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001, Wanli Ouyang
IEEE Trans. Circuits Syst. Video Technol.3
2024 Multi-View Vision Fusion Network: Can 2D Pre-Trained Model Boost 3D Point Cloud Data-Scarce Learning?
abstract
Point cloud based 3D deep model has wide applications in many applications such as autonomous driving, house robot, etc. Inspired by the recent prompt learning in natural language processing, this work proposes a novel Multi-view Vision Fusion Network (MvNet) for few-shot 3D point cloud classification. MvNet investigates the possibility of leveraging the off-the-shelf 2D pre-trained models to achieve the few-shot classification, which can alleviate the over-dependence issue of the existing baseline models towards the large-scale annotated 3D point cloud data. Specifically, MvNet first encodes a 3D point cloud into multi-view image features for a number of different views. Then, a novel multi-view prompt fusion module is developed to fuse information from different views effectively to bridge the gap between 3D point cloud data and 2D pre-trained models. A set of 2D image prompts can then be derived to better describe the suitable prior knowledge for a large-scale pre-trained image model for few-shot 3D point cloud classification. Extensive experiments on ModelNet, ScanObjectNN, and ShapeNet datasets demonstrate that MvNet achieves new state-of-the-art performance for 3D few-shot point cloud image classification. The source code of this work is available at https://github.com/invictus717/MetaTransformer.
Haoyang Peng, Baopu Li, Bo Zhang 0069, Xin Chen 0040, Tao Chen 0003, Hongyuan Zhu 0002
IEEE Trans. Circuits Syst. Video Technol.2
2024 Joint Distribution Adaptive-Alignment for Cross-Domain Segmentation of High-Resolution Remote Sensing Images
abstract
Although existing unsupervised domain adaptation (UDA) methods have successfully applied to semantic segmentation tasks for high-resolution remote sensing (HRS) images, they still have some limitations that need to be addressed: 1) they mainly focus on aligning the marginal distributions while ignoring the interdomain differences in the conditional distributions, which may be suboptimal because they assume that the boundaries of category decision are identical across domains; and 2) they depend on self-supervised learning for easy-to-hard alignment, which may result in model learning erroneous knowledge from the pseudo labels. To address the above limitations, we propose a joint distribution adaptive-alignment framework (JDAF) to eliminate the distribution difference between the source and target domains, which is mainly composed of a marginal distribution alignment (MDA) module, a conditional distribution alignment (CDA) module, and an improved easy-to-hard adaptation strategy. The MDA module is used to narrow local semantic and global spatial differences between domains and first advance, and then, the CDA module that includes a category-invariant feature alignment (CFA) block and a dataset-level context aggregation (DCA) block is presented and designed, which can dynamically update and align the feature representations that are invariant to category change and adaptively incorporate dataset-level context into the features of source domain to enhance the pixel-level representation. An uncertainty-adaptive learning (UAL) method is, moreover, proposed to improve the easy-to-hard adaptation strategy by enabling the model to learn accurate knowledge from the pseudo labels, which can boost the adaptive performance of the whole JDAF. Comprehensive experiments with four cross-domain tasks on two benchmark datasets of aerospace HRS images demonstrate that the proposed JDAF achieves significant performance gains compared to the state-of-the-art cross-domain semantic segmentation methods. Our code is available at:https://github.com/maple-hx/JDAF.
Baopu Li, Tao Chen 0003, Bin Wang 0008
IEEE Trans. Geosci. Remote. Sens.2
2023 ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design
abstract
Vision Transformers (ViTs) have achieved state-of-the-art performance on various vision tasks. However, ViTs’ self-attention module is still arguably a major bottleneck, limiting their achievable hardware efficiency and more extensive applications to resource constrained platforms. Meanwhile, existing accelerators dedicated to NLP Transformers are not optimal for ViTs. This is because there is a large difference between ViTs and Transformers for natural language processing (NLP) tasks: ViTs have a relatively fixed number of input tokens, whose attention maps can be pruned by up to 90% even with fixed sparse patterns, without severely hurting the model accuracy (e.g.,=50%). To this end, we propose a dedicated algorithm and accelerator co-design framework dubbed ViTCoD for accelerating ViTs. Specifically, on the algorithm level, ViTCoD prunes and polarizes the attention maps to have either denser or sparser fixed patterns for regularizing two levels of workloads without hurting the accuracy, largely reducing the attention computations while leaving room for alleviating the remaining dominant data movements; on top of that, we further integrate a lightweight and learnable auto-encoder module to enable trading the dominant high-cost data movements for lower-cost computations. On the hardware level, we develop a dedicated accelerator to simultaneously coordinate the aforementioned enforced denser and sparser workloads for boosted hardware utilization, while integrating on-chip encoder and decoder engines to leverage ViTCoD’s algorithm pipeline for much reduced data movements. Extensive experiments and ablation studies validate that ViTCoD largely reduces the dominant data movement costs, achieving speedups of up to 235.3×, 142.9×, 86.0×, 10.1×, and 6.8× over general computing platforms CPUs, EdgeGPUs, GPUs, and prior-art Transformer accelerators SpAtten and Sanger under an attention sparsity of 90%, respectively. Our code implementation is available at https://github.com/GATECH-EIC/ViTCoD.
Haoran You, Zhanyi Sun, Huihong Shi, Zhongzhi Yu, Yang Zhao 0013, Yongan Zhang, Chaojian Li, Baopu Li, Yingyan (Celine) Lin
HPCA8
2023 MRM: Masked Relation Modeling for Medical Image Pre-Training with Genetics
abstract
Modern deep learning techniques on automatic multi-modal medical diagnosis rely on massive expert annotations, which is time-consuming and prohibitive. Recent masked image modeling (MIM)-based pre-training methods have witnessed impressive advances for learning meaningful representations from unlabeled data and transferring to downstream tasks. However, these methods focus on natural images and ignore the specific properties of medical data, yielding unsatisfying generalization performance on downstream medical diagnosis. In this paper, we aim to leverage genetics to boost image pre-training and present a masked relation modeling (MRM) framework. Instead of explicitly masking input data in previous MIM methods leading to loss of disease-related semantics, we design relation masking to mask out token-wise feature relation in both self- and cross-modality levels, which preserves intact semantics within the input and allows the model to learn rich disease-related information. Moreover, to enhance semantic relation modeling, we propose relation matching to align the sample-wise relation between the intact and masked features. The relation matching exploits inter-sample relation by encouraging global constraints in the feature space to render sufficient semantic relation for feature representation. Extensive experiments demonstrate that the proposed framework is simple yet powerful, achieving state-of-the-art transfer performance on various downstream diagnosis tasks. Codes are available at https://github.com/CityU-AIM-Group/MRM.
Qiushi Yang, Wuyang Li, Baopu Li, Yixuan Yuan
ICCV3
2023 Rethinking Pseudo-Label-Based Unsupervised Person Re-ID with Hierarchical Prototype-based Graph
abstract
Unsupervised person re-identification (Re-ID) aims to match individuals without manual annotations. However, existing methods often struggle with intra-class variations due to differences in person poses and camera styles such as resolution and environment information. Additionally, clustering may produce incorrect pseudo-labels, compounding the issue. To address these challenges, we propose a novel hierarchical prototype-based graph network (HPG-Net) for unsupervised person Re-ID. Our approach uses a hierarchical prototype-based graph structure to describe person images by attributes of poses and camera styles, with each graph node representing the average of image features as a prototype. We then apply a hierarchical contrastive learning module to enhance the feature learning at each level, reducing the impact of intra-class differences caused by extraneous attributes. We also calculate the similarity between samples and each level of prototypes, maintaining prototype-based graph consistency with the mean-teacher network to mitigate the accumulation errors caused by pseudo-labels. Experimental results on three benchmarks show that our method outperforms state-of-the-art (SOTA) works. Moreover, we achieve promising performance on an occluded dataset.
Ben Sha, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001
ACM Multimedia2
2023 Effective Invertible Arbitrary Image Rescaling
abstract
Great successes have been achieved using deep learning techniques for image super-resolution (SR) with fixed scales. To increase its real world applicability, numerous models have also been proposed to restore SR images with arbitrary scale factors, including asymmetric ones where images are resized to different scales along horizontal and vertical directions. Though most models are only optimized for the unidirectional upscaling task while assuming a predefined downscaling kernel for low-resolution (LR) inputs, recent models based on Invertible Neural Networks (INN) are able to increase upscaling accuracy significantly by optimizing the downscaling and upscaling cycle jointly. However, limited by the INN architecture, it is constrained to fixed integer scale factors and requires one model for each scale. Without increasing model complexity, a simple and effective invertible arbitrary rescaling network (IARN) is proposed to achieve arbitrary image rescaling by training only one model in this work. Using innovative components like position-aware scale encoding and preemptive channel splitting, the network is optimized to convert the non-invertible rescaling cycle to an effectively invertible process. It is shown to achieve a state-of-the-art (SOTA) performance in bidirectional arbitrary rescaling without compromising perceptual quality in LR outputs. It is also demonstrated to perform well on tests with asymmetric scales using the same network architecture.
Zhihong Pan 0001, Baopu Li, Dongliang He, Errui Ding
WACV2
2023 Adversarial learning based intermediate feature refinement for semantic segmentation
Dongli Wang, Zhitian Yuan, Wanli Ouyang, Baopu Li, Yan Zhou 0003
Appl. Intell.4
2023 Generalized Gradient Flow Based Saliency for Pruning Deep Convolutional Neural Networks
Xinyu Liu 0001, Baopu Li, Zhen Chen 0013, Yixuan Yuan
Int. J. Comput. Vis.2
2023 A Strip Dilated Convolutional Network for Semantic Segmentation
Yan Zhou 0003, Xihong Zheng, Wanli Ouyang, Baopu Li
Neural Process. Lett.4
2023 Automatic Loss Function Search for Adversarial Unsupervised Domain Adaptation
abstract
Unsupervised domain adaption (UDA) aims to reduce the domain gap between labeled source and unlabeled target domains. Many prior works exploit adversarial learning that leverages pre-designed discriminators to drive the network for aligning distributions between domains. However, most of them do not consider the degeneration of the domain discriminators caused by the gradually dominating gradients of aligned target samples during training, and they still suffer from the cross-domain semantic mismatch problem in the learned feature space. Hence, this paper attempts to understand and solve both issues from the lens of optimization loss and propose an automatic loss function search for adversarial domain adaptation (ALSDA). First, we extend the common adversarial loss by adding an adjustable hyper-parameter that can re-weight the gradients assigned to target samples, so that the domain discriminator can impose consecutive and influential driving forces for domain alignment. Meanwhile, we upgrade the traditional orthogonality loss with class-wisely adjustable hyper-parameters that can strengthen the cross-domain feature separation. Since manually determining the optimal loss functions requires expensive expert efforts, we leverage the popular AutoML to automatically search for the optimal loss functions from a pre-defined novel and unique search space for UDA. Further, to enable the loss function search when the target domain is unlabeled, we introduce a simple-but-effective entropy-guided search strategy with the aid of REINFORCE learning. Extensive experiments on various typical baselines and benchmark datasets such as Office-Home, Office-31, and Birds-31 have been conducted, and the results validate the generalization and superiority of the proposed ALSDA.
Peng Ye 0006, Hancheng Ye, Baopu Li, Jinyang Guo 0002, Tao Chen 0003, Wanli Ouyang
IEEE Trans. Circuits Syst. Video Technol.4
2023 Iterative Class Prototype Calibration for Transductive Zero-Shot Learning
abstract
Zero-shot learning (ZSL) typically suffers from the domain shift issue since the projected feature embedding of unseen samples mismatch with the corresponding class semantic prototypes, making it very challenging to fine-tune an optimal visual-semantic mapping for the unseen domain. Some existing transductive ZSL methods solve this problem by introducing unlabeled samples of the unseen domain, in which the projected features of unseen samples are still not discriminative and tend to be distributed around prototypes of seen classes. Therefore, how to effectively align the projection features of samples in unseen classes with corresponding predefined class prototypes is crucial for promoting the generalization of ZSL models. In this paper, we propose a novel Iterative Class Prototype Calibration (ICPC) framework for transductive ZSL which consists of a pseudo-labeling stage and a model retraining stage to address the above key issue. First, in the labeling stage, we devise a Class Prototype Calibration (CPC) module to calibrate the predefined class prototypes of the unseen domain by estimating the real center of projected feature distribution, which achieves better matching of sample points and class prototypes. Next, in the retraining stage, we devise a Certain Samples Screening (CSS) module to select relatively certain unseen samples with high confidence and align them with predefined class prototypes in the embedding space. A progressive training strategy is adopted to select more certain samples and update the proposed model with augmented training data. Extensive experiments on AwA2, CUB, and SUN datasets demonstrate that the proposed scheme achieves new state-of-the-art in the conventional setting under both standard split (SS) and proposed split (PS).
Hairui Yang, Baoli Sun, Baopu Li, Caifei Yang, Zhihui Wang 0001, Jenhui Chen, Lei Wang 0005
IEEE Trans. Circuits Syst. Video Technol.3
2023 Rethinking Cross-Domain Pedestrian Detection: A Background-Focused Distribution Alignment Framework for Instance-Free One-Stage Detectors
abstract
Cross-domain pedestrian detection aims to generalize pedestrian detectors from one label-rich domain to another label-scarce domain, which is crucial for various real-world applications. Most recent works focus on domain alignment to train domain-adaptive detectors either at the instance level or image level. From a practical point of view, one-stage detectors are faster. Therefore, we concentrate on designing a cross-domain algorithm for rapid one-stage detectors that lacks instance-level proposals and can only perform image-level feature alignment. However, pure image-level feature alignment causes the foreground-background misalignment issue to arise, i.e., the foreground features in the source domain image are falsely aligned with background features in the target domain image. To address this issue, we systematically analyze the importance of foreground and background in image-level cross-domain alignment, and learn that background plays a more critical role in image-level cross-domain alignment. Therefore, we focus on cross-domain background feature alignment while minimizing the influence of foreground features on the cross-domain alignment stage. This paper proposes a novel framework, namely, background-focused distribution alignment (BFDA), to train domain adaptive one-stage pedestrian detectors. Specifically, BFDA first decouples the background features from the whole image feature maps and then aligns them via a novel long-short-range discriminator. Extensive experiments demonstrate that compared to mainstream domain adaptation technologies, BFDA significantly enhances cross-domain pedestrian detection performance for either one-stage or two-stage detectors. Moreover, by employing the efficient one-stage detector (YOLOv5), BFDA can reach 217.4 FPS ( 640×480 pixels) on NVIDIA Tesla V100 (7~12 times the FPS of the existing frameworks), which is highly significant for practical applications. The code from this study will be made publicly available.
Yancheng Cai, Bo Zhang 0069, Baopu Li, Tao Chen 0003, Hongliang Yan, Jingdong Zhang 0003
IEEE Trans. Image Process.3
2023 A Closer Look at the Joint Training of Object Detection and Re-Identification in Multi-Object Tracking
abstract
Unifying object detection and re-identification (ReID) into a single network enables faster multi-object tracking (MOT), while this multi-task setting poses challenges for training. In this work, we dissect the joint training of detection and ReID from two dimensions: label assignment and loss function. We find previous works generally overlook them and directly borrow the practices from object detection, inevitably causing inferior performance. Specifically, we identify a qualified label assignment for MOT should: 1) have the assignment cost aware of ReID cost, not just detection cost; 2) provide sufficient positive samples for robust feature learning while avoiding ambiguous positives (i.e., the positives shared by different ground-truth objects). To achieve the above goals, we first propose Identity-aware Label Assignment, which jointly considers the assignment cost of detection and ReID to select positive samples for each instance without ambiguities. Moreover, we advance a novel Discriminative Focal Loss that integrates ReID predictions with Focal Loss to focus the training on the discriminative samples. Finally, we upgrade the strong baseline FairMOT with our techniques and achieve up to 7.0 MOTA / 54.1% IDs improvements on MOT16/17/20 benchmarks under favorable inference speed, which verifies our tailored label assignment and loss function for MOT are superior to those inherited from object detection.
Tianyi Liang 0001, Baopu Li, Mengzhu Wang, Huibin Tan, Zhigang Luo
IEEE Trans. Image Process.2
2023 OTP-NMS: Toward Optimal Threshold Prediction of NMS for Crowded Pedestrian Detection
abstract
Pedestrian detection is still a challenging task for computer vision, especially in crowded scenes where the overlaps between pedestrians tend to be large. The non-maximum suppression (NMS) plays an important role in removing the redundant false positive detection proposals while retaining the true positive detection proposals. However, the highly overlapped results may be suppressed if the threshold of NMS is lower. Meanwhile, a higher threshold of NMS will introduce a larger number of false positive results. To solve this problem, we propose an optimal threshold prediction (OTP) based NMS method that predicts a suitable threshold of NMS for each human instance. First, a visibility estimation module is designed to obtain the visibility ratio. Then, we propose a threshold prediction subnet to determine the optimal threshold of NMS automatically according to the visibility ratio and classification score. Finally, we re-formulate the objective function of the subnet and utilize the reward-guided gradient estimation algorithm to update the subnet. Comprehensive experiments on CrowdHuman and CityPersons show the superior performance of the proposed method in pedestrian detection, especially in crowded scenes.
Min Liu 0008, Baopu Li, Yaonan Wang 0001, Wanli Ouyang
IEEE Trans. Image Process.3
2022 Towards Bidirectional Arbitrary Image Rescaling: Joint Optimization and Cycle Idempotence
abstract
Deep learning based single image super-resolution models have been widely studied and superb results are achieved in upscaling low-resolution images with fixed scale factor and downscaling degradation kernel. To improve real world applicability of such models, there are growing interests to develop models optimized for arbitrary upscaling factors. Our proposed method is the first to treat arbitrary rescaling, both upscaling and downscaling, as one unified process. Using joint optimization of both directions, the proposed model is able to learn upscaling and downscaling simultaneously and achieve bidirectional arbitrary image rescaling. It improves the performance of current arbitrary upscaling models by a large margin while at the same time learns to maintain visual perception quality in downscaled images. The proposed model is further shown to be robust in cycle idempotence test, free of severe degradations in reconstruction accuracy when the downscaling-to-upscaling cycle is applied repetitively. This robustness is beneficial for image rescaling in the wild when this cycle could be applied to one image for multiple times. It also performs well on tests with arbitrary large scales and asymmetric scales, even when the model is not trained with such tasks. Extensive experiments are conducted to demonstrate the superior performance of our model.
Zhihong Pan 0001, Baopu Li, Dongliang He, Mingde Yao, Xin Li 0106, Errui Ding
CVPR2
2022 Towards Robust Adaptive Object Detection under Noisy Annotations
abstract
Domain Adaptive Object Detection (DAOD) models a joint distribution of images and labels from an annotated source domain and learns a domain-invariant transformation to estimate the target labels with the given target domain images. Existing methods assume that the source domain labels are completely clean, yet large-scale datasets often contain error-prone annotations due to instance ambiguity, which may lead to a biased source distribution and severely degrade the performance of the domain adaptive detector de facto. In this paper, we represent the first effort to formulate noisy DAOD and propose a Noise Latent Transferability Exploration (NLTE) framework to address this issue. It is featured with 1) Potential Instance Mining (PIM), which leverages eligible proposals to recapture the miss-annotated instances from the background; 2) Morphable Graph Relation Module (MGRM), which models the adaptation feasibility and transition probability of noisy samples with relation matrices; 3) Entropy-Aware Gradient Reconcilement (EAGR), which incorporates the semantic information into the discrimination process and enforces the gradients provided by noisy and clean samples to be consistent towards learning domain-invariant representations. A thorough evaluation on benchmark DAOD datasets with noisy source annotations validates the effectiveness of NLTE. In particular, NLTE improves the mAP by 8.4% under 60% corrupted annotations and even approaches the ideal upper bound of training on a clean source dataset.11Code is available at https://github.com/CityU-AIM-Group/NLTE.
Xinyu Liu 0001, Wuyang Li, Qiushi Yang, Baopu Li, Yixuan Yuan
CVPR4
2022 β-DARTS: Beta-Decay Regularization for Differentiable Architecture Search
abstract
Neural Architecture Search (NAS) has attracted increasingly more attention in recent years because of its capability to design deep neural network automatically. Among them, differential NAS approaches such as DARTS, have gained popularity for the search efficiency. However, they suffer from two main issues, the weak robustness to the performance collapse and the poor generalization ability of the searched architectures. To solve these two problems, a simple-but-efficient regularization method, termed as Beta-Decay, is proposed to regularize the DARTS-based NAS searching process. Specifically, Beta-Decay regularization can impose constraints to keep the value and variance of activated architecture parameters from too large. Furthermore, we provide in-depth theoretical analysis on how it works and why it works. Experimental results on NAS-Bench-201 show that our proposed method can help to stabilize the searching process and makes the searched network more transferable across different datasets. In addition, our search scheme shows an outstanding property of being less dependent on training time and data. Comprehensive experiments on a variety of search spaces and datasets validate the effectiveness of the proposed method. The code is available at https://github.com/Sunshine-Ye/Beta-DARTS.
Peng Ye 0006, Baopu Li, Yikang Li 0002, Tao Chen 0003, Jiayuan Fan 0001, Wanli Ouyang
CVPR2
2022 SuperTickets: Drawing Task-Agnostic Lottery Tickets from Supernets via Jointly Architecture Searching and Parameter Pruning
Haoran You, Baopu Li, Zhanyi Sun, Xu Ouyang, Yingyan (Celine) Lin
ECCV (11)2
2022 Field-aware Variational Autoencoders for Billion-scale User Representation Learning
abstract
User representation learning plays an essential role in Internet applications, such as recommender systems. Though developing a universal embedding for users is demanding, only few previous works are conducted in an unsupervised learning manner. The unsupervised method is however important as most of the user data is collected without specific labels. In this paper, we harness the unsupervised advantages of Variational Autoencoders (VAEs), to learn user representation from large-scale, high-dimensional, and multi-field data. We extend the traditional VAE by developing Field-aware VAE (FVAE) to model each feature field with an independent multinomial distribution. To reduce the complexity in training, we employ dynamic hash tables, a batched softmax function, and a feature sampling strategy to improve the efficiency of our method. We conduct experiments on multiple datasets, showing that the proposed FVAE significantly outperforms baselines on several tasks of data reconstruction and tag prediction. Moreover, we deploy the proposed method in real-world applications and conduct online A/B tests in a look-alike system. Results demonstrate that our method can effectively improve the quality of recommendation. To the best of our knowledge, it is the first time that the VAE-based user representation learning model is applied to real-world recommender systems.
Ge Fan, Chaoyun Zhang, Junyang Chen 0001, Baopu Li, Zenglin Xu, Luyu Peng, Zhiguo Gong
ICDE4
2022 ShiftAddNAS: Hardware-Inspired Search for More Accurate and Efficient Neural Networks
abstract
Neural networks (NNs) with intensive multiplications (e.g., convolutions and transformers) are powerful yet power hungry, impeding their more extensive deployment into resource-constrained edge devices. As such, multiplication-free networks, which follow a common practice in energy-efficient hardware implementation to parameterize NNs with more efficient operators (e.g., bitwise shifts and additions), have gained growing attention. However, multiplication-free networks in general under-perform their vanilla counterparts in terms of the achieved accuracy. To this end, this work advocates hybrid NNs that consist of both powerful yet costly multiplications and efficient yet less powerful operators for marrying the best of both worlds, and proposes ShiftAddNAS, which can automatically search for more accurate and more efficient NNs. Our ShiftAddNAS highlights two enablers. Specifically, it integrates (1) the first hybrid search space that incorporates both multiplication-based and multiplication-free operators for facilitating the development of both accurate and efficient hybrid NNs; and (2) a novel weight sharing strategy that enables effective weight sharing among different operators that follow heterogeneous distributions (e.g., Gaussian for convolutions vs. Laplacian for add operators) and simultaneously leads to a largely reduced supernet size and much better searched networks. Extensive experiments and ablation studies on various models, datasets, and tasks consistently validate the effectiveness of ShiftAddNAS, e.g., achieving up to a +7.7% higher accuracy or a +4.9 better BLEU score as compared to state-of-the-art expert-designed and neural architecture searched NNs, while leading to up to 93% or 69% energy and latency savings, respectively. Codes and pretrained models are available at https://github.com/RICE-EIC/ShiftAddNAS.
Haoran You, Baopu Li, Huihong Shi, Yonggan Fu, Yingyan (Celine) Lin
ICML2
2022 Bayesian based Re-parameterization for DNN Model Pruning
abstract
Filter pruning, as an effective strategy to obtain efficient compact structures from over-parametric deep neural networks(DNN), has attracted a lot of attention. Previous pruning methods select channels for pruning by developing different criteria, yet little attention has been devoted to whether these criteria can represent correlations between channels. Meanwhile, most existing methods generally ignore the parameters being pruned and only perform additional training on the retained network to reduce accuracy loss. In this paper, we present a novel perspective of re-parametric pruning by Bayesian estimation. First, we estimate the probability distribution of different channels based on Bayesian estimation and indicate the importance of the channels by the discrepancy in the distribution before and after channel pruning. Second, to minimize the variation in distribution after pruning, we re-parameterize the pruned network based on the probability distribution to pursue optimal pruning. We evaluate our approach on popular datasets with some typical network architectures, and comprehensive experimental results validate that this method illustrates better performance compared to the state-of-the-art approaches.
Xiaotong Lu, Teng Xi, Baopu Li, Weisheng Dong, Guangming Shi
ACM Multimedia3
2022 DEAL: An Unsupervised Domain Adaptive Framework for Graph-level Classification
abstract
Graph neural networks (GNNs) have achieved state-of-the-art results on graph classification tasks. They have been primarily studied in cases of supervised end-to-end training, which requires abundant task-specific labels. Unfortunately, annotating labels of graph data could be prohibitively expensive or even impossible in many applications. An effective solution is to incorporate labeled graphs from a different, but related source domain, to develop a graph classification model for the target domain. However, the problem of unsupervised domain adaptation for graph classification is challenging due to potential domain discrepancy in graph space as well as the label scarcity in the target domain. In this paper, we present a novel GNN framework named DEAL by incorporating both source graphs and target graphs, which is featured by two modules, i.e., adversarial perturbation and pseudo-label distilling. Specifically, to overcome domain discrepancy, we equip source graphs with target semantics by applying to them adaptive perturbations which are adversarially trained against a domain discriminator. Additionally, DEAL explores distinct feature spaces at different layers of the GNN encoder, which emphasize global and local semantics respectively. Then, we distill the consistent predictions from two spaces to generate reliable pseudo-labels for sufficiently utilizing unlabeled data, which further improves the performance of graph classification. Extensive experiments on a wide range of graph classification datasets reveal the effectiveness of our proposed DEAL.
Li Shen 0008, Baopu Li, Mengzhu Wang, Xiao Luo 0001, Chong Chen 0002, Zhigang Luo, Xian-Sheng Hua 0001
ACM Multimedia3
2022 Learning Cross-Image Object Semantic Relation in Transformer for Few-Shot Fine-Grained Image Classification
abstract
Few-shot fine-grained learning aims to classify a query image into one of a set of support categories with fine-grained differences. Although learning different objects' local differences via Deep Neural Networks has achieved success, how to exploit the query-support cross-image object semantic relations in Transformer-based architecture remains under-explored in the few-shot fine-grained scenario. In this work, we propose a Transformer-based double-helix model, namely HelixFormer, to achieve the cross-image object semantic relation mining in a bidirectional and symmetrical manner. The HelixFormer consists of two steps: 1) Relation Mining Process (RMP) across different branches, and 2) Representation Enhancement Process (REP) within each individual branch. By the designed RMP, each branch can extract fine-grained object-level Cross-image Semantic Relation Maps (CSRMs) using information from the other branch, ensuring better cross-image interaction in semantically related local object regions. Further, with the aid of CSRMs, the developed REP can strengthen the extracted features for those discovered semantically-related local regions in each branch, boosting the model's ability to distinguish subtle feature differences of fine-grained objects. Extensive experiments conducted on five public fine-grained benchmarks demonstrate that HelixFormer can effectively enhance the cross-image object semantic relation matching for recognizing fine-grained objects, achieving much better performance over most state-of-the-art methods under 1-shot and 5-shot scenarios.
Bo Zhang 0069, Jiakang Yuan, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001, Botian Shi
ACM Multimedia3
2022 Stimulative Training of Residual Networks: A Social Psychology Perspective of Loafing
abstract
Residual networks have shown great success and become indispensable in today’s deep models. In this work, we aim to re-investigate the training process of residual networks from a novel social psychology perspective of loafing, and further propose a new training strategy to strengthen the performance of residual networks. As residual networks can be viewed as ensembles of relatively shallow networks (i.e., unraveled view) in prior works, we also start from such view and consider that the final performance of a residual network is co-determined by a group of sub-networks. Inspired by the social loafing problem of social psychology, we find that residual networks invariably suffer from similar problem, where sub-networks in a residual network are prone to exert less effort when working as part of the group compared to working alone. We define this previously overlooked problem as network loafing. As social loafing will ultimately cause the low individual productivity and the reduced overall performance, network loafing will also hinder the performance of a given residual network and its sub-networks. Referring to the solutions of social psychology, we propose stimulative training, which randomly samples a residual sub-network and calculates the KL-divergence loss between the sampled sub-network and the given residual network, to act as extra supervision for sub-networks and make the overall goal consistent. Comprehensive empirical results and theoretical analyses verify that stimulative training can well handle the loafing problem, and improve the performance of a residual network by improving the performance of its sub-networks. The code is available at https://github.com/Sunshine-Ye/NIPS22-ST.
Peng Ye 0006, Shengji Tang, Baopu Li, Tao Chen 0003, Wanli Ouyang
NeurIPS3
2022 Efficient Joint-Dimensional Search with Solution Space Regularization for Real-Time Semantic Segmentation
Peng Ye 0006, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001, Chen Lin 0003, Chongyan Zuo, Qinghua Chi, Wanli Ouyang
Int. J. Comput. Vis.2
2022 NOSnoop: An Effective Collaborative Meta-Learning Scheme Against Property Inference Attack
abstract
Collaborative learning has been used to train a joint model on geographically diverse data through periodically sharing knowledge. Although participants keep the data locally in collaborative learning, the adversary can still launch inference attacks through participants’ shared information. In this article, we focus on the property inference attack during model training and design a novel defense mechanism, namely, NOSnoop, to defend such an attack. We propose a collaborative meta-learning architecture to learn the common knowledge over all participants and utilize the natural advantage of meta-learning to hide the sensitive property data. We consider both irrelevant property and relevant property preservation in NOSnoop. For irrelevant property preservation, we utilize the inherent advantage of meta-learning to hide the sensitive property data in meta-training support data set. Thus, the adversary cannot capture the key information related to the sensitive properties and cannot infer victim’s private property successfully. For relevant property preservation, an adversarial game is further proposed to reduce the inference success rate of the adversary. We conduct comprehensive experiments to evaluate the effectiveness of NOSnoop. When hiding the sensitive property data in meta-training support data set, NOSnoop achieves an inference AUC score as low as 0.4984 for irrelevant property preservation, meaning the adversary cannot distinguish whether the training batch has the sensitive property data or not. When preserving the relevant property, NOSnoop is able to achieve an inference AUC score of 0.5091 without compromising model utility.
XinDi Ma, Baopu Li, Qi Jiang 0001, Yimin Chen 0004, Sheng Gao 0002, Jianfeng Ma 0001
IEEE Internet Things J.2
2022 Confidence Regularized Label Propagation Based Domain Adaptation
abstract
In domain adaptation (DA), label-induced losses generally occupy a dominant position and most previous models regard hard or soft labels as their inputs. However, these two types of labels may mislead the modeling process of label-induced losses since hard label is sensitive to a wrongly-predicted sample while soft label may introduce label noise, thus they may cause negative transfer. To relieve this problem, we propose a novel label learning approach namely confidence regularized label propagation (CRLP) that regularizes the confidence of predicted soft labels with constraints of F-norm or L21-norm. It is validated that maximizing either one of these two constraints equals to minimizing entropy loss. Specially, we illustrate that L21-norm is more suitable for DA than F-norm when the dataset contain a large number of categories. Then, we leverage the regularized soft labels produced by CRLP to reformulate some popular label-induced losses that consider feature transferability and discriminability such as class-wise maximum mean discrepancy, intra-class compactness and inter-class dispersion in a probability manner to present a novel DA method (i.e., CRLP-DA). Comprehensive analysis and experiments on four cross-domain object recognition datasets verify that the proposed CRLP-DA outperforms some state-of-the-art methods, especially 59.5% for Office10+Caltech10 dataset with SURF features. For others to better reproduce, our preliminary Matlab code will be available athttps://github.com/WWLoveTransfer/CRLP-DA/.
Wei Wang 0335, Baopu Li, Mengzhu Wang, Feiping Nie 0001, Zhihui Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Video Action Recognition with Neural Architecture Search
abstract
Recently, deep convolutional neural networks have been widely used in the field of videoaction recognition. Current approaches tend to concentrate on the structure design fordifferent backbone networks, but what kind of network structures can process video botheffectively and quickly still remains to be solved despite the encouraging progress. With thehelp of neural architecture search (NAS), we search for three hyperparameters in the videoprocessing network, which are the number of frames, the number of layers per residual stageand the channel number for all layers. We relax the entire search space into a continuoussearch space, and search for a set of network architectures that balance accuracy andcomputational efficiency by considering accuracy as the primary optimization goal andcomputational complexity as the secondary optimization goal. We conduct experiments onUCF101 and Kinetics400 datasets, validating new state-of-the-art results of the proposedNAS based scheme for video action recognition.
Yuanding Zhou, Baopu Li, Zhihui Wang 0001
ACML2
2021 MetaCorrection: Domain-Aware Meta Loss Correction for Unsupervised Domain Adaptation in Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) aims to transfer the knowledge from the labeled source domain to the unlabeled target domain. Existing self-training based UDA approaches assign pseudo labels for target data and treat them as ground truth labels to fully leverage unlabeled target data for model adaptation. However, the generated pseudo labels from the model optimized on the source domain inevitably contain noise due to the domain gap. To tackle this issue, we advance a MetaCorrection framework, where a Domain-aware Meta-learning strategy is devised to benefit Loss Correction (DMLC) for UDA semantic segmentation. In particular, we model the noise distribution of pseudo labels in target domain by introducing a noise transition matrix (NTM) and construct meta data set with domain-invariant source data to guide the estimation of NTM. Through the risk minimization on the meta data set, the optimized NTM thus can correct the noisy issues in pseudo labels and enhance the generalization ability of the model on the target data. Considering the capacity gap between shallow and deep features, we further employ the proposed DMLC strategy to provide matched and compatible supervision signals for different level features, thereby ensuring deep adaptation. Extensive experimental results highlight the effectiveness of our methodaagainst existing state-of-the-art methods on three benchmarks.
Xiaoqing Guo, Chen Yang 0026, Baopu Li, Yixuan Yuan
CVPR3
2021 Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-Resolution
abstract
Existing color-guided depth super-resolution (DSR) approaches require paired RGB-D data as training samples where the RGB image is used as structural guidance to recover the degraded depth map due to their geometrical similarity. However, the paired data may be limited or expensive to be collected in actual testing environment. Therefore, we explore for the first time to learn the cross-modality knowledge at training stage, where both RGB and depth modalities are available, but test on the target dataset, where only single depth modality exists. Our key idea is to distill the knowledge of scene structural guidance from RGB modality to the single DSR task without changing its network architecture. Specifically, we construct an auxiliary depth estimation (DE) task that takes an RGB image as input to estimate a depth map, and train both DSR task and DE task collaboratively to boost the performance of DSR. Upon this, a cross-task interaction module is proposed to realize bilateral cross-task knowledge transfer. First, we design a cross-task distillation scheme that encourages DSR and DE networks to learn from each other in a teacher-student role-exchanging fashion. Then, we advance a structure prediction (SP) task that provides extra structure regularization to help both DSR and DE networks learn more informative structure representations for depth recovery. Extensive experiments demonstrate that our scheme achieves superior performance in comparison with other DSR methods.
Baoli Sun, Xinchen Ye, Baopu Li, Zhihui Wang 0001, Rui Xu 0002
CVPR3
2021 Cost Affinity Learning Network for Stereo Matching
abstract
Existing stereo matching methods mainly tend to directly aggregate features output from Convolutional Neural Network to obtain more discriminative cost features, but ignore the affinity of each element in the cost feature which also plays a key role in enhancing the cost feature. In this work, we propose a novel cost affinity learning network(CAL-Net) whose Affinity Enhanced Module(AEM) extracts the affinity of the elements in the cost feature and reconstructs a more discriminative feature. In addition, CAL-Net designs a Disparity Weight Loss(DWL) to guide training. Specifically, AEM takes the advatange of the self-attention mechanism to learn internal affinity between different elements and exploits it to reconstruct the cost feature for emphasizing informative elements. DWL calculates the adaptive weight according to disparity error. As the error decreases, the weight gradually increases and enables the network to gradually transit from the pixel level disparity to sub-pixel level. Experiments demonstrate that CAL-Net boosts the performance, especially in textureless and reflective regions, and achieves better results on Scene Flow and KITTI 2012 benchmarks than some typical related methods.
Shenglun Chen, Baopu Li, Wei Wang 0335, Hong Zhang 0011, Zhihui Wang 0001
ICASSP2
2021 Real Image Super-Resolution Using Token Based Contextual Attention
Zhihong Pan 0001, Baopu Li
ICASSP2
2021 BN-NAS: Neural Architecture Search with Batch Normalization
abstract
We present BN-NAS, neural architecture search with Batch Normalization (BN-NAS), to accelerate neural architecture search (NAS). BN-NAS can significantly reduce the time required by model training and evaluation in NAS. Specifically, for fast evaluation, we propose a BN-based indicator for predicting subnet performance at a very early training stage. The BN-based indicator further facilitates us to improve the training efficiency by only training the BN parameters during the supernet training. This is based on our observation that training the whole supernet is not necessary while training only BN parameters accelerates network convergence for network architecture search. Extensive experiments show that our method can significantly shorten the time of training supernet by more than 10 times and shorten the time of evaluating subnets by more than 600,000 times without losing accuracy. The source codes are available at https://github.com/bychen515/BNNAS.
Peixia Li, Baopu Li, Chen Lin 0003, Chuming Li, Ming Sun 0008, Wanli Ouyang
ICCV3
2021 GLiT: Neural Architecture Search for Global and Local Image Transformer
abstract
We introduce the first Neural Architecture Search (NAS) method to find a better transformer architecture for image recognition. Recently, transformers without CNN-based backbones are found to achieve impressive performance for image recognition. However, the transformer is designed for NLP tasks and thus could be sub-optimal when directly used for image recognition. In order to improve the visual representation ability for transformers, we propose a new search space and searching algorithm. Specifically, we introduce a locality module that models the local correlations in images explicitly with fewer computational cost. With the locality module, our search space is defined to let the search algorithm freely trade off between global and local information as well as optimizing the low-level design choice in each module. To tackle the problem caused by huge search space, a hierarchical neural architecture search method is proposed to search the optimal vision transformer from two levels separately with the evolutionary algorithm. Extensive experiments on the ImageNet dataset demonstrate that our method can find more discriminative and efficient trans-former variants than the ResNet family (e.g., ResNet101) and the baseline ViT for image classification. The source codes are available at https://github.com/bychen515/GLiT.
Peixia Li, Chuming Li, Baopu Li, Lei Bai 0001, Chen Lin 0003, Ming Sun 0008, Wanli Ouyang
ICCV4
2021 AutoSampling: Search for Effective Data Sampling Schedules
abstract
Data sampling acts as a pivotal role in training deep learning models. However, an effective sampling schedule is difficult to learn due to its inherent high-dimension as a hyper-parameter. In this paper, we propose an AutoSampling method to automatically learn sampling schedules for model training, which consists of the multi-exploitation step aiming for optimal local sampling schedules and the exploration step for the ideal sampling distribution. More specifically, we achieve sampling schedule search with shortened exploitation cycle to provide enough supervision. In addition, we periodically estimate the sampling distribution from the learned sampling schedules and perturb it to search in the distribution space. The combination of two searches allows us to learn a robust sampling schedule. We apply our AutoSampling method to a variety of image classification tasks illustrating the effectiveness of the proposed method.
Ming Sun 0008, Haoxuan Dou, Baopu Li, Wanli Ouyang
ICML3
2021 No Need for Interactions: Robust Model-Based Imitation Learning using Neural ODE
abstract
Interactions with either environments or expert policies during training are needed for most of the current imitation learning (IL) algorithms. For IL problems with no interactions, a typical approach is Behavior Cloning (BC). However, BC-like methods tend to be affected by distribution shift. To mitigate this problem, we come up with a Robust Model-Based Imitation Learning (RMBIL) framework that casts imitation learning as an end-to-end differentiable nonlinear closed-loop tracking problem. RMBIL applies Neural ODE to learn a precise multi-step dynamics and a robust tracking controller via Nonlinear Dynamics Inversion (NDI) algorithm. Then, the learned NDI controller will be combined with a trajectory generator, a conditional VAE, to imitate an expert’s behavior. Theoretical derivation shows that the controller network can approximate an NDI when minimizing the training loss of Neural ODE. Experiments on Mujoco tasks also demonstrate that RMBIL is competitive to the state-of-the-art generative adversarial method (GAIL) and achieves at least 30% performance gain over BC in uneven surfaces.
HaoChih Lin, Baopu Li, Jiankun Wang 0001, Max Q.-H. Meng
ICRA2
2021 Cursor-based Adaptive Quantization for Deep Convolutional Neural Network
Baopu Li, Yanwen Fan, Zhihong Pan 0001, Zhiyu Cheng
IJCNN1
2021 COINet: Adaptive Segmentation with Co-Interactive Network for Autonomous Driving
abstract
Semantic segmentation serves as a cornerstone for safety autonomous driving and has been achieved remarkable progress at the price of dense annotations. Unsupervised domain adaptation was widely utilized to addresses this labor-intensive problem, which transfers the knowledge learned from labeled synthetic datset to real-world without any annotations. However, most existing adaptation works predict the segmentation results and domain identification results separately only with the last-layer feature, and ignore the intrinsic relationship among these two tasks. To address this issue, we present a CO-Interactive Network (COINet) for unsupervised adaptive segmentation. In particular, we propose a scale-aware distilled decoder to integrate multi-scale features dynamically through the designed inter-distilled module (IDM) and obtain fine-grained feature representations. A dual-task classifier is advanced with this decoder, to jointly predict the segmentation results and pixel-wise domain prediction results, which extracts shared complementary information for accurate segmentation. We further devise a co-interactive loss to explicitly model the intrinsic relationship among the segmentation and domain prediction, enabling the feature distribution alignment in pixel-level and an optimal segmentation decision boundary. We demonstrate the effectiveness of the proposed COINet on benchmark adaptation settings with extensive experimental and ablation results, and our model shows favorable performance against existing algorithms.
Jie Liu 0044, Xiaoqing Guo, Baopu Li, Yixuan Yuan
IROS3
2021 Automatic Channel Pruning with Hyper-parameter Search and Dynamic Masking
abstract
Modern deep neural network models tend to be large and computationally intensive. One typical solution to this issue is model pruning. However, most current model pruning algorithms depend on hand crafted rules or need to input the pruning ratio beforehand. To overcome this problem, we propose a learning based automatic channel pruning algorithm for deep neural network, which is inspired by recent automatic machine learning (Auto ML). A two objectives' pruning problem that aims for the weights and the remaining channels for each layer is first formulated. An alternative optimization approach is then proposed to derive the channel numbers and weights simultaneously. In the process of pruning, we utilize a searchable hyper-parameter, remaining ratio, to denote the number of channels in each convolution layer, and then a dynamic masking process is proposed to describe the corresponding channel evolution. To adjust the trade-off between accuracy of a model and the pruning ratio of floating point operations, a new loss function is further introduced. Extensive experimental results on benchmark datasets demonstrate that our scheme achieves competitive results for neural network pruning.
Baopu Li, Yanwen Fan, Zhihong Pan 0001, Yuchen Bian
ACM Multimedia1
2021 Exploring Gradient Flow Based Saliency for DNN Model Compression
abstract
Model pruning aims to reduce the deep neural network (DNN) model size or computational overhead. Traditional model pruning methods such as l-1 pruning that evaluates the channel significance for DNN pay too much attention to the local analysis of each channel and make use of the magnitude of the entire feature while ignoring its relevance to the batch normalization (BN) and ReLU layer after each convolutional operation. To overcome these problems, we propose a new model pruning method from a new perspective of gradient flow in this paper. Specifically, we first theoretically analyze the channel's influence based on Taylor expansion by integrating the effects of BN layer and ReLU activation function. Then, the incorporation of the first-order Talyor polynomial of the scaling parameter and the shifting parameter in the BN layer is suggested to effectively indicate the significance of a channel in a DNN. Comprehensive experiments on both image classification and image denoising tasks demonstrate the superiority of the proposed novel theory and scheme. Code is available at https://github.com/CityU-AIM-Group/GFBS.
Xinyu Liu 0001, Baopu Li, Zhen Chen 0013, Yixuan Yuan
ACM Multimedia2
2021 InterBN: Channel Fusion for Adversarial Unsupervised Domain Adaptation
abstract
A classifier trained on one dataset rarely works on other datasets obtained under different conditions because of domain shifting. Such a problem is usually solved by domain adaptation methods. In this paper, we propose a novel unsupervised domain adaptation (UDA) method based on Interchangeable Batch Normalization (InterBN) to fuse different channels in deep neural networks for adversarial domain adaptation.Specifically, we first observe that the channels with small batch normalization scaling factor have less influence on the whole domain adaption, followed by a theoretical proof that the scaling factors for some channels will definitely come close to zero when imposing a sparsity regularization. Then, we replace the channels that have smaller scaling factors in the source domain with the mean of the channels which have larger scaling factors in the target domain or vice versa. Such a simple but effective channel fusion scheme can drastically increase the domain adaption ability.Extensive experimental results show that our InterBN significantly outperforms the current adversarial domain adaptation methods by a large margin on four visual benchmarks. In particular, InterBN achieves a remarkable improvement of 7.7% over the conditional adversarial adaptation networks (CDAN) on VisDA-2017 benchmark.
Mengzhu Wang, Wei Wang 0335, Baopu Li, Xiang Zhang 0008, Long Lan, Huibin Tan, Tianyi Liang 0001, Wei Yu 0029, Zhigang Luo
ACM Multimedia3
2021 Kinematic Constrained Bi-directional RRT with Efficient Branch Pruning for robot path planning
Jiankun Wang 0001, Baopu Li, Max Q.-H. Meng
Expert Syst. Appl.2
2021 Sparsely-labeled source assisted domain adaptation
Wei Wang 0335, Shenglun Chen, Yuankai Xiang, Jing Sun 0012, Zhihui Wang 0001, Fuming Sun, Zhengming Ding, Baopu Li
Pattern Recognit.9
2021 Dual Color Space Guided Sketch Colorization
abstract
Automatic sketch colorization is a challenging task in both computer graphics and computer vision since all the color, texture, shading generation have to be created based on the abstract sketch. Besides, it is a subjective task in painting process, which needs illustrators to comprehend drawing priori (DP), such as hue variation, saturation contrast and gray contrast and utilize them in the HSV color space which is closer to human visual cognition system. As such, incorporating supplementary supervision in the HSV color space may be beneficial to sketch colorization. However, previous methods improve the colorization quality only in the RGB color space without considering the HSV color space, often causing results with dull color, inappropriate saturation contrast, and artifacts. To address this issue, we propose a novel sketch colorization method, dual color space guided generative adversarial network (DCSGAN), that considers the complementary information contained in both the RGB and HSV color space. Specifically, we incorporate the HSV color space to construct dual color spaces for supervising our method with a color space transformation (CST) network that learns transformation from the RGB to HSV color space. Then, we propose a DP loss that enables the DCSGAN to generate vivid color images with pixel level supervision. Additionally, a novel dual color space adversarial (DCSA) loss is designed to guide the generator at global level to reduce the artifacts to meet audiences' aesthetic expectations. Extensive experiments and ablation studies demonstrate the superiority of the proposed method over previous state-of-the-art (SOTA) methods.
Zhi Dou, Ning Wang 0025, Baopu Li, Zhihui Wang 0001, Bin Liu 0040
IEEE Trans. Image Process.3
2021 HTD: Heterogeneous Task Decoupling for Two-Stage Object Detection
abstract
Decoupling the sibling head has recently shown great potential in relieving the inherent task-misalignment problem in two-stage object detectors. However, existing works design similar structures for the classification and regression, ignoring task-specific characteristics and feature demands. Besides, the shared knowledge that may benefit the two branches is neglected, leading to potential excessive decoupling and semantic inconsistency. To address these two issues, we propose Heterogeneous task decoupling (HTD) framework for object detection, which utilizes a Progressive Graph (PGraph) module and a Border-aware Adaptation (BA) module for task-decoupling. Specifically, we first devise a Semantic Feature Aggregation (SFA) module to aggregate global semantics with image-level supervision, serving as the shared knowledge for the task-decoupled framework. Then, the PGraph module performs progressive graph reasoning, including local spatial aggregation and global semantic interaction, to enhance semantic representations of region proposals for classification. The proposed BA module integrates multi-level features adaptively, focusing on the low-level border activation to obtain representations with spatial and border perception for regression. Finally, we utilize the aggregated knowledge from SFA to keep the instance-level semantic consistency (ISC) of decoupled frameworks. Extensive experiments demonstrate that HTD outperforms existing detection works by a large margin, and achieves single-model 50.4%AP and 33.2% APs on COCO test-dev set using ResNet-101-DCN backbone, which is the best entry among state-of-the-arts under the same configuration. Our code is available at https://github.com/CityU-AIM-Group/HTD.
Wuyang Li, Zhen Chen 0013, Baopu Li, Dingwen Zhang, Yixuan Yuan
IEEE Trans. Image Process.3
2021 AutoPedestrian: An Automatic Data Augmentation and Loss Function Search Scheme for Pedestrian Detection
abstract
Pedestrian detection is a challenging and hot research topic in the field of computer vision, especially for the crowded scenes where occlusion happens frequently. In this paper, we propose a novel AutoPedestrian scheme that automatically augments the pedestrian data and searches for suitable loss functions, aiming for better performance of pedestrian detection especially in crowded scenes. To our best knowledge, it is the first work to automatically search the optimal policy of data augmentation and loss function jointly for the pedestrian detection. To achieve the goal of searching the optimal augmentation scheme and loss function jointly, we first formulate the data augmentation policy and loss function as probability distributions based on different hyper-parameters. Then, we apply a double-loop scheme with importance-sampling to solve the optimization problem of data augmentation and loss function types efficiently. Comprehensive experiments on two popular benchmarks of CrowdHuman and CityPersons show the effectiveness of our proposed method. In particular, we achieve 40.58% in MR on CrowdHuman datasets and 11.3% in MR on CityPersons reasonable subset, yielding new state-of-the-art results on these two datasets.
Baopu Li, Min Liu 0008, Yaonan Wang 0001, Wanli Ouyang
IEEE Trans. Image Process.2
2020 Action Segmentation With Joint Self-Supervised Temporal Domain Adaptation
abstract
Despite the recent progress of fully-supervised action segmentation techniques, the performance is still not fully satisfactory. One main challenge is the problem of spatiotemporal variations (e.g. different people may perform the same activity in various ways). Therefore, we exploit unlabeled videos to address this problem by reformulating the action segmentation task as a cross-domain problem with domain discrepancy caused by spatio-temporal variations. To reduce the discrepancy, we propose SelfSupervised Temporal Domain Adaptation (SSTDA), which contains two self-supervised auxiliary tasks (binary and sequential domain prediction) to jointly align cross-domain feature spaces embedded with local and global temporal dynamics, achieving better performance than other Domain Adaptation (DA) approaches. On three challenging benchmark datasets (GTEA, 50Salads, and Breakfast), SSTDA outperforms the current state-of-the-art method by large margins (e.g. for the F1@25 score, from 59.6% to 69.1% on Breakfast, from 73.4% to 81.5% on 50Salads, and from 83.6% to 89.1% on GTEA), and requires only 65% of the labeled training data for comparable performance, demonstrating the usefulness of adapting to unlabeled target videos across variations. The source code is available at https://github.com/cmhungsteve/SSTDA.
Min-Hung Chen, Baopu Li, Sid Ying-Ze Bao, Ghassan Al-Regib, Zsolt Kira
CVPR2
2020 Cross-Modality Person Re-Identification With Shared-Specific Feature Transfer
abstract
Cross-modality person re-identification (cm-ReID) is a challenging but key technology for intelligent video analysis. Existing works mainly focus on learning modality-shared representation by embedding different modalities into a same feature space, lowering the upper bound of feature distinctiveness. In this paper, we tackle the above limitation by proposing a novel cross-modality shared-specific feature transfer algorithm (termed cm-SSFT) to explore the potential of both the modality-shared information and the modality-specific characteristics to boost the reidentification performance. We model the affinities of different modality samples according to the shared features and then transfer both shared and specific features among and across modalities. We also propose a complementary feature learning strategy including modality adaption, project adversarial learning and reconstruction enhancement to learn discriminative and complementary shared and specific features of each modality, respectively. The entire cmSSFTalgorithm can be trained in an end-to-end manner. We conducted comprehensive experiments to validate the superiority ofthe overall algorithm and the effectiveness ofeach component. The proposed algorithm significantly outperforms state-of-the-arts by 22.5% and 19.3% mAP on the two mainstream benchmark datasets SYSU-MM01 and RegDB, respectively.
Yan Lu 0001, Bin Liu 0016, Tianzhu Zhang 0001, Baopu Li, Qi Chu 0001, Nenghai Yu
CVPR5
2020 A Novel Rank Selection Scheme in Tensor Ring Decomposition Based on Reinforcement Learning for Deep Neural Networks
abstract
Tensor decomposition has been proved to be effective for solving many problems in signal processing and machine learning[1]. Recently, tensor decomposition finds its advantage for compressing deep neural networks. In many applications of deep neural networks, it is critical to reduce the number of parameters and computation workload to accelerate inference speed in deployment of the network. Modern deep neural network consists of multiple layers with multi-array weights where tensor decomposition is a natural way to perform compression. It is achieved by decomposing the weight tensors in convolutional layers or fully-connected layers with specified tensor ranks (e.g. canonical ranks, tensor train ranks). Conventional tensor decomposition in compressing deep neural networks selects the ranks manually that requires tedious human efforts to finetune the performance. To overcome this issue, we propose a novel rank selection scheme, which is inspired by reinforcement learning, to automatically select ranks in recently studied tensor ring decomposition in each convolutional layer. Experimental results validate that our learning based rank selection significantly outperforms hand-crafted rank selection heuristics on a number of benchmark datasets, for the purpose of effectively compressing deep neural networks while maintaining comparable accuracy.
Zhiyu Cheng, Baopu Li, Yanwen Fan, Sid Ying-Ze Bao
ICASSP2
2020 Deep Residual Network for MSFA Raw Image Denoising
abstract
Multispectral filter arrays (MSFA) is increasingly used in multispectral imaging. While many previous works studied the denoising algorithms for CFA based cameras, denoising MSFA raw images is little discussed. The major challenges for denoising MSFA data include 1) more channels than CFA and no predominant channel; 2) compatibility between denoising and the subsequent demosaicking process. To overcome these challenges, we propose a new deep residual network designed to account for the uniqueness of MSFA mosaic patterns. First, a split and stride convolution layer is innovated to match the mosaic pattern of the MSFA raw image. Then, data augmentation using MSFA shifting and dynamic noise is proposed to make the model robust to different noise levels. In addition, a new network optimization criteria is also suggested by using the noise standard deviation to normalize the L1 loss function. Comprehensive experiments demonstrate that the proposed deep residual network outperforms the state-of-the-art denoising algorithms in MSFA field.
Zhihong Pan 0001, Baopu Li, Hsuchun Cheng, Sid Ying-Ze Bao
ICASSP2
2020 Depth Super-Resolution via Deep Controllable Slicing Network
abstract
Due to the imaging limitation of depth sensors, high-resolution (HR) depth maps are often difficult to be acquired directly, thus effective depth super-resolution (DSR) algorithms are needed to generate HR output from its low-resolution (LR) counterpart. Previous methods treat all depth regions equally without considering different extents of degradation at region-level, and regard DSR under different scales as independent tasks without considering the modeling of different scales, which impede further performance improvement and practical use of DSR. To alleviate these problems, we propose a deep controllable slicing network from a novel perspective. Specifically, our model is to learn a set of slicing branches in a divide-and-conquer manner, parameterized by a distance-aware weighting scheme to adaptively aggregate different depths in an ensemble. Each branch that specifies a depth slice (e.g., the region in some depth range) tends to yield accurate depth recovery. Meanwhile, a scale-controllable module that extracts depth features under different scales is proposed and inserted into the front of slicing network, and enables finely-grained control of the depth restoration results of slicing network with a scale hyper-parameter. Extensive experiments on synthetic and real-world benchmark datasets demonstrate that our method achieves superior performance.
Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Jing-Yu Yang 0002, Rui Xu 0002, Baopu Li
ACM Multimedia7
2020 Action Segmentation with Mixed Temporal Domain Adaptation
abstract
The main progress for action segmentation comes from densely-annotated data for fully-supervised learning. Since manual annotation for frame-level actions is time-consuming and challenging, we propose to exploit auxiliary unlabeled videos, which are much easier to obtain, by shaping this problem as a domain adaptation (DA) problem. Although various DA techniques have been proposed in recent years, most of them have been developed only for the spatial direction. Therefore, we propose Mixed Temporal Domain Adaptation (MTDA) to jointly align frame-and video-level embedded feature spaces across domains, and further integrate with the domain attention mechanism to focus on aligning the frame-level features with higher domain discrepancy, leading to more effective domain adaptation. Finally, we evaluate our proposed methods on three challenging datasets (GTEA, 50Salads, and Breakfast), and validate that MTDA outperforms the current state-of-the-art methods on all three datasets by large margins (e.g. 6.4% gain on F1@50 and 6.8% gain on the edit score for GTEA).
Min-Hung Chen, Baopu Li, Sid Ying-Ze Bao, Ghassan Al-Regib
WACV2
2020 PMBANet: Progressive Multi-Branch Aggregation Network for Scene Depth Super-Resolution
abstract
Depth map super-resolution is an ill-posed inverse problem with many challenges. First, depth boundaries are generally hard to reconstruct particularly at large magnification factors. Second, depth regions on fine structures and tiny objects in the scene are destroyed seriously by downsampling degradation. To tackle these difficulties, we propose a progressive multi-branch aggregation network (PMBANet), which consists of stacked MBA blocks to fully address the above problems and progressively recover the degraded depth map. Specifically, each MBA block has multiple parallel branches: 1) The reconstruction branch is proposed based on the designed attention-based error feed-forward/-back modules, which iteratively exploits and compensates the downsampling errors to refine the depth map by imposing the attention mechanism on the module to gradually highlight the informative features at depth boundaries. 2) We formulate a separate guidance branch as prior knowledge to help to recover the depth details, in which the multi-scale branch is to learn a multi-scale representation that pays close attention at objects of different scales, while the color branch regularizes the depth map by using auxiliary color information. Then, a fusion block is introduced to adaptively fuse and select the discriminative features from all the branches. The design methodology of our whole network is well-founded, and extensive experiments on benchmark datasets demonstrate that our method achieves superior performance in comparison with the state-of-the-art methods. Our code and models are available athttps://github.com/Sunbaoli/PMBANet_DSR/.
Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Jing-Yu Yang 0002, Rui Xu 0002, Baopu Li
IEEE Trans. Image Process.7
2017 WCE Abnormality Detection Based on Saliency and Adaptive Locality-Constrained Linear Coding
abstract
Wireless capsule endoscopy (WCE) has become a widely used diagnostic technique for the digestive tract, at the price of a large volume of data that needs to be analyzed. To tackle this problem, a new computer-aided system using novel features is proposed in this paper to classify WCE images automatically. In the feature learning stage, to obtain the representative visual words, we first calculate the color scale invariant feature transform from the bleeding, polyp, ulcer, and normal WCE image samples separately and then apply K -means clustering on these features to obtain visual words. These four types of visual words are combined together to composite the representative visual words for classifying the WCE images. In the feature coding stage, we propose a novel saliency and adaptive locality-constrained linear coding (SALLC) algorithm to encode the images. The SALLC encodes patch features based on adaptive coding bases, which are calculated by the distance differences among the features and the visual words. Moreover, it imposes the patch saliency constraint on the feature coding process to emphasize the important information in the images. The experimental results exhibit a promising overall recognition accuracy of 88.61%, validating the effectiveness of the proposed method.
Yixuan Yuan, Baopu Li, Max Q.-H. Meng
IEEE Trans Autom. Sci. Eng.2
2016 Improved Bag of Feature for Automatic Polyp Detection in Wireless Capsule Endoscopy Images
abstract
Wireless capsule endoscopy (WCE) needs computerized method to reduce the review time for its large image data. In this paper, we propose an improved bag of feature (BoF) method to assist classification of polyps in WCE images. Instead of utilizing a single scale-invariant feature transform (SIFT) feature in the traditional BoF method, we extract different textural features from the neighborhoods of the key points and integrate them together as synthetic descriptors to carry out classification tasks. Specifically, we study influence of the number of visual words, the patch size and different classification methods in terms of classification performance. Comprehensive experimental results reveal that the best classification performance is obtained with the integrated feature strategy using the SIFT and the complete local binary pattern (CLBP) feature, the visual words with a length of 120, the patch size of 8*8, and the support vector machine (SVM). The achieved classification accuracy reaches 93.2%, confirming that the proposed scheme is promising for classification of polyps in WCE images.
Yixuan Yuan, Baopu Li, Max Q.-H. Meng
IEEE Trans Autom. Sci. Eng.2
2016 Bleeding Frame and Region Detection in the Wireless Capsule Endoscopy Video
abstract
Wireless capsule endoscopy (WCE) enables noninvasive and painless direct visual inspection of a patient's whole digestive tract, but at the price of long time reviewing large amount of images by clinicians. Thus, an automatic computer-aided technique to reduce the burden of physicians is highly demanded. In this paper, we propose a novel color feature extraction method to discriminate the bleeding frames from the normal ones, with further localization of the bleeding regions. Our proposal is based on a twofold system. First, we make full use of the color information of WCE images and utilize K-means clustering method on the pixel represented images to obtain the cluster centers, with which we characterize WCE images as words-based color histograms. Then, we judge the status of a WCE frame by applying the support vector machine (SVM) and K-nearest neighbor methods. Comprehensive experimental results reveal that the best classification performance is obtained with YCbCr color space, cluster number 80 and the SVM. The achieved classification performance reaches 95.75% in accuracy, 0.9771 for AUC, validating that the proposed scheme provides an exciting performance for bleeding classification. Second, we propose a two-stage saliency map extraction method to highlight bleeding regions, where the first-stage saliency map is created by means of different color channels mixer and the second-stage saliency map is obtained from the visual contrast. Followed by an appropriate fusion strategy and threshold, we localize the bleeding areas. Quantitative as well as qualitative results show that our methods could differentiate the bleeding areas from neighborhoods correctly.
Yixuan Yuan, Baopu Li, Max Q.-H. Meng
IEEE J. Biomed. Health Informatics2
2015 A novel license plate location method based on wavelet transform and EMD analysis
Shouyuan Yu, Baopu Li, Qi Zhang 0049, Max Q.-H. Meng
Pattern Recognit.2
2015 Saliency Based Ulcer Detection for Wireless Capsule Endoscopy Diagnosis
abstract
Ulcer is one of the most common symptoms of many serious diseases in the human digestive tract. Especially for the ulcers in the small bowel where other procedures cannot adequately visualize, wireless capsule endoscopy (WCE) is increasingly being used in the diagnosis and clinical management. Because WCE generates large amount of images from the whole process of inspection, computer-aided detection of ulcer is considered an indispensable relief to clinicians. In this paper, a two-staged fully automated computer-aided detection system is proposed to detect ulcer from WCE images. In the first stage, we propose an effective saliency detection method based on multi-level superpixel representation to outline the ulcer candidates. To find the perceptually and semantically meaningful salient regions, we first segment the image into multi-level superpixel segmentations. Each level corresponds to different initial region sizes of the superpixels. Then we evaluate the corresponding saliency according to the color and texture features in superpixel region of each level. In the end, we fuse the saliency maps from all levels together to obtain the final saliency map. In the second stage, we apply the obtained saliency map to better encode the image features for the ulcer image recognition tasks. Because the ulcer mainly corresponds to the saliency region, we propose a saliency max-pooling method integrated with the Locality-constrained Linear Coding (LLC) method to characterize the images. Experiment results achieve promising 92.65% accuracy and 94.12% sensitivity, validating the effectiveness of the proposed method. Moreover, the comparison results show that our detection system outperforms the state-of-the-art methods on the ulcer classification task.
Yixuan Yuan, Jiaole Wang, Baopu Li, Max Q.-H. Meng
IEEE Trans. Medical Imaging3
2014 Robust Object Tracking With Reacquisition Ability Using Online Learned Detector
abstract
Long term tracking is a challenging task for many applications. In this paper, we propose a novel tracking approach that can adapt various appearance changes such as illumination, motion, and occlusions, and owns the ability of robust reacquisition after drifting. We utilize a condensation-based method with an online support vector machine as a reliable observation model to realize adaptive tracking. To redetect the target when drifting, a cascade detector based on random ferns is proposed. It can detect the target robustly in real time. After redetection, we also come up with a new refinement strategy to improve the tracker's performance by removing the support vectors corresponding to possible wrong updates by a matching template. Extensive comparison experiments on typical and challenging benchmark dataset illustrate a robust and encouraging performance of the proposed approach.
Baopu Li, Max Q.-H. Meng
IEEE Trans. Cybern.2
2013 Adaptive visual tracking with reacquisition ability for arbitrary objects
abstract
This paper introduces a novel tracking framework for robots that can adapt various appearance changes of object and also owns the ability of reacquisition after drift. Two classifiers, LaRank and Online Random Ferns, are adopted to realize this tracking algorithm. The former one maintains the adaptive tracking using a Condensation-based method with an online support vector machine (SVM) as observation model, which also provides the reliable image patch samples to detector for updating. The other one is in charge of the task of detection in order to redetect the object when the target drifts. We also present a refinement strategy to improve the tracker's performance by discarding the support vector corresponding to possible wrong updates by a matching template after re-initialization. The experiments on benchmark dataset compare our tracking method with several other state-of-the-art algorithms, demonstrating a promising performance of the proposed framework.
Baopu Li, Max Q.-H. Meng
ICRA2
2013 A novel method for capsule endoscopy video automatic segmentation
abstract
Wireless capsule endoscopy (WCE) is a recently developed revolutionary medical technology which records the video of human's digestive tract noninvasively. However, reviewing a WCE video is a tired and time-consuming task for clinicians. Thus, WCE video automatic segmentation methods are emerging to reduce the review time for clinicians. In our previous work, a two-level WCE video segmentation approach has been proposed, which provides a novel approach to localize the boundaries more exactly and efficiently. However, it has an unsatisfactory performance in the small intestine/large intestine boundary detection. In this paper, we propose new features and an improved classifier to improve the previous two-level segmentation algorithm. In the rough level, color feature is utilized to draw a dissimilarity curve and an approximate boundary has been obtained. At the same time, training data for fine level can be directly labeled and collected between the two approximate boundaries of organs to overcome the difficulty of training data acquisition. In the fine level, a novel color uniform local binary pattern (CULBP) algorithm is proposed, which includes two kinds of patterns, color norm patterns and color angle patterns. The CULBP feature is more robust to variation of illumination and more discriminative for classification. Moreover, in order to elevate the performance of SVM classifier we proposed the Ada-SVM classifier which using RBFSVMs as component of Adaboost classifier. At last, an analysis of classification results of the Ada-SVM classifier is carried out to segment the WCE video into several meaningful parts, stomach, small intestine and large intestine. The experiments demonstrate a promising performance of the proposed method. The average precision and recall are as high as 91.37% and 88.50% in stomach/small intestine classification, 90.35% and 97.28% in small intestine/ large intestine classification.
Ran Zhou 0002, Baopu Li, Hongmei Zhu, Max Q.-H. Meng
IROS2
2012 A novel correspondence searching strategy in multiocular vision
abstract
Correspondence searching among different images is a fundamental problem in computer vision. It is important to find correspondences correctly and rapidly, especially for real-time tracking systems. Therefore, the definition of search areas in images is crucial. Traditional epipolar constraint is not noise-enduring; some reformative methods lack explicit geometric meanings. All of them cannot help defining rational search areas under noises. This paper proposes two new binocular imaging constraints with clear geometric meanings and strong restraining forces. Based on them, a novel searching strategy among multiimages is developed which can define optimal search areas with smallest sizes but best reliability. Practical algorithms for implementation are presented and experiments with real images are performed, validating the effectiveness of the proposed strategy.
Baopu Li, Max Q.-H. Meng
ICRA2
2012 Automatic polyp detection for wireless capsule endoscopy images
Baopu Li, Max Q.-H. Meng
Expert Syst. Appl.1
2012 Wireless capsule endoscopy images enhancement via adaptive contrast diffusion
Baopu Li, Max Q.-H. Meng
J. Vis. Commun. Image Represent.1
2012 Tumor Recognition in Wireless Capsule Endoscopy Images Using Textural Features and SVM-Based Feature Selection
abstract
Tumor in digestive tract is a common disease and wireless capsule endoscopy (WCE) is a relatively new technology to examine diseases for digestive tract especially for small intestine. This paper addresses the problem of automatic recognition of tumor for WCE images. Candidate color texture feature that integrates uniform local binary pattern and wavelet is proposed to characterize WCE images. The proposed features are invariant to illumination change and describe multiresolution characteristics of WCE images. Two feature selection approaches based on support vector machine, sequential forward floating selection and recursive feature elimination, are further employed to refine the proposed features for improving the detection accuracy. Extensive experiments validate that the proposed computer-aided diagnosis system achieves a promising tumor recognition accuracy of 92.4% in WCE images on our collected data.
Baopu Li, Max Q.-H. Meng
IEEE Trans. Inf. Technol. Biomed.1
2011 Comparison of several image features for WCE video abstract
abstract
The direct view of the inner tract of the small intestine is not feasible until a recently revolutionary imaging technology, wireless capsule endoscopy (WCE), appeared in 2001. However, interpretation of the produced video data for the digestive tract on each patient is left to naked eyes of medical staffs. Such a process is very tedious and time-consuming with the average inspection time about two hours for a whole WCE video. To overcome this big problem, automatic WCE video analysis is required. In this paper, we propose a comparative study of several image based features that may be suitable for WCE video abstract, which may be a good candidate to reduce the burden of physicians. Color, texture and motion features built from images are investigated and compared to show their performance in representing video content for a WCE video abstract. Preliminary experimental results of these features for WCE video abstract are also demonstrated and discussed. It is found that textural and motion features may be suitable candidates for visual frame depiction for WCE video abstract in terms of visual content representation and compression ratio. Clinical validation of our work remains to be implemented in the near future.
Baopu Li, Max Q.-H. Meng
IROS1
2011 Computer-aided small bowel tumor detection for capsule endoscopy
Baopu Li, Max Q.-H. Meng, James Yun Wong Lau
Artif. Intell. Medicine1
2010 Tumor CE image classification using SVM-based feature selection
abstract
In this paper, we propose a new scheme aimed for gastrointestinal (GI) tumor capsule endoscopy (CE) images classification, which utilizes sequential forward floating selection (SFFS) together with support vector machine (SVM). To achieve this goal, candidate features related to texture characteristics of CE images are extracted. With these candidate features, SFFS based on SVM is applied to select the most discriminative features that can separate normal CE images from tumor CE images. Comprehensive experiments on our present CE image data verify that it is promising to employ the proposed scheme to recognize tumor CE images.
Baopu Li, Max Q.-H. Meng
IROS1
2009 Small bowel tumor detection for wireless capsule endoscopy images using textural features and support vector machine
abstract
Wireless capsule endoscopy (WCE) has been gradually applied in hospitals due to its great advantage that it can directly view the entire small bowel in human body compared with traditional endoscopies and other imaging techniques for gastrointestinal diseases. However, a challenging problem with this new technology is that too many images produced by WCE causes a tough task to doctors, so it is very significant to help and relief the clinicians if we can develop computer based automatic detection system to prescreen the collected large amount of images and identify the images with potential problems. In this paper, we propose a new scheme aimed for small bowel tumor detection of WCE images. This new scheme utilizes texture feature, also a powerful clue used by physicians, to detect tumor images with support vector machine. We put forward a new idea of wavelet based local binary pattern as the textural features to discriminate tumor regions from normal regions, which take advantage of wavelet transform and uniform local binary pattern. With support vector machine as the classifier, three-fold cross validation experiments on our present image data verify that it is promising to employ the proposed texture features to recognize the small bowel tumor regions.
Baopu Li, Max Q.-H. Meng
IROS1
2009 Texture analysis for ulcer detection in capsule endoscopy images
Baopu Li, Max Q.-H. Meng
Image Vis. Comput.1
2007 Wireless Capsule Endoscopy Images Enhancement using Contrast Driven Forward and Backward Anisotropic Diffusion
abstract
The wireless capsule endoscopy (WCE) has been widely used to detect the diseases in gastrointestinal tract. However, the contrast of many images it produced is rather dark due to some reasons, which causes some difficulties to diagnosis and computer aided diagnosis. To overcome this shortcoming, we propose a forward and backward anisotropic diffusion method based on the contrast space to enhance the capsule endoscopy images. Experimental results show that this new method can provide a better visualization of the wireless capsule endoscopy images than the forward and backward Beltrami flow and the contrast-limited adaptive histogram equalization so as to assist the inspection of WCE images.
Baopu Li, Max Q.-H. Meng
ICIP (2)1