VLDB 2026 Research / reviewers in the wild / expert
Bin Li 0038
dblp:89/6764-38
· DBLP profile ↗
12ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 5%>100%: Breaking Performance Shackles of Full Fine-Tuning on Visual Recognition TasksabstractPre-training & fine-tuning can enhance the transferring efficiency and performance in visual tasks. Recent delta-tuning methods provide more options for visual classification tasks. Despite their success, existing visual delta-tuning art fails to exceed the upper limit of full fine-tuning on challenging tasks. To find a competitive alternative to full fine-tuning, we propose the Multi-cognitive Visual Adapter (Mona) tuning, a novel adapter-based tuning method. First, we introduce multiple vision-friendly filters into the adapter to enhance its ability for processing visual signals, while previous methods mainly rely on language-friendly linear filters. Second, we add the scaled layer-norm in the adapter to regulate the distribution of input features for visual filters. To fully demonstrate the practicality and generality of Mona, we conduct experiments on representative visual tasks, including instance segmentation on COCO, semantic segmentation on ADE20K, object detection on Pascal VOC, oriented object detection on DOTA/STAR, and image classification on three common datasets. Exciting results illustrate that Mona surpasses full fine-tuning on all these tasks by tuning less than 5% params of the backbone, and is the only delta-tuning method outperforming full fine-tuning on all tasks. For example, Mona achieves 1% performance gain on the COCO compared to full fine-tuning. Comprehensive results suggest that Mona-tuning is more suitable for retaining and utilizing the capabilities of pre-trained models than full fine-tuning. The code is publicly available on https://github.com/Leiyi-Hu/mona. Dongshuo Yin, Leiyi Hu, Bin Li 0038, Youqun Zhang, Xue Yang 0005 |
CVPR | 3 |
| 2024 | Parameter-efficient is not Sufficient: Exploring Parameter, Memory, and Time Efficient Adapter Tuning for Dense PredictionsabstractPre-training & fine-tuning is a prevalent paradigm in computer vision (CV). Recently, parameter-efficient transfer learning (PETL) methods have shown promising performance in adapting to downstream tasks with only a few trainable parameters. Despite their success, the existing PETL methods in CV can be computationally expensive and require large amounts of memory and time cost during training, which limits low-resource users from conducting research and applications on large models. In this work, we propose Parameter, Memory, and Time Efficient Visual Adapter (E3VA) tuning to address this issue. We provide a gradient backpropagation highway for low-rank adapters which eliminates the need for expensive backpropagation through the frozen pre-trained model, resulting in substantial savings of training memory and training time. Furthermore, we optimise the E3VA structure for CV tasks to promote model performance. Extensive experiments on COCO, ADE20K, and Pascal VOC benchmarks show that E3VA can save up to 62.2% training memory and 26.2% training time on average, while achieving comparable performance to full fine-tuning and better performance than most PETL methods. Note that we can even train the Swin-Large-based Cascade Mask RCNN on GTX 1080Ti GPUs with less than 1.5% trainable parameters. Dongshuo Yin, Xueting Han, Bin Li 0038, Jing Bai 0010 |
ACM Multimedia | 3 |
| 2024 | Epoch-Evolving Gaussian Process Guided Learning for ClassificationabstractThe conventional mini-batch gradient descent algorithms are usually trapped in the local batch-level distribution information, resulting in the ``zig-zag'' effect in the learning process. To characterize the correlation information between the batch-level distribution and the global data distribution, we propose a novel learning scheme called epoch-evolving Gaussian process guided learning (GPGL) to encode the global data distribution information in a non-parametric way. Upon a set of class-aware anchor samples, our GP model is built to estimate the class distribution for each sample in mini-batch through label propagation from the anchor samples to the batch samples. The class distribution, also named the context label, is provided as a complement for the ground-truth one-hot label. Such a class distribution structure has a smooth property and usually carries a rich body of contextual information that is capable of speeding up the convergence process. With the guidance of the context label and ground-truth label, the GPGL scheme provides a more efficient optimization through updating the model parameters with a triangle consistency loss. Furthermore, our GPGL scheme can be generalized and naturally applied to the current deep models, outperforming the state-of-the-art optimization methods on six benchmark datasets. Jiabao Cui, Xuewei Li 0003, Hanbin Zhao, Hui Wang 0107, Bin Li 0038, Xi Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Forgery face detection via adaptive learning from multiple experts
Xinghe Fu, Shengming Li, Yike Yuan, Bin Li 0038, Xi Li 0001 |
Neurocomputing | 4 |
| 2023 | Knowledge Distillation Classifier Generation Network for Zero-Shot LearningabstractIn this article, we present a conceptually simple but effective framework called knowledge distillation classifier generation network (KDCGN) for zero-shot learning (ZSL), where the learning agent requires recognizing unseen classes that have no visual data for training. Different from the existing generative approaches that synthesize visual features for unseen classifiers' learning, the proposed framework directly generates classifiers for unseen classes conditioned on the corresponding class-level semantics. To ensure the generated classifiers to be discriminative to the visual features, we borrow the knowledge distillation idea to both supervise the classifier generation and distill the knowledge with, respectively, the visual classifiers and soft targets trained from a traditional classification network. Under this framework, we develop two, respectively, strategies, i.e., class augmentation and semantics guidance, to facilitate the supervision process from the perspectives of improving visual classifiers. Specifically, the class augmentation strategy incorporates some additional categories to train the visual classifiers, which regularizes the visual classifier weights to be compact, under supervision of which the generated classifiers will be more discriminative. The semantics-guidance strategy encodes the class semantics into the visual classifiers, which would facilitate the supervision process by minimizing the differences between the generated and the real-visual classifiers. To evaluate the effectiveness of the proposed framework, we have conducted extensive experiments on five datasets in image classification, i.e., AwA1, AwA2, CUB, FLO, and APY. Experimental results show that the proposed approach performs best in the traditional ZSL task and achieves a significant performance improvement on four out of the five datasets in the generalized ZSL task. Yunlong Yu 0001, Bin Li 0038, Zhong Ji, Jungong Han, Zhongfei Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Self-Paced Knowledge Distillation for Real-Time Image Guided Depth CompletionabstractImage guided depth completion aims to generate a dense depth map from a sparse one with the guidance of a color image. Previous high-accuracy methods often rely on complex networks that are large in size and expensive in computational cost, making them inapplicable to real-time platforms. In this letter, we propose a self-paced knowledge distillation method, which obtains a lightweight but accurate depth completion model via distilling knowledge from a complex teacher network. Specifically, by taking advantage of the easy-to-hard learning curriculum in deep networks, we first design a groundtruth-free hard-pixel mining module to tell hard and noisy pixels in the teacher’s output. Then, we design two self-paced distillation losses, which gradually introduce hard pixels to distill the depth and structure knowledge from the teacher to the compact student network. Experiments on the KITTI benchmark show that the proposed method can improve the original student model by a considerable margin. The distilled compact and real-time student model outperforms all previous lightweight networks, mitigating the performance gap with state-of-the-art high-accuracy but complex models. Shuling Wang 0002, Mu Hu, Bin Li 0038, Xiaojin Gong |
IEEE Signal Process. Lett. | 3 |
| 2021 | PENet: Towards Precise and Efficient Image Guided Depth CompletionabstractImage guided depth completion is the task of generating a dense depth map from a sparse depth map and a high quality image. In this task, how to fuse the color and depth modalities plays an important role in achieving good performance. This paper proposes a two-branch backbone that consists of a color-dominant branch and a depth-dominant branch to exploit and fuse two modalities thoroughly. More specifically, one branch inputs a color image and a sparse depth map to predict a dense depth map. The other branch takes as inputs the sparse depth map and the previously predicted depth map, and outputs a dense depth map as well. The depth maps predicted from two branches are complimentary to each other and therefore they are adaptively fused. In addition, we also propose a simple geometric convolutional layer to encode 3D geometric cues. The geometric encoded backbone conducts the fusion of different modalities at multiple stages, leading to good depth completion results. We further implement a dilated and accelerated CSPN++ to refine the fused depth map efficiently. The proposed full model ranks 1st in the KITTI depth completion online leaderboard at the time of submission. It also infers much faster than most of the top ranked methods. The code of this work is available at https://github.com/JUGGHM/PENet_ICRA2021. Mu Hu, Shuling Wang 0002, Bin Li 0038, Shiyu Ning, Xiaojin Gong |
ICRA | 3 |
| 2021 | Self-supervised Visual-LiDAR Odometry with Flip ConsistencyabstractMost learning-based methods estimate ego-motion by utilizing visual sensors, which suffer from dramatic lighting variations and textureless scenarios. In this paper, we incorporate sparse but accurate depth measurements obtained from lidars to overcome the limitation of visual methods. To this end, we design a self-supervised visual-lidar odometry (Self-VLO) framework. It takes both monocular images and sparse depth maps projected from 3D lidar points as input, and produces pose and depth estimations in an end-to-end learning manner, without using any ground truth labels. To effectively fuse two modalities, we design a two-pathway encoder to extract features from visual and depth images and fuse the encoded features with those in decoders at multiple scales by our fusion module. We also adopt a siamese architecture and design an adaptively weighted flip consistency loss to facilitate the self-supervised learning of our VLO. Experiments on the KITTI odometry benchmark show that the proposed approach out-performs all self-supervised visual or lidar odometries. It also performs better than fully supervised VOs, demonstrating the power of fusion. Bin Li 0038, Mu Hu, Shuling Wang 0002, Lianghao Wang, Xiaojin Gong |
WACV | 1 |
| 2021 | Multitask Non-Autoregressive Model for Human Motion PredictionabstractHuman motion prediction, which aims at predicting future human skeletons given the past ones, is a typical sequence-to-sequence problem. Therefore, extensive efforts have been devoted to exploring different RNN-based encoder-decoder architectures. However, by generating target poses conditioned on the previously generated ones, these models are prone to bringing issues such as error accumulation problem. In this paper, we argue that such issue is mainly caused by adopting autoregressive manner. Hence, a novel Non-AuToregressive model (NAT) is proposed with a complete non-autoregressive decoding scheme, as well as a context encoder and a positional encoding module. More specifically, the context encoder embeds the given poses from temporal and spatial perspectives. The frame decoder is responsible for predicting each future pose independently. The positional encoding module injects positional signal into the model to indicate the temporal order. Besides, a multitask training paradigm is presented for both low-level human skeleton prediction and high-level human action recognition, resulting in the considerable improvement for the prediction task. Our approach is evaluated on Human3.6M and CMU-Mocap benchmarks and outperforms state-of-the-art autoregressive methods. Bin Li 0038, Zhongfei Zhang, Hailin Feng, Xi Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Condition-Aware Comparison Scheme for Gait RecognitionabstractAs an important and challenging problem, gait recognition has gained considerable attention. It suffers from confounding conditions, that is, it is sensitive to camera views, dressing types and so on. Interestingly, it is observed that, under different conditions, local body parts contribute differently to recognition performance. In this paper, we propose a condition-aware comparison scheme to measure gait pairs' similarity via a novel module named Instructor. Also, we present a geometry-guided data augmentation approach (Dresser) to enrich dressing conditions. Furthermore, to enhance the gait representation, we propose to model temporal local information from coarse to fine. Our model is evaluated on two popular benchmarks, CASIA-B and OULP. Results show that our method outperforms current state-of-the-art methods, especially in the cross-condition scenario. Haoqian Wu, Yongjian Fu 0002, Bin Li 0038, Xi Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Spatio-Temporal Graph Routing for Skeleton-Based Action RecognitionabstractWith the representation effectiveness, skeleton-based human action recognition has received considerable research attention, and has a wide range of real applications. In this area, many existing methods typically rely on fixed physicalconnectivity skeleton structure for recognition, which is incapable of well capturing the intrinsic high-order correlations among skeleton joints. In this paper, we propose a novel spatio-temporal graph routing (STGR) scheme for skeletonbased action recognition, which adaptively learns the intrinsic high-order connectivity relationships for physicallyapart skeleton joints. Specifically, the scheme is composed of two components: spatial graph router (SGR) and temporal graph router (TGR). The SGR aims to discover the connectivity relationships among the joints based on sub-group clustering along the spatial dimension, while the TGR explores the structural information by measuring the correlation degrees between temporal joint node trajectories. The proposed scheme is naturally and seamlessly incorporated into the framework of graph convolutional networks (GCNs) to produce a set of skeleton-joint-connectivity graphs, which are further fed into the classification networks. Moreover, an insightful analysis on receptive field of graph node is provided to explain the necessity of our method. Experimental results on two benchmark datasets (NTU-RGB+D and Kinetics) demonstrate the effectiveness against the state-of-the-art. Bin Li 0038, Xi Li 0001, Zhongfei Zhang, Fei Wu 0001 |
AAAI | 1 |
| 2019 | Text Guided Person Image SynthesisabstractThis paper presents a novel method to manipulate the visual appearance (pose and attribute) of a person image according to natural language descriptions. Our method can be boiled down to two stages: 1) text guided pose generation and 2) visual appearance transferred image synthesis. In the first stage, our method infers a reasonable target human pose based on the text. In the second stage, our method synthesizes a realistic and appearance transferred person image according to the text in conjunction with the target pose. Our method extracts sufficient information from the text and establishes a mapping between the image space and the language space, making generating and editing images corresponding to the description possible. We conduct extensive experiments to reveal the effectiveness of our method, as well as using the VQA Perceptual Score as a metric for evaluating the method. It shows for the first time that we can automatically edit the person image from the natural language descriptions. Xingran Zhou, Siyu Huang, Bin Li 0038, Yingming Li, Zhongfei Zhang |
CVPR | 3 |