VLDB 2026 Research / reviewers in the wild / expert
Longrong Yang
dblp:253/7951
· DBLP profile ↗
11ranked-venue papers
10as first author
8since 2021 · last 2025
0000-0002-2433-2099ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Libra-Merging: Importance-redundancy and Pruning-merging Trade-off for Acceleration Plug-in in Large Vision-Language ModelabstractLarge Vision-Language Models (LVLMs) have achieved significant progress in recent years. However, the expensive inference cost limits the realistic deployment of LVLMs. Some works find that visual tokens are redundant and compress tokens to reduce the inference cost. These works identify important non-redundant tokens as target tokens, then prune the remaining tokens (non-target tokens) or merge them into target tokens. However, target token identification faces the token importance-redundancy dilemma. Besides, token merging and pruning face a dilemma between disrupting target token information and losing non-target token information. To solve these problems, we propose a novel visual token compression scheme, named Libra-Merging. In target token identification, Libra-Merging selects the most important tokens from spatially discrete intervals, achieving a more robust token importance-redundancy trade-off than relying on a hyper-parameter. In token compression, when non-target tokens are dissimilar to target tokens, Libra-Merging does not merge them into the target tokens, thus avoiding disrupting target token information. Meanwhile, Libra-Merging condenses these non-target tokens into an information compensation token to prevent losing important non-target token information. Our method can serve as a plug-in for diverse LVLMs, and extensive experimental results demonstrate its effectiveness. The code will be publicly available at https://github.com/longrongyang/Libra-Merging. Longrong Yang, Dong Shen 0003, Chaoxiang Cai, Kaibing Chen, Fan Yang 0094, Tingting Gao, Di Zhang 0026 |
CVPR | 1 |
| 2025 | Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language ModelabstractThe Mixture-of-Experts (MoE) has gained increasing attention in studying Large Vision-Language Models (LVLMs). It uses a sparse model to replace the dense model, achieving comparable performance while activating fewer parameters during inference, thus significantly reducing the inference cost. Existing MoE methods in LVLM encourage different experts to specialize in different tokens, and they usually employ a router to predict the routing of each token. However, the router is not optimized concerning distinct parameter optimization directions generated from tokens within an expert. This may lead to severe interference between tokens within an expert. To address this problem, we propose to use the token-level gradient analysis to Solving Token Gradient Conflict (STGC) in this paper. Specifically, we first use token-level gradients to identify conflicting tokens in experts. After that, we add a regularization loss tailored to encourage conflicting tokens routing from their current experts to other experts, for reducing interference between tokens within an expert. Our method can serve as a plug-in for diverse LVLM methods, and extensive experimental results demonstrate its effectiveness. demonstrate its effectiveness.
The code will be publicly available at https://github.com/longrongyang/STGC. Longrong Yang, Dong Shen 0003, Chaoxiang Cai, Fan Yang 0094, Tingting Gao, Di Zhang 0026 |
ICLR | 1 |
| 2025 | GCSTG: Generating Class-Confusion-Aware Samples With a Tree-Structure Graph for Few-Shot Object DetectionabstractFew-Shot Object Detection (FSOD) aims to detect the objects of novel classes using only a few manually annotated samples. With the few novel class samples, learning the inter-class relationships among foreground and constructing the corresponding class hierarchy in FSOD is a challenging task. The poor construction of the class hierarchy will result in the inter-class confusion problem, which has been identified as a primary cause of inferior performance in novel classes by recent FSOD methods. In this work, we further find that the intra-super-class confusion, where samples are misclassified as classes within their associated super-classes, is the main challenge in solving the confusion problem. To solve this issue, this work generates class-confusion-aware samples with a pre-defined tree-structure graph, for helping models to construct a precise class hierarchy. In precise, for generating class-confusion-aware samples, we add the noise into available samples and update the noise to maximize confidence scores on associated confusion categories of samples. Then, a confusion-aware curriculum learning strategy is proposed to make generated samples gradually participate in the training, which benefits the model convergence while learning the generated samples. Experimental results show that our method can be used as a plug-in in recent FSOD methods and consistently improve the model performance. Longrong Yang, Hanbin Zhao, Hongliang Li 0001, Liang Qiao 0001, Xi Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | RCS-Prompt: Learning Prompt to Rearrange Class Space for Prompt-Based Continual Learning
Longrong Yang, Hanbin Zhao, Yunlong Yu 0001, Xiaodong Zeng, Xi Li 0001 |
ECCV (47) | 1 |
| 2023 | Bridging Cross-task Protocol Inconsistency for Distillation in Dense Object DetectionabstractKnowledge distillation (KD) has shown potential for learning compact models in dense object detection. However, the commonly used softmax-based distillation ignores the absolute classification scores for individual categories. Thus, the optimum of the distillation loss does not necessarily lead to the optimal student classification scores for dense object detectors. This cross-task protocol inconsistency is critical, especially for dense object detectors, since the foreground categories are extremely imbalanced. To address the issue of protocol differences between distillation and classification, we propose a novel distillation method with cross-task consistent protocols, tailored for the dense object detection. For classification distillation, we address the cross-task protocol inconsistency problem by formulating the classification logit maps in both teacher and student models as multiple binary-classification maps and applying a binary-classification distillation loss to each map. For localization distillation, we design an IoU-based Localization Distillation Loss that is free from specific network structures and can be compared with existing localization distillation losses. Our proposed method is simple but effective, and experimental results demonstrate its superiority over existing methods. Code is available at https://github.com/TinyTigerPan/BCKD. Longrong Yang, Xianpan Zhou, Xuewei Li 0003, Liang Qiao 0001, Zheyang Li, Ziwei Yang 0004, Gaoang Wang, Xi Li 0001 |
ICCV | 1 |
| 2023 | Task-Specific Loss for Robust Instance Segmentation With Noisy Class LabelsabstractDeep learning methods have achieved significant progress in the presence of correctly annotated datasets in instance segmentation. However, object classes in large-scale datasets are sometimes ambiguous, which easily causes confusion. Besides, limited experience and knowledge of annotators can lead to mislabeled object semantic classes. To solve this issue, a novel method is proposed in this paper, which considers different roles of noisy class labels in different sub-tasks. Our method is based on two basic observations: firstly, the foreground-background annotation of a sample is correct even though its class label is noisy. Secondly, symmetric loss benefits the model robustness to noisy labels but harms the learning of hard samples, while cross entropy loss is the opposite. Based on the two basic observations, in the foreground-background sub-task, cross entropy loss is used to fully exploit correct gradient guidance. In the foreground-instance sub-task, symmetric loss is used to prevent incorrect gradient guidance provided by noisy class labels. Furthermore, we apply contrastive self-supervised loss to update features of all foreground, to compensate for insufficient guidance provided by partially correct labels especially in the highly noisy setting. Extensive experiments conducted with three popular datasets (i.e., Pascal VOC, Cityscapes and COCO) have demonstrated the effectiveness of our method in a wide range of noisy class label scenarios. Longrong Yang, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Bias-Correction Feature Learner for Semi-Supervised Instance SegmentationabstractInstance segmentation is heavily reliant on large-scale annotated datasets to yield an ideal accuracy. However, annotated data are difficult to collect. To expand the annotated data, a straightforward idea is to introduce semi-supervised learning, which uses a trained model to obtain initial proposals on unlabeled images and then use initial proposals to generate pseudo labels. However, existing methods inevitably introduce the bias for the model learning, i.e., the foreground in initial low-confident proposals (low-confident foreground) is arbitrarily assigned as background. This bias makes the foreground and background closer in the feature space, which degenerates the model accuracy. To address this issue, this paper discards incorrect supervision and designs a bias-correction feature learner. Specifically, on the one hand, low-confident foreground does not participate in supervised learning. On the other hand, we extract possible foreground regions from all initial proposals to construct high-quality positive pairs which depict objects of the same category in contrastive learning. Then, positive pairs are pulled closer in the feature space. This helps models extract closely clustered foreground features. Experimental results demonstrate the effectiveness of our method on the public datasets (i.e., COCO, Cityscapes and Pascal VOC). Longrong Yang, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, Heqian Qiu, Linfeng Xu 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | DE-CrossDet: Divisible and Extensible Crossline Representation for Object DetectionabstractObject detection aims to localize and classify objects. Suitable object representation plays an important role in accurate detection. Because a complete crossline inevitably passes through the noise of backgrounds or other objects, object features directly extracted by the whole crossline are often confused. In this paper, we present a new feature extraction method, DE-Crossline, which can enhance the original crossline representation to capture more accurate object information. Specifically, we divide the crossline into several segments, each of which extracts the maximum activation key point respectively to reduce the impact of noise mentioned above. Furthermore, considering various shapes and sizes of objects, we design a Deformable Width Extension Module to learn a suitable width of each crossline, so as to capture richer object information. Extensive experiments prove the effectiveness of our proposed method. The total performance of our proposed detector can reach 49.0% AP, using ResNet-101 as backbone on the MS-COCO dataset. Hefei Mei, Hongliang Li 0001, Heqian Qiu, Jianhua Cui, Longrong Yang |
VCIP | 5 |
| 2020 | Learning with Noisy Class Labels for Instance Segmentation
Longrong Yang, Fanman Meng, Hongliang Li 0001, Qingbo Wu 0001, Qishang Cheng |
ECCV (14) | 1 |
| 2020 | Mono is Enough: Instance Segmentation from Single Annotated SampleabstractWith the help of various Deep Neural Networks, instance segmentation has achieved significant progress. How-ever, these successes are heavily reliant on large-scale manually annotated samples, which are extremely time-consuming and expensive. To address this issue, we propose a highly efficient anisotropic data augmentation method, which generates high quality training data from a single manually annotated sample. Instead of equivalently modifying foreground and background like traditional data augmentation methods, we focus on enriching the diversities of foreground appearance and positional relation between foreground and background, which are beneficial for the classification and localization sub-tasks respectively. All foreground instances of the source annotated sample undergo various rotation, brightness change, rescale, distortion and frequency-component mixup (FCM). Then, these modified instances are randomly embedded into background, which serve as new training samples. Experiments on Cityscapes dataset show that our method significantly outperforms traditional data augmentation methods. Longrong Yang, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng, King Ngi Ngan |
VCIP | 1 |
| 2019 | Incorporating Non-local and Task-specific Features for Instance SegmentationabstractThis paper proposes a novel instance segmentation model, which improves the instance segmentation by considering two aspects. One is a new non-local features module to recover detailed information that is lost in the deep convolutional operations. The other is to introduce attention mechanism to generate specific features adaptive to each task. The proposed method is verified on three well-known datasets, namely Pascal VOC, Cityscapes and COCO. The experiments show that the method using the proposed modules outperforms baseline Mask R-CNN on all of the datasets without bells and whistles. Longrong Yang, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001 |
MMSP | 1 |