EDBT 2026 Demo / reviewers in the wild / expert
Xuewei Li 0003
dblp:43/3869-3
· DBLP profile ↗
15ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-3414-8754ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CWPS: Efficient Channel-Wise Parameter Sharing for Knowledge TransferabstractKnowledge transfer aims to apply existing knowledge to different tasks or new data, and it has extensive applications in multi-domain and Multi-Task Learning. The key to this task is quickly identifying a fine-grained object for knowledge sharing and efficiently transferring knowledge. Current methods, such as fine-tuning, layer-wise parameter sharing, and task-specific adapters, only offer coarse-grained sharing solutions and struggle to effectively search for shared parameters, thus hindering the performance and efficiency of knowledge transfer. To address these issues, we propose Channel-Wise Parameter Sharing (CWPS), a novel fine-grained parameter-sharing method for knowledge transfer, which is efficient for parameter sharing, comprehensive, and plug-and-play. For the coarse-grained problem, we first achieve fine-grained parameter sharing by refining the granularity of shared parameters from the level of layers to the level of neurons. The knowledge learned from previous tasks can be utilized through the explicit composition of the model neurons. Besides, we promote an effective search strategy to minimize computational costs, simplifying the selection of shared weights. In addition, our CWPS has strong composability and generalization ability, which theoretically can be applied to any network consisting of linear and convolution layers. We introduce several datasets in both Incremental Learning and Multi-Task Learning scenarios. Our method has achieved state-of-the-art precision-to-parameter ratio performance with various backbones, demonstrating its efficiency and versatility. Mingxuan Cui, Xuewei Li 0003, Cunzheng Wang, Gaoang Wang, Chenyi Zhuang, Jinjie Gu, Xiubo Liang, Xi Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Exploring Vision-Based Active 3D Object Detection by Informativeness CharacterizationabstractVision-based 3D object detection (3DOD) gains lots of attention due to its low cost for deployment compared to Lidar-based tasks, while it suffers from labor-expensive data annotations. At the same time, active learning (AL) has shown great potential in reducing annotation costs in related tasks, which can maximize model performance within very limited labeled data. In this paper, we explore active learning for vision-based 3DOD for the first time. Inspired by the entropy analysis, we involve three concerns to characterize the sample informativeness: sample diversity in input space, feature informativeness in BEV space, and result distribution in prediction space. Based on these concerns, we propose a novel AL framework named HMAD, which utilizes Height Modeling and Adaptive Diversity-based sampling for comprehensive informativeness characterization. In HMAD, we first propose a novel height-guided adversarial module in BEV space, which measures the informativeness of height modeling for 2D-to-3D mapping in an adversarial manner. Furthermore, Budget-aware SpatioTemporal diversity Sampling (BSTS) and Class Balance Sampling (CBS) are proposed to adaptively measure the sample informativeness in input and prediction space, respectively. Finally, the three components are integrated into a two-stage sampling strategy, with which the most informative samples can be selected and annotated for the next iteration. Experiments evidence that HMAD achieves comparable performances by only using 50% annotated training data, and can generalize well on different conditions. Yiming Wu 0006, Yehao Lu, Xuewei Li 0003, Xiubo Liang, Xi Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Enhancing Tiny Object Detection Using Guided Object Inference Slicing (GOIS): An efficient dynamic adaptive framework for fine-tuned and non-fine-tuned deep learning models
Muhammad Muzammul, Xuewei Li 0003, Xi Li 0001 |
Neurocomputing | 2 |
| 2025 | Relationship-Incremental Scene Graph Generation by a Divide-and-Conquer Pipeline With Feature AdapterabstractAs a challenging computer vision task, Scene Graph Generation (SGG) finds the latent semantic relationships among objects from a given image, which may be limited by the datasets and real-world scenarios. In this paper, we consider a novel incremental learning task called Relationship-Incremental Scene Graph Generation (RISGG) that learns the semantic relationships among objects in an incremental way. Compared with classic Class-Incremental Learning (CIL) problem, RISGG suffers from its special issues: 1) Old class shift - the relationship-labeled object pair may have different labels during different learning sessions; 2) Background shift - the relationship-unlabeled object pair may not be a real unlabeled one. In this work, we address the above issues from the following aspects. First, we present a Divide-and-Conquer (DaC) pipeline to deal with the old class shift via decoupling the recognition of relationship classes and recognizing relationships individually. In this way, label confusion and interaction among different relationships are eliminated during training. Second, we propose a Feature Adapter (FA) to bridge the feature space gap between the current session and the previous one and use our extra supervision to mine old relationship information in the current session. Our proposed network combined DaC and FA, abbreviated DaCFA-Net, for RISGG. Experimental results on the benchmark dataset demonstrate the significant performance gain of DaCFA-Net in RISGG. It gains about 20% improvement against the SGG baselines on the popular VG dataset. Xuewei Li 0003, Guangcong Zheng, Yunlong Yu 0001, Naye Ji, Xi Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Decoupling Discriminative Attributes for Few-Shot Fine-Grained RecognitionabstractFew-shot fine-tuning of pre-trained vision-language models (VLMs) for downstream tasks has gained widespread attention for reducing data annotation efforts while maintaining high performance. However, we observe that VLMs excel in excluding most incorrect classes in fine-grained recognition tasks, but struggles with a small set of confusing categories, which are typically highly similar subspecies. Existing few-shot fine-tuning methods attempt to directly recognize the correct category among all predefined classes, limiting their ability to capture discriminative features for those confusing categories. This raises an intriguing question: Can we specifically extract useful information from confusing classes to enhance fine-grained recognition performance? Based on this insight, we propose a hierarchical few-shot fine-tuning framework to address the severe confusion problem while ensuring the interpretability, namely Attribute-Decoupled Discriminator (AttrDD). Instead of thinking once among all classes, AttrDD employs a two-stage recognition, "think through" then "think smart". Specifically, in the first phase, a representative VLM, CLIP, is fine-tuned to select the Top-K confusing classes. In the second phase, we leverage the knowledge of large language models (LLMs) to generate fixed format descriptions of attribute differences between these confusing classes via in-context learning. Attribute-decoupled classifications are then conducted to capture fine-grained discriminative features. To achieve parameter-efficient fine-tuning, we introduce a lightweight attention adapter for each phase to align image features with task-specific textual features and LLM-generated textual features. Extensive experiments on 9 fine-grained recognition benchmarks demonstrate that AttrDD consistently outperforms existing baselines by wide margins. Yehao Lu, Chaoxiang Cai, Wei Su 0009, Guangcong Zheng, Xuewei Li 0003, Xi Li 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | GAMA-Pose: Graph-Aware Multi-Representation Aggregation for 3D Human Pose EstimationabstractMonocular 3D human pose estimation presents a considerable challenge owing to the intrinsic depth ambiguity associated with single-camera observations. Existing methods primarily rely on mean per joint position error (MPJPE) loss to train models for the conversion from 2D to 3D coordinates. However, empirical analysis reveals that models trained solely with point-based supervision may produce biomechanically implausible poses or exhibit significant depth ambiguity, even when achieving low MPJPE. This limitation arises from the fact that point-based loss only considers individual joint locations without accounting for inter-joint relationships. Fortunately, edges of human pose encode critical prior knowledge, including skeleton connectivity and biomechanical distributions. Explicitly modeling edge representations enables the model to overcome the constraints associated with point-only approaches, reducing the uncertainty in the optimization process of the 2D-3D inverse mapping and directly constraining depth ambiguity. Therefore, we propose the Graph-Aware Multi-Representation Aggregation (GAMA-Pose) framework that jointly predicts points and edges, with their fusion serving as the final output. To ensure the accuracy of edge predictions and mitigate depth ambiguity, Anti-Depth-Ambiguity Loss (ADA-Loss) is introduced to supervise the properties of edges and give direct supervision on depth ambiguity. Correspondingly, edge-based metrics are proposed to quantify the error of predicted edges. Experiments conducted on Human3.6M and MPI-INF-3DHP datasets demonstrate that GAMA-Pose effectively addresses the limitations of models relying solely on point constraints, mitigates depth ambiguity, enhances the accuracy of both point and edge predictions, and achieves state-of-the-art (SOTA) performance on both datasets. Songran Zhou, Xuewei Li 0003, Xiubo Liang, Naye Ji, Xi Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion ModelabstractControllable spherical panoramic image generation holds substantial applicative potential across a variety of domains. However, it remains a challenging task due to the inherent spherical distortion and geometry characteristics, resulting in low-quality content generation. In this paper, we introduce a novel framework of SphereDiffusion to address these unique challenges, for better generating high-quality and precisely controllable spherical panoramic images. For the spherical distortion characteristic, we embed the semantics of the distorted object with text encoding, then explicitly construct the relationship with text-object correspondence to better use the pre-trained knowledge of the planar images. Meanwhile, we employ a deformable technique to mitigate the semantic deviation in latent space caused by spherical distortion. For the spherical geometry characteristic, in virtue of spherical rotation invariance, we improve the data diversity and optimization objectives in the training process, enabling the model to better learn the spherical geometry characteristic. Furthermore, we enhance the denoising process of the diffusion model, enabling it to effectively use the learned geometric characteristic to ensure the boundary continuity of the generated images. With these specific techniques, experiments on Structured3D dataset show that SphereDiffusion significantly improves the quality of controllable spherical image generation and relatively reduces around 35% FID on average. Xuewei Li 0003, Zhongang Qi, Xintao Wang 0002, Ying Shan, Xi Li 0001 |
AAAI | 2 |
| 2024 | MLMG-SGG: Multilabel Scene Graph Generation With Multigrained FeaturesabstractAs an important and challenging problem in computer vision, scene graph generation (SGG) aims to find out the underlying semantic relationships among objects from a given image for scene understanding. Usually, prevalent SGG approaches adopt a learning pipeline with the assumption that there exists only a single relationship for a particular object pair. Considering the common phenomenon that a pair of objects can be attached by multiple relationships, we propose a multi-label scene graph generation pipeline with multi-grained features (MLMG-SGG), which formulates the relationship detection as a multi-label classification problem during training while generating multigraphs at inference time. In order to better model the fine-grained relationships, the proposed pipeline encodes the feature representation of SGG on different spatial scales by a specially designed Multi-Grained Module (MGM), resulting in the multi-grained (i.e., object-level and region-level) features of objects. Experimental results over the benchmark dataset demonstrate the significant performance gain of the proposed pipeline used as a plug-in for the state-of-the-art methods. Xuewei Li 0003, Peihan Miao 0002, Xi Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Self-Paced Multi-Grained Cross-Modal Interaction Modeling for Referring Expression ComprehensionabstractAs an important and challenging problem in vision-language tasks, referring expression comprehension (REC) generally requires a large amount of multi-grained information of visual and linguistic modalities to realize accurate reasoning. In addition, due to the diversity of visual scenes and the variation of linguistic expressions, some hard examples have much more abundant multi-grained information than others. How to aggregate multi-grained information from different modalities and extract abundant knowledge from hard examples is crucial in the REC task. To address aforementioned challenges, in this paper, we propose a Self-paced Multi-grained Cross-modal Interaction Modeling framework, which improves the language-to-vision localization ability through innovations in network structure and learning mechanism. Concretely, we design a transformer-based multi-grained cross-modal attention, which effectively utilizes the inherent multi-grained information in visual and linguistic encoders. Furthermore, considering the large variance of samples, we propose a self-paced sample informativeness learning to adaptively enhance the network learning for samples containing abundant multi-grained information. The proposed framework significantly outperforms state-of-the-art methods on widely used datasets, such as RefCOCO, RefCOCO+, RefCOCOg, and ReferItGame datasets, demonstrating the effectiveness of our method. Peihan Miao 0002, Wei Su 0009, Gaoang Wang, Xuewei Li 0003, Xi Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Epoch-Evolving Gaussian Process Guided Learning for ClassificationabstractThe conventional mini-batch gradient descent algorithms are usually trapped in the local batch-level distribution information, resulting in the ``zig-zag'' effect in the learning process. To characterize the correlation information between the batch-level distribution and the global data distribution, we propose a novel learning scheme called epoch-evolving Gaussian process guided learning (GPGL) to encode the global data distribution information in a non-parametric way. Upon a set of class-aware anchor samples, our GP model is built to estimate the class distribution for each sample in mini-batch through label propagation from the anchor samples to the batch samples. The class distribution, also named the context label, is provided as a complement for the ground-truth one-hot label. Such a class distribution structure has a smooth property and usually carries a rich body of contextual information that is capable of speeding up the convergence process. With the guidance of the context label and ground-truth label, the GPGL scheme provides a more efficient optimization through updating the model parameters with a triangle consistency loss. Furthermore, our GPGL scheme can be generalized and naturally applied to the current deep models, outperforming the state-of-the-art optimization methods on six benchmark datasets. Jiabao Cui, Xuewei Li 0003, Hanbin Zhao, Hui Wang 0107, Bin Li 0038, Xi Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | LayoutDiffusion: Controllable Diffusion Model for Layout-to-Image GenerationabstractRecently, diffusion models have achieved great success in image synthesis. However, when it comes to the layout-to-image generation where an image often has a complex scene of multiple objects, how to make strong control over both the global layout map and each detailed object remains a challenging task. In this paper, we propose a diffusion model named LayoutDiffusion that can obtain higher generation quality and greater controllability than the previous works. To overcome the difficult multimodal fusion of image and layout, we propose to construct a structural image patch with region information and transform the patched image into a special layout to fuse with the normal layout in a unified form. Moreover, Layout Fusion Module (LFM) and Object-aware Cross Attention (OaCA) are proposed to model the relationship among multiple objects and designed to be object-aware and position-sensitive, allowing for precisely controlling the spatial related information. Extensive experiments show that our LayoutDiffusion out-performs the previous SOTA methods on FID, CAS by relatively 46.35%,26.70% on COCO-stuff and 44.29%,41.82% on VG. Code is available at https://github.com/ZGCTroy/LayoutDiffusion. Guangcong Zheng, Xianpan Zhou, Xuewei Li 0003, Zhongang Qi, Ying Shan, Xi Li 0001 |
CVPR | 3 |
| 2023 | Bridging Cross-task Protocol Inconsistency for Distillation in Dense Object DetectionabstractKnowledge distillation (KD) has shown potential for learning compact models in dense object detection. However, the commonly used softmax-based distillation ignores the absolute classification scores for individual categories. Thus, the optimum of the distillation loss does not necessarily lead to the optimal student classification scores for dense object detectors. This cross-task protocol inconsistency is critical, especially for dense object detectors, since the foreground categories are extremely imbalanced. To address the issue of protocol differences between distillation and classification, we propose a novel distillation method with cross-task consistent protocols, tailored for the dense object detection. For classification distillation, we address the cross-task protocol inconsistency problem by formulating the classification logit maps in both teacher and student models as multiple binary-classification maps and applying a binary-classification distillation loss to each map. For localization distillation, we design an IoU-based Localization Distillation Loss that is free from specific network structures and can be compared with existing localization distillation losses. Our proposed method is simple but effective, and experimental results demonstrate its superiority over existing methods. Code is available at https://github.com/TinyTigerPan/BCKD. Longrong Yang, Xianpan Zhou, Xuewei Li 0003, Liang Qiao 0001, Zheyang Li, Ziwei Yang 0004, Gaoang Wang, Xi Li 0001 |
ICCV | 3 |
| 2023 | SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic SegmentationabstractAs an important and challenging problem in computer vision, PAnoramic Semantic Segmentation (PASS) gives complete scene perception based on an ultra-wide angle of view. Usually, prevalent PASS methods with 2D panoramic image input focus on solving image distortions but lack consideration of the 3D properties of original 360 degree data. Therefore, their performance will drop a lot when inputting panoramic images with the 3D disturbance. To be more robust to 3D disturbance, we propose our Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation (SGAT4PASS), considering 3D spherical geometry knowledge. Specifically, a spherical geometry-aware framework is proposed for PASS. It includes three modules, i.e., spherical geometry-aware image projection, spherical deformable patch embedding, and a panorama-aware loss, which takes input images with 3D disturbance into account, adds a spherical geometry-aware constraint on the existing deformable patch embedding, and indicates the pixel density of original 360 degree data, respectively. Experimental results on Stanford2D3D Panoramic datasets show that SGAT4PASS significantly improves performance and robustness, with approximately a 2% increase in mIoU, and when small 3D disturbances occur in the data, the stability of our performance is improved by an order of magnitude. Our code and supplementary material are available at https://github.com/TencentARC/SGAT4PASS. Xuewei Li 0003, Zhongang Qi, Gaoang Wang, Ying Shan, Xi Li 0001 |
IJCAI | 1 |
| 2023 | Uncertainty-Aware Scene Graph Generation
Xuewei Li 0003, Guangcong Zheng, Yunlong Yu 0001, Xi Li 0001 |
Pattern Recognit. Lett. | 1 |
| 2021 | ResKD: Residual-Guided Knowledge DistillationabstractKnowledge distillation, aimed at transferring the knowledge from a heavy teacher network to a lightweight student network, has emerged as a promising technique for compressing neural networks. However, due to the capacity gap between the heavy teacher and the lightweight student, there still exists a significant performance gap between them. In this article, we see knowledge distillation in a fresh light, using the knowledge gap, or the residual, between a teacher and a student as guidance to train a much more lightweight student, called a res-student. We combine the student and the res-student into a new student, where the res-student rectifies the errors of the former student. Such a residual-guided process can be repeated until the user strikes the balance between accuracy and cost. At inference time, we propose a sample-adaptive strategy to decide which res-students are not necessary for each sample, which can save computational cost. Experimental results show that we achieve competitive performance with 18.04%, 23.14%, 53.59%, and 56.86% of the teachers' computational costs on the CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet datasets. Finally, we do thorough theoretical and empirical analysis for our method. Xuewei Li 0003, Omar El Farouk Bourahla, Fei Wu 0001, Xi Li 0001 |
IEEE Trans. Image Process. | 1 |