Kai Xu 0012

dblp:30/495-12 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-4334-2444ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Deep Residual Discriminative Dictionary Learning for image classification
Lichun Wang 0002, Jianjia Xin, Kai Xu 0012, Huiyong Zhang, Shaofan Wang 0001, Dehui Kong
Knowl. Based Syst.4
2026 VRGraspNet: Toward Viewpoint Robust 6-DoF Grasp Pose Estimation
abstract
6-DoF grasp pose estimation is crucial for achieving robust robot manipulation. Despite significant progress in data-driven methods, the cross-view adaptability of 6-DoF grasp pose estimation remains insufficient. Under the condition of observing some certain viewpoints, the performance of pose estimation is relatively low. In response, this paper introduces VRGraspNet, a novel 6-DoF grasp pose estimation model designed to enhance the robustness and performance of robotic grasping across diverse viewpoints. The key to the VRGraspNet lies in filling in the holes in the scene point cloud and learning multimodal features for seed points. The former provides dense neighborhood points for the seed point, while the latter provides richer information for extracting geometric features of the local area to which the seed point belongs. Furthermore, this paper proposes a Performance-Viewpoint jointly Weighted Loss (PVWL), with the key being two weight factors: a static viewpoint position dependent weight factor that focuses on challenging samples, and a dynamic performance related weight factor that focuses on samples difficult to learn. Extensive experiments on the GraspNet-1Billion dataset demonstrate that VRGraspNet achieves SOTA performance and strong cross-view robustness. Real-world robot experiments further validate its practicality in robotic manipulation tasks. Our source code is available at https://github.com/huamo555/VRGraspNet.
Yuming Gao, Lichun Wang 0002, Jiaqi Zheng 0020, Kai Xu 0012, Huayang Yao
IEEE Trans. Circuits Syst. Video Technol.4
2026 Self-Knowledge Distillation and Its Application: A Survey
abstract
As an efficient model compression technique, knowledge distillation has become an important research topic in the field of deep learning. However, the requirement of pre-trained teacher networks makes the process cumbersome and inefficient, which prompted researchers to propose a more efficient mechanism. Therefore, self-knowledge distillation is proposed, which does not require assistance from additional teacher networks. In previous surveys, self-knowledge distillation has usually been considered a special case of knowledge distillation. In recent years, significant progress has been made in self-knowledge distillation, which has evolved beyond the functions or roles of traditional knowledge distillation. However, there is no individual and intensive survey of self-knowledge distillation methods up to now. Therefore, this paper reviews and investigates existing self-knowledge distillation methods from a comprehensive perspective. Specifically, first, according to different sources of knowledge, this paper categorizes self-knowledge distillation methods into three types, label knowledge-based, feature knowledge-based and data knowledge-based. Then, this paper introduces the evaluation protocol and performance of SKD. In particular, the commonly used experimental datasets and evaluation networks are summarized, aiming to encourage researchers to choose common network architectures and evaluation datasets for promoting the standardization of self-knowledge distillation's comparison. Finally, this paper introduces the applications of self-knowledge distillation in different task scenarios, enabling researchers to quickly locate the relevant task fields. Furthermore, this paper introduces the technical purposes of applying self-knowledge distillation and explores the core motivations for using SKD across different tasks, promoting researchers' extensive attempts at self-distillation technology in their own research fields that have not yet applied self-knowledge distillation.
Kai Xu 0012, Lichun Wang 0002, Huiyong Zhang
IEEE Trans. Knowl. Data Eng.1
2025 Ensemble predicate decoding for unbiased scene graph generation
Jiasong Feng, Lichun Wang 0002, Kai Xu 0012
Neurocomputing4
2025 Bcgn: BLIP-based cross-modal grasping network for language-conditioned robotic grasping
Kai Xu 0012, Lichun Wang 0002, Jianjia Xin
Multim. Syst.1
2025 Scene Adaptive Context Modeling and Balanced Relation Prediction for Scene Graph Generation
abstract
Scene graph generation (SGG) aims to perceive objects and their relations in images, which can bridge the gap between upstream detection tasks and downstream high-level visual understanding tasks. For SGG models, over-fitting head predicates can lead to bias in the generated scene graph, which has become a consensus. A series of debiasing methods have been proposed to solve the problem. However, some existing debiasing SGG methods have a tendency to over-fit tail predicates, which is another type of bias. In order to eliminate the one-way over-fitting of head or tail predicates, this article proposes a balanced relation prediction (BRP) module which is model-agnostic and compatible with existing re-balancing methods. Moreover, because the relation prediction is based on object feature representation, this article proposes a scene adaptive context fusion (SACF) module to refine the object feature representation. Specifically, SACF models the context based on a chain structure, where the order of objects in the chain structure is adaptively arranged according to the scene content, achieving visual information fusion that adapts to the scene where the objects are located. Experiments on VG and GQA datasets show that the proposed method achieves competitive results on the comprehensive metric of R@K and mR@K.
Kai Xu 0012, Lichun Wang 0002
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Self-Knowledge Distillation with Learning from Role-Model Samples
abstract
Self-knowledge distillation does not require a pre-trained teacher network like traditional knowledge distillation. Existing methods either require additional parameters or require additional memory consumption. To alleviate this problem, this paper proposes a more efficient self-knowledge distillation method, named LRMS (learning from role-model samples). In every mini-batch, LRMS selects out a role-model sample for each sampled category, and takes its prediction as the proxy semantic for the corresponding category. Then, predictions of the other samples are constrained to be consistent with the proxy semantics, which makes the distribution of predictions for samples within the same category more compact. Meanwhile, the regularization targets corresponding to proxy semantics are set with a higher distillation temperature to better utilize the classificatory information about the categories. Experimental results show that diverse architectures achieve improvements on four image classification datasets by using LRMS. Code is acaliable: https://github.com/KAI1179/LRMS
Kai Xu 0012, Lichun Wang 0002, Huiyong Zhang
ICASSP1
2024 Area-keywords cross-modal alignment for referring image segmentation
Huiyong Zhang, Lichun Wang 0002, Kai Xu 0012
Neurocomputing4
2024 A New Training Data Organization Form and Training Mode for Unbiased Scene Graph Generation
abstract
The current mainstream studies on Scene Graph Generation (SGG) devote to the long-tailed predicate distribution problem to generate unbiased scene graph. The long-tailed predicate distribution exists in VG dataset and is more severe during the SGG network training process. Most existing de-biasing methods solve the problem by applying re-sampling or re-weighting in a mini-batch, with the main idea being to provide unbiased attention to different predicate categories based on prior predicate distributions. During the training process of SGG models, existing training mode samples several images into a mini-batch to obtain training data, thus providing sparse and scattered predicate instances for training. However, sampling predicate instances from a limited set of predicate samples in terms of quantity and category poses difficulties in training unbiased SGG models. In order to provide a wider range for sampling predicate instances, this paper reorganizes the images in VG training set with a new form, i.e. object-pairs, and constructs VG-OP (VG Object-Pair) training set to save object-pairs. Meanwhile, this paper introduces a new SGG network training mode, which can realize unbiased SGG without resampling or re-weighting. In particular, a Predicate-balanced Sampling Network (PS-Net) is designed to validate the new training mode. Extensive experiments on VG test set demonstrate that our method achieves competitive or state-of-the-art unbiased SGG performance.
Lichun Wang 0002, Kai Xu 0012, Fangyu Fu, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.3
2024 Self-Distillation With Augmentation in Feature Space
abstract
Compared with traditional knowledge distillation, self-distillation does not require a pre-trained teacher network, which is more concise. Among them, data augmentation-based methods provide an elegant solution without modifying the network structure or additional memory consumption. However, when employing data augmentation in the input space, the forward propagations for augmented data bring additional computation costs and the augmentation methods need be adaptive to the modality of input data. Meanwhile, we note that from a generalization perspective, under the condition of being able to distinguish from other classes, a dispersed intra-class feature distribution is superior to compact intra-class feature distribution, especially for categories with larger sample differences. Based on the above considerations, this paper proposes a feature augmentation based self-distillation method (FASD) based on the idea of feature extrapolation. For each source feature, two augmentations are generated by subtraction between features. The one is subtracting the temporary class center computed with samples belonging to the same category, and another one is subtracting a sample feature belonging to other categories with the closest distance. Then, the predicted outputs of the augmented features are constrained to be consistent with that of the source feature. The consistent constraint on the previous augmented feature expands the learned class feature distribution, leading to greater overlap with the unknown feature distribution of test samples, thereby improving the generalization performance of the network. The consistent constraint on the latter augmented feature increases the distance between samples from different categories, which enhances the distinguishability between categories. Experimental results on image classification task demonstrate the effectiveness and efficiency of the proposed method. Meanwhile, experiments on text and audio tasks prove the universality of the method for classification tasks with different modalities.
Kai Xu 0012, Lichun Wang 0002, Jianjia Xin
IEEE Trans. Circuits Syst. Video Technol.1
2024 Learning From Teacher's Failure: A Reflective Learning Paradigm for Knowledge Distillation
abstract
Knowledge Distillation transfers knowledge learned by a teacher network to a student network. A common mode of knowledge transfer is directly using the teacher network’s experience for all samples without differentiating whether the experience of teacher is successful or not. According to common sense, experience varies with its nature. Successful experience is used for guidance, and failed experience is used for correction. Inspired by that, this paper analyzes the failure of teacher and proposes a reflective learning paradigm, which additionally uses heuristic knowledge extracted from the teacher’s failure besides following the authority of teacher. Specifically, this paper defines Mutual Error Distance (MED) based on the teacher’s wrong predictions. MED measures the adequacy of the decision boundary learned by teacher, which concretizes the failure of teacher. Then, this paper proposes DCGD (divide-and-conquer grouping distillation) to critically transfer the teacher’s knowledge by grouping the target task into small-scale subtasks and designing multi-branch networks on the basis of MED. Finally, a switchable training mechanism is designed to integrate a regular student which provides an option of student network without parameter addition compared with the multi-branch student network. Extensive experiments on three image classification benchmarks (CIFAR-10, CIFAR-100 and TinyImageNet) show the effectiveness of the proposed paradigm. Especially on CIFAR-100 dataset, the average error of students using DCGD+DKD decreased by 4.28%. In addition, the experiment results show that the paradigm is also applicable to self-distillation.
Kai Xu 0012, Lichun Wang 0002, Jianjia Xin
IEEE Trans. Circuits Syst. Video Technol.1
2023 A Balanced Relation Prediction Framework for Scene Graph Generation
Kai Xu 0012, Lichun Wang 0002, Huiyong Zhang
ICANN (4)1
2023 A Novel Encoder and Label Assignment for Instance Segmentation
Huiyong Zhang, Lichun Wang 0002, Kai Xu 0012
ICANN (6)4
2023 Scale-Adaptive Multi-area Representation for Instance Segmentation
Huiyong Zhang, Lichun Wang 0002, Kai Xu 0012
ICIG (4)4
2023 Augmented Spatial Context Fusion Network for Scene Graph Generation
abstract
Scene graph generation provides high-order semantic information by understanding the objects and their relations in images. In order to improve the performance of scene graph generation, context fusion has been widely used in scene graph generation tasks, LSTM and Vision-Transformer are commonly used fusion modules. Both LSTM and Vision-Transformer realize context fusion by stacking multiple basic units, which needs to learn a large number of parameters of the units. However, the model computational efficiency of scene graph generation as a mid-level semantic understanding task to support downstream tasks is crucial. To simplify the context fusion computation, this paper proposes ASCF -Net (Augmented Spatial Context Fusion Network) which computes the spatial context of designated object by searching the nearest neighbor objects with high relevance and strengthens the context with random noise. Without learning parameters, the above computational process essentially simulates the attention mechanism. Experiments on VG dataset show that ASCF -Net uses 15.26% of the parameters of Bi-LSTM and 13.34% of the parameters of Vision-Transformer for context fusion based on the same baseline and achieves higher performance than using the two fusion modules. At the same time, ASCF -Net uses simple fusion module to obtain competitive results on VG dataset compared with the mainstream scene generation models.
Lichun Wang 0002, Kai Xu 0012, Fangyu Fu, Qingming Huang
IJCNN3