Kaiyou Song

dblp:216/9384 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-8999-2680ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
YearPublicationVenuePosition
2026 VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
abstract
Retrieval Augmented Generation enhances the response accuracy of Large Language Models (LLMs) by integrating retrieval and generation modules with external knowledge, demonstrating particular strength in real-time queries and Visual Question Answering tasks. However, the effectiveness of RAG is frequently hindered by the precision of the retriever: many retrieved samples fed into the generation phase are irrelevant or misleading, posing a critical bottleneck to LLMs’ performance. To address this challenge, we introduce \textbf{VaccineRAG}, a novel Chain-of-Thought-based retrieval-augmented generation dataset. On one hand, VaccineRAG employs a benchmark to evaluate models using data with varying positive/negative sample ratios, systematically exposing inherent weaknesses in current LLMs. On the other hand, it enhances models’ sample-discrimination capabilities by prompting LLMs to generate explicit Chain-of-Thought (CoT) analysis for each sample before producing final answers. Furthermore, to enhance the model’s ability to learn long-sequence complex CoT content, we propose \textbf{Partial-GRPO}. By modeling the outputs of LLMs as multiple components rather than a single whole, our model can make more informed preference selections for complex sequences, thereby enhancing its capacity to learn complex CoT. Comprehensive evaluations and ablation studies on VaccineRAG validate the effectiveness of the proposed scheme.
Qixin Sun, Ziqin Wang, Hengyuan Zhao, Kaiyou Song, Si Liu 0001, Xiaolin Hu 0001, Qingpei Guo, Linjiang Huang
AAAI5
2025 Bootstrap Masked Visual Modeling via Hard Patch Mining
abstract
Masked visual modeling has attracted much attention due to its promising potential in learning generalizable representations. Typical approaches urge models to predict specific contents of masked tokens, which can be intuitively considered as teaching a student (the model) to solve given problems (predicting masked contents). Under such settings, the performance is highly correlated with mask strategies (the difficulty of provided problems). We argue that it is equally important for the model to stand in the shoes of a teacher to produce challenging problems by itself. Intuitively, patches with high values of reconstruction loss can be regarded as hard samples, and masking those hard patches naturally becomes a demanding reconstruction task. To empower the model as a teacher, we propose Hard Patch Mining (HPM), predicting patch-wise losses and subsequently determining where to mask. Technically, we introduce an auxiliary loss predictor, which is trained with a relative objective to prevent overfitting to exact loss values. To gradually guide the training procedure, we propose an easy-to-hard mask strategy. Empirically, HPM brings significant improvements under both image and video benchmarks. Interestingly, solely incorporating the extra loss prediction objective leads to better representations, verifying the efficacy of determining where is hard to reconstruct.
Junsong Fan, Yuxi Wang 0001, Kaiyou Song, Tiancai Wang, Xiangyu Zhang 0005, Zhaoxiang Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Semantic-Aware Autoregressive Image Modeling for Visual Representation Learning
abstract
The development of autoregressive modeling (AM) in computer vision lags behind natural language processing (NLP) in self-supervised pre-training. This is mainly caused by the challenge that images are not sequential signals and lack a natural order when applying autoregressive modeling. In this study, inspired by human beings’ way of grasping an image, i.e., focusing on the main object first, we present a semantic-aware autoregressive image modeling (SemAIM) method to tackle this challenge. The key insight of SemAIM is to autoregressively model images from the semantic patches to the less semantic patches. To this end, we first calculate a semantic-aware permutation of patches according to their feature similarities and then perform the autoregression procedure based on the permutation. In addition, considering that the raw pixels of patches are low-level signals and are not ideal prediction targets for learning high-level semantic representation, we also explore utilizing the patch features as the prediction targets. Extensive experiments are conducted on a broad range of downstream tasks, including image classification, object detection, and instance/semantic segmentation, to evaluate the performance of SemAIM. The results demonstrate SemAIM achieves state-of-the-art performance compared with other self-supervised methods. Specifically, with ViT-B, SemAIM achieves 84.1% top-1 accuracy for fine-tuning on ImageNet, 51.3% AP and 45.4% AP for object detection and instance segmentation on COCO, which outperforms the vanilla MAE by 0.5%, 1.0%, and 0.5%, respectively. Code is available at https://github.com/skyoux/SemAIM.
Kaiyou Song
AAAI1
2023 Multi-Mode Online Knowledge Distillation for Self-Supervised Visual Representation Learning
abstract
Self-supervised learning (SSL) has made remarkable progress in visual representation learning. Some studies combine SSL with knowledge distillation (SSL-KD) to boost the representation learning performance of small models. In this study, we propose a Multi-mode Online Knowledge Distillation method (MOKD) to boost self-supervised visual representation learning. Different from existing SSL-KD methods that transfer knowledge from a static pre-trained teacher to a student, in MOKD, two different models learn collaboratively in a self-supervised manner. Specifically, MOKD consists of two distillation modes: self-distillation and cross-distillation modes. Among them, self-distillation performs self-supervised learning for each model independently, while cross-distillation realizes knowledge interaction between different models. In cross-distillation, a cross-attention feature search strategy is proposed to enhance the semantic feature alignment between different models. As a result, the two models can absorb knowledge from each other to boost their representation learning performance. Extensive experimental results on different backbones and datasets demonstrate that two heterogeneous models can benefit from MOKD and outperform their independently trained baseline. In addition, MOKD also outperforms existing SSL-KD methods for both the student and teacher models.
Kaiyou Song, Zimeng Luo
CVPR1
2023 Hard Patches Mining for Masked Image Modeling
abstract
Masked image modeling (MIM) has attracted much research attention due to its promising potential for learning scalable visual representations. In typical approaches, models usually focus on predicting specific contents of masked patches, and their performances are highly related to pre-defined mask strategies. Intuitively, this procedure can be considered as training a student (the model) on solving given problems (predict masked patches). However, we argue that the model should not only focus on solving given problems, but also stand in the shoes of a teacher to produce a more challenging problem by itself. To this end, we propose Hard Patches Mining (HPM), a brand-new framework for MIM pre-training. We observe that the reconstruction loss can naturally be the metric of the difficulty of the pretraining task. Therefore, we introduce an auxiliary loss predictor, predicting patch-wise losses first and deciding where to mask next. It adopts a relative relationship learning strategy to prevent overfitting to exact reconstruction loss values. Experiments under various settings demonstrate the effectiveness of HPM in constructing masked images. Furthermore, we empirically find that solely introducing the loss prediction objective leads to powerful representations, verifying the efficacy of the ability to be aware of where is hard to reconstruct.11Code: https://github.com/Haochen-wang409/HPM
Kaiyou Song, Junsong Fan, Yuxi Wang 0001, Zhaoxiang Zhang 0001
CVPR2
2023 Semantics-Consistent Feature Search for Self-Supervised Visual Representation Learning
abstract
In contrastive self-supervised learning, the common way to learn discriminative representation is to pull different augmented "views" of the same image closer while pushing all other images further apart, which has been proven to be effective. However, it is unavoidable to construct undesirable views containing different semantic concepts during the augmentation procedure. It would damage the semantic consistency of representation to pull these augmentations closer in the feature space indiscriminately. In this study, we introduce feature-level augmentation and propose a novel semantics-consistent feature search (SCFS) method to mitigate this negative effect. The main idea of SCFS is to adaptively search semantics-consistent features to enhance the contrast between semantics-consistent regions in different augmentations. Thus, the trained model can learn to focus on meaningful object regions, improving the semantic representation ability. Extensive experiments conducted on different datasets and tasks demonstrate that SCFS effectively improves the performance of self-supervised learning and achieves state-of-the-art performance on different downstream tasks.1
Kaiyou Song, Zimeng Luo
ICCV1
2023 DropPos: Pre-Training Vision Transformers by Reconstructing Dropped Positions
abstract
As it is empirically observed that Vision Transformers (ViTs) are quite insensitive to the order of input tokens, the need for an appropriate self-supervised pretext task that enhances the location awareness of ViTs is becoming evident. To address this, we present DropPos, a novel pretext task designed to reconstruct Dropped Positions. The formulation of DropPos is simple: we first drop a large random subset of positional embeddings and then the model classifies the actual position for each non-overlapping patch among all possible positions solely based on their visual appearance. To avoid trivial solutions, we increase the difficulty of this task by keeping only a subset of patches visible. Additionally, considering there may be different patches with similar visual appearances, we propose position smoothing and attentive reconstruction strategies to relax this classification problem, since it is not necessary to reconstruct their exact positions in these cases. Empirical evaluations of DropPos show strong capabilities. DropPos outperforms supervised pre-training and achieves competitive results compared with state-of-the-art self-supervised alternatives on a wide range of downstream benchmarks. This suggests that explicitly encouraging spatial reasoning abilities, as DropPos does, indeed contributes to the improved location awareness of ViTs. The code is publicly available at https://github.com/Haochen-Wang409/DropPos.
Junsong Fan, Yuxi Wang 0001, Kaiyou Song, Zhaoxiang Zhang 0001
NeurIPS4
2021 Multi-Scale Boosting Feature Encoding Network for Texture Recognition
abstract
Texture recognition remains a challenging visual task due to the complex appearance variations caused by scale changes in the real world. In most existing texture recognition methods, textures are represented at a single scale; thus, multi-scale texture information is not fully utilized, resulting in insufficient representation and inaccurate recognition. In this study, with the goal of addressing the challenge of scale changes, we propose a novel multi-scale boosting feature encoding network (MSBFEN) for accurate texture recognition. MSBFEN first extracts multi-scale features with multi-scale texture structure information under the guidance of texture priors using a novel prior-guided feature extraction (PFE) method. Then, a multi-scale texture encoding (MSTE) method is devised to capture discriminative multi-scale texture representations by encoding the extracted features. Finally, to fully utilize the multi-scale texture representations for accurate texture recognition, a novel multi-scale boosting learning (MSBL) method is proposed. In MSBL, the learning procedure for multi-scale texture recognition is boosted in a hierarchical, progressively reinforced manner, significantly addressing the challenge of scale changes and greatly enhancing the recognition accuracy. In addition, a novel outlier-aware texture encoding (OTE) method is proposed for robust texture encoding at each scale of MSTE. OTE can resist the influence of background interference and can further enhance the robustness of MSBFEN. In extensive experiments conducted on six challenging texture recognition datasets, namely, KTH-TIPS2b, FMD, DTD, MINC, GTOS and GTOS-mobile, MSBFEN achieves accuracies of 86.2%, 86.4%, 77.8%, 85.3%, 86.4% and 87.57%, respectively, representing state-of-the-art texture recognition performance.
Kaiyou Song, Hua Yang 0002, Zhou-Ping Yin
IEEE Trans. Circuits Syst. Video Technol.1
2021 An Anomaly Feature-Editing-Based Adversarial Network for Texture Defect Visual Inspection
abstract
Establishing a unified model for the defect inspection of different texture surfaces remains a challenge in the industrial automation field because these surfaces can vary in regular and irregular ways. Current unsupervised learning methods are trained on defect-free samples only and cannot directly address anomalies during testing, which precludes these methods from simultaneously inspecting for various texture defects. In this article, we propose a novel unsupervised anomaly feature-editing-based adversarial network (AFEAN) to accurately inspect various texture defects. To impart the AFEAN with the ability to address anomalies, a paired input, consisting of a defect-free image and an artificially defective image, is utilized for training. First, the AFEAN employs a feature extraction module (FEM) to extract latent features for the paired input. Subsequently, a novel anomaly feature detection module (AFDM) is proposed to detect anomaly features of the artificially defective image in the latent space. In the proposed AFDM, a novel central-constraint-based clustering method is proposed to detect anomaly features by learning the distribution of the latent features. Next, a novel global context feature editing module (GCFEM) is proposed to convert the detected anomaly features to normal features to suppress the reconstruction of defects. Finally, a feature decoding module (FDM) utilizes the edited features to reconstruct the texture background. Through the AFDM and GCFEM, the AFEAN achieves the ability to address anomaly features, effectively suppressing the reconstruction of defects on the texture background. In addition, to further improve the texture reconstruction accuracy, a pixel-level discrimination module (PDM) is employed to reconstruct texture details. In the testing phase, the defects are segmented by the residual image between the input image and the reconstructed texture background. The extensive experimental results demonstrate that the AFEAN achieves the state-of-the-art inspection accuracy.
Hua Yang 0002, Qinyuan Zhou, Kaiyou Song, Zhou-Ping Yin
IEEE Trans. Ind. Informatics3
2021 Weighted Feature Histogram of Multi-Scale Local Patch Using Multi-Bit Binary Descriptor for Face Recognition
abstract
Most face recognition methods employ single-bit binary descriptors for face representation. The information from these methods is lost in the process of quantization from real-valued descriptors to binary descriptors, which greatly limits their robustness for face recognition. In this study, we propose a novel weighted feature histogram (WFH) method of multi-scale local patches using multi-bit binary descriptors for face recognition. First, to obtain multi-scale information of the face image, the local patches are extracted using a multi-scale local patch generation (MSLPG) method. Second, with the goal of reducing the quantization information loss of binary descriptors, a novel multi-bit local binary descriptor learning (MBLBDL) method is proposed to extract multi-bit local binary descriptors (MBLBDs). In MBLBDL, a learned mapping matrix and novel multi-bit coding rules are employed to project pixel difference vectors (PDVs) into the MBLBDs in each local patch. Finally, a novel robust weight learning (RWL) method is proposed to learn a set of robust weights for each patch to integrate the MBLBDs into the final face representation. In RWL, a codebook is first constructed by clustering MBLBDs on each local patch to extract a feature histogram. Then, considering that different parts of the face have different degrees of robustness to local changes, a set of weights is learned to concatenate the feature histograms of all local patches into the final representation of a face image. In addition, to further improve the performance for heterogeneous face recognition, a coupled WFH (C-WFH) method is proposed. C-WFH maintains the similarity of the corresponding MBLBDs and feature histograms for a pair of heterogeneous face images by means of a novel coupled feature learning (CFL) method to reduce the modality gap. A series of experiments are conducted on widely used face datasets to analyze the performance of WFH and C-WFH. Extensive experimental results show that WFH and C-WFH outperform state-of-the-art face recognition methods.
Hua Yang 0002, Chenting Gong, Kaiji Huang, Kaiyou Song, Zhou-Ping Yin
IEEE Trans. Image Process.4
2019 Large-scale and rotation-invariant template matching using adaptive radial ring code histograms
Hua Yang 0002, Chenghui Huang, Feiyue Wang 0003, Kaiyou Song, Shijiao Zheng, Zhou-Ping Yin
Pattern Recognit.4
2019 Multiscale Feature-Clustering-Based Fully Convolutional Autoencoder for Fast Accurate Visual Inspection of Texture Surface Defects
abstract
Visual inspection of texture surface defects is still a challenging task in the industrial automation field due to the tremendous changes in the appearance of various surface textures. Current visual inspection methods cannot simultaneously and efficiently inspect various types of texture defects due to either the low discriminative capabilities of handcrafted features or their time-consuming sliding-window strategy. In this paper, we present a novel unsupervised multiscale feature-clustering-based fully convolutional autoencoder (MS-FCAE) method that efficiently and accurately inspects various types of texture defects based on a small number of defect-free texture samples. The proposed MS-FCAE method utilizes multiple FCAE subnetworks at different scale levels to reconstruct several textured background images. The residual images are obtained by subtracting these texture backgrounds from the input image individually; then, they are fused into one defect image. To maximize the efficiency, each FCAE subnetwork utilizes fully convolutional neural networks to extract the original feature maps directly from the input images. Meanwhile, each FCAE subnetwork performs feature clustering to improve the discriminant power of the encoded feature maps. The proposed MS-FCAE method is evaluated on several texture surface inspection data sets both qualitatively and quantitatively. This method achieves a Precision of 92.0% while requiring only 82 ms for input images of $1920\times 1080$ pixels. The extensive experimental results demonstrate that MS-FCAE achieves highly efficient and state-of-the-art inspection accuracy. Note to Practitioners-Most conventional visual inspection methods can address only one specific type of texture defect, while multiscale feature-clustering-based fully convolutional autoencoder (MS-FCAE) can simultaneously and accurately inspect various types of texture surface defects, such as those of thin-film transistor liquid crystal displays, wood, fabrics, and ceramic tiles. Furthermore, MS-FCAE requires only a small number of surface texture samples to learn a robust network model, and its training requires no defect samples. This is extremely important for industrial applications because identifying and labeling defect samples is difficult. Moreover, MS-FCAE can be applied to online visual inspection utilizing a graphics processing unit-based parallel processing strategy.
Hua Yang 0002, Kaiyou Song, Zhou-Ping Yin
IEEE Trans Autom. Sci. Eng.3
2019 Multi-Scale Attention Deep Neural Network for Fast Accurate Object Detection
abstract
Object detection remains a challenging task in computer vision due to the tremendous extent of changes in the appearances of objects caused by clustered backgrounds, occlusion, truncation, and scale change. Current deep neural network (DNN)-based object detection methods cannot simultaneously achieve a high accuracy and a high efficiency. To overcome this limitation, in this paper, we propose a novel multi-scale attention (MSA) DNN for accurate object detection with high efficiency. The proposed MSA-DNN method utilizes a novel multi-scale feature fusion module (MSFFM) to construct high-level semantic features. Subsequently, a novel MSA module (MSAM) based on the fused layers of the MSFFM is introduced to exploit the global semantic information of image-level labels to guide detection. On the one hand, MSAM can capture global semantic information to further enhance the semantic feature representation of the fused layers constructed by the MSFFM, thereby improving the detection accuracy. On the other hand, the MSA maps generated by MSAM can be employed to rapidly and coarsely locate objects at different scales. In addition, an attention-based hard negative mining strategy is introduced to filter out negative samples to reduce the search space, dramatically alleviating the severe class imbalance problem. Extensive experimental results on the challenging PASCAL VOC 2007, PASCAL VOC 2012, and MS COCO datasets demonstrate that MSA-DNN achieves a state-of-the-art detection accuracy while maintaining a high efficiency. Furthermore, MSA-DNN significantly improves the small-object detection accuracy.
Kaiyou Song, Hua Yang 0002, Zhou-Ping Yin
IEEE Trans. Circuits Syst. Video Technol.1
2019 Robust Semantic Template Matching Using a Superpixel Region Binary Descriptor
abstract
Almost all conventional template-matching methods employ low-level image features to measure the similarity between a template image and a scene image using similarity measures such as pixel intensity and pixel gradient. Although these methods have been widely used in many applications, they cannot simultaneously address all types of robustness challenges. In this study, with the goal of simultaneously addressing the various challenges, we present a robust semantic template-matching approach (RSTM). Inspired by the local binary descriptor, we propose a novel superpixel region binary descriptor (SRBD) to construct a multilevel semantic fusion feature vector for RSTM. SRBD uses a new kernel-distance-based simple linear iterative clustering (KD-SLIC) method to extract the stable superpixels from the template image; Then, based on the average intensity difference between each superpixel region and its neighbors, the dominant gradient orientation of each superpixel can be obtained, and the semantic features of each superpixel can be described as the dominant orientation difference vector, which is coded as the rotation-invariant SRBD. In the off-line matching phase, the fusion semantic feature vector of RSTM combines the multilevel SRBD features with different numbers of superpixels. In the online matching phase, to cope with rotation invariance, a marginal probability model is proposed and applied to locate the positions of template images in the scene image. Moreover, to accelerate computation, an image pyramid is employed. We conduct a series of experiments on a large dataset randomly selected from the MS COCO dataset to fully analyze the robustness of this approach. The experimental results show that RSTM simultaneously addresses rotation changes, scale changes, noise, occlusions, blur, nonlinear illumination changes and deformation with high time efficiency while also outperforming previous stateof- the-art template-matching methods.
Hua Yang 0002, Chenghui Huang, Feiyue Wang 0003, Kaiyou Song, Zhou-Ping Yin
IEEE Trans. Image Process.4
2018 An Accurate Mura Defect Vision Inspection Method Using Outlier-Prejudging-Based Image Background Construction and Region-Gradient-Based Level Set
abstract
The visual inspection of Mura defects is still a challenging task in the quality control of panel displays because of the intrinsically nonuniform brightness and blurry contours of these defects. The current methods cannot detect all Mura defect types simultaneously, especially small defects. In this paper, we introduce an accurate Mura defect visual inspection (AMVI) method for the fast simultaneous inspection of various Mura defect types. The method consists of two parts: an outlier-prejudging-based image background construction (OPBC) algorithm is proposed to quickly reduce the influence of image backgrounds with uneven brightness and to coarsely estimate the candidate regions of Mura defects. Then, a novel region-gradient-based level set (RGLS) algorithm is applied only to these candidate regions to quickly and accurately segment the contours of the Mura defects. To demonstrate the performance of AMVI, several experiments are conducted to compare AMVI with other popular visual inspection methods are conducted. The experimental results show that AMVI tends to achieve better inspection performance and can quickly and accurately inspect a greater number of Mura defect types, especially for small and large Mura defects with uneven backlight.
Hua Yang 0002, Kaiyou Song, Shuang Mei, Zhou-Ping Yin
IEEE Trans Autom. Sci. Eng.2