Peipei Kang

dblp:143/0426 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
17since 2021 · last 2025
0000-0001-8637-051XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Multimodal Pseudo-label Guided Semantic Enhanced Hashing Learning for Cross-modal Retrieval
Changhong Wu, Shaohua Teng, Zefeng Zheng, Wei Zhang 0005, Peipei Kang
PRCV (1)5
2025 Global and local semantic enhancement of samples for cross-modal hashing
Shaohua Teng, Zefeng Zheng, Wei Zhang 0005, Peipei Kang
Neurocomputing5
2024 Local residual preserving non-negative matrix factorization for multi-view clustering
Peipei Kang, Weijun Sun, Zhikun Jiang
Neurocomputing2
2024 SCH: Symmetric Consistent Hashing for cross-modal retrieval
Haomin Ni, Xiaozhao Fang, Peipei Kang, Hongbo Gao 0001, Guoxu Zhou, Shengli Xie 0001
Signal Process.3
2024 Kernel-Based Sparse Representation Learning With Global and Local Low-Rank Label Constraint
abstract
Due to the large-scale and multiscale natures of social media data, sparse representation (SR) learning methods are widely followed. However, there are three problems associated with the existing SR methods: 1) they neglect the fact that the semantic features of data may change during iterative learning, which leads to weak semantic learning; 2) they often assume that the data are linearly separable, while the data might be nonlinear in many real-world applications; and 3) they cannot ensure the low-rank and discriminative properties of the data at the same time and might neglect the global properties of the data, leading to suboptimal solutions. To solve these problems, we propose a novel method, named kernel-based SR learning with global and local low-rank label (KSR-GL3) constraint, which strengthens the semantic information and ensures the semantic features invariant during learning. First, we map the data into a high-dimensional feature space to learn the linear representation of samples. Second, global and local low-rank label (GL3) constraint is used to ensure the semantic invariance, low-rankness, and discrimination of features during learning. Third, an$\ell _{2,1}$is imposed to explore the sparseness of the subspace. Mathematical analyses show that GL3 can retain the intrinsic properties of data during learning. By combining the above three components, a generalized power iteration (GPI) approach is applied to build the model and deal with the tricky optimization problem. By KSR-GL 3, a sparse, low-rank, and discriminative subspace is produced from the high-dimensional and orthogonal representation of the data under the guidance of semantics, while the intrinsic properties of data are preserved. Extensive experiments on six datasets compared with five advanced algorithms demonstrate its promising prospects.
Luyao Teng, Feiyi Tang, Zefeng Zheng, Peipei Kang, Shaohua Teng
IEEE Trans. Comput. Soc. Syst.4
2024 Two-Step Strategy for Domain Adaptation Retrieval
abstract
Conventional hash-based retrieval method rely on the assumption that the query and database are of the identical domain. However, cross-domain problem often occurs in real-world applications, leading to the unsatisfactory performance of existing hashing methods. Recently, some researchers have put forward domain adaptation retrieval (DAR) under the perspective of domain adaptation (DA) and achieved promising results. But the following limitations still exist: 1) a single function is used to handle two challenges, i.e., domain adaptation and hashing, which is not flexible to explore enough underlying information for simultaneously accomplishing these challenges well; 2) non-dominant features in the sample are ignored; 3) the dissimilarity structure of dissimilar samples is not taken into account. To address the above problems, we propose a novel framework named two-step strategy (TSS) for domain adaptation retrieval, which advocates dividing DAR into two steps: DA step and hashing step. A DA function and a hash function are learned to handle the above two challenges, respectively, making the process more reasonable. Additionally, a discriminant semantic fusion loss is proposed to improve the discriminative ability among classes. Unlike other works that focus on discovering dominant features, we exploit the neglected non-dominant features and assign them attention with sinusoidal semantic embedding, actively creating a clear separation between classes. At last, we present an adaptive similarity preserving loss to preserve the similarity structure of the original data in all intra-domain and inter-domain hash codes. Extensive experiments on various datasets demonstrate that the proposed TSS achieves state-of-the-art performance.
Yonghao Chen, Xiaozhao Fang, Peipei Kang, Na Han, Shengli Xie 0001
IEEE Trans. Knowl. Data Eng.5
2024 Two-Stage Asymmetric Similarity Preserving Hashing for Cross-Modal Retrieval
abstract
Hashing-based techniques present appealing solutions for cross-modal retrieval due to its low storage requirements and excellent query efficiency. The majority of cross-modal hashing methods typically adopt equal-length encoding scheme to represent multimodal data and achieve cross-modal similarity search. However, such scheme can be regarded as a relatively strict limitation, because it sacrifices the flexible representation of multimodal data in reality and cannot always guarantee the optimal retrieval performance. To address the challenge, this paper focuses on encoding heterogeneous data with varying hash lengths. To achieve this purpose, we propose a flexible cross-modal hashing approach, named Two-stage Asymmetric Similarity Preserving Hashing, TASPH for short, which can be applied to both unequal-length and equal-length retrieval scenarios. Specifically, in the first stage, TASPH designs a novel discrete asymmetric strategy to learn the modality-specific hash codes with varying lengths, enabling a flexible representation of heterogeneous data. Simultaneously, TASPH utilizes two semantic transformation matrices to establish the semantic correlations between varying hash codes. Different from most of the existing approaches that employ relaxation solutions, TASPH satisfies the discrete constraints without any relaxation. In the second stage, the learned semantic transformation matrices are employed to alleviate cross-modal heterogeneity, which guarantees that TASPH can learn more powerful hash functions to improve the discriminative ability of hash codes. Abundant experiments conducted on three benchmark datasets demonstrate encouraging results compared with the state-of-the-art approaches under different retrieval scenarios.
Junfan Huang, Peipei Kang, Na Han, Yonghao Chen, Xiaozhao Fang, Hongbo Gao 0001, Guoxu Zhou
IEEE Trans. Knowl. Data Eng.2
2024 Dual Noise Elimination and Dynamic Label Correlation Guided Partial Multi-Label Learning
abstract
Partial multi-label learning (PML) needs to address the problem of multi-label learning when the dataset contains redundant information. PML is more challenging compared to traditional multi-label learning, because PML needs not only to perform the multi classification task, but also to reduce the impact of noise information on the model. Existing PML methods suffer from the following problems. (1) Only single source of noise is considered. (2) Some methods ignore the label correlations. To solve the above problems, we proposes a new dual noise elimination and dynamic label correlation guided partial multi-label learning (PML-DNDC). Specifically, the hidden ground-truth label matrix is decomposed into two compressed matrices of instance and classifier, which are used to approximate the candidate label matrix, to eliminate the negative effects of label noise on the model. On one hand, the compressed instance matrix maintains local structural consistency with the original instances, eliminating noise in the feature. On the other hand, dynamic label correlation guidance is designed to help classifier training by dynamically exploring the potential label correlations, which encourages relevant labels to obtain similar classifiers. After extensive experiments and analyses, we conclude that the proposed PML-DNDC is superior to the state-of-the-art methods.
Xiaozhao Fang, Peipei Kang, Yonghao Chen, Yuting Fang, Shengli Xie 0001
IEEE Trans. Multim.3
2024 Efficient Discriminative Hashing for Cross-Modal Retrieval
abstract
Hashing techniques have been extensively studied in cross-modal retrieval due to their advantages in high computational efficiency and low storage cost. However, existing methods unconsciously ignore the complementary information of multimodal data, thus failing to consider learning discriminative hash codes from the perspective of information complementarity while often involving time-consuming training overhead. To tackle the above issues, we propose an efficient discriminative hashing (EDH) with information complementarity consideration. Specifically, we reckon that multimodal features and their corresponding semantic labels describe heterogeneous data viewed from low-and high-level structures, which owns complementarity. To this end, low-level latent representation and high-level semantics representation are simply derived. Then, a joint learning strategy is formulated to simultaneously exploit the above two representations for generating discriminative hash codes, which is quite computationally efficient. Besides, EDH decomposes hash learning into two steps. To obtain powerful hash functions which are conductive to retrieval, a regularization term considering pairwise semantic similarity is introduced into hash functions learning. In addition, an efficient optimization algorithm is designed to solve the optimization problem in EDH. Extensive experiments conducted on benchmark datasets demonstrate the superiority of our EDH in terms of retrieval performance and training efficiency. The source code is available at https://github.com/hjf-hjf/EDH.
Junfan Huang, Peipei Kang, Xiaozhao Fang, Na Han, Shengli Xie 0001, Hongbo Gao 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2023 Partial multi-label learning: exploration of binary ground-truth labels
abstract
Partial multi-label learning (PML) aims to accurately predict multi-labels for unknown instances with noise labels in training set. Accurate identification of ground-truth labels is essential to optimize performance. The accuracy of ground-truth labels affects label correlation capture, which guides classifier learning. However, current methods often rely on approximation of ground-truth labels through intermediate variables, rather than restoring binary ground-truth labels directly. To solve this problem, we propose a partial multi-label learning algorithm by exploring binary ground-truth labels (PML-EBGL). First, the candidate label matrix is decomposed into a ground-truth label matrix and a noise matrix. Then, the noise matrix is constrained using the l1norm for sparsity, and the binary ground-truth label matrix is used to explore label correlation, which is transferred to classifier by using the Laplacian term. Finally, the rotation matrix is added after the classifier, and the instance is projected onto the binary label matrix. Extensive experiments show that PML-EBGL outperforms state-of-the-art methods.
Xiaozhao Fang, Weijun Lv, Peipei Kang
ICME4
2023 MKB: Multi-Kernel Bures Metric for Nighttime Aerial Tracking
Peipei Kang, Qintai Hu, Xiaozhao Fang
PRCV (10)2
2023 Cross-modal hashing with missing labels
Haomin Ni, Peipei Kang, Xiaozhao Fang, Weijun Sun, Shengli Xie 0001, Na Han
Neural Networks3
2023 Adaptive Visual Field Multi-scale Generative Adversarial Networks Image Inpainting Base on Coordinate-Attention
Peipei Kang, Xingcai Wu, Zhenguo Yang, Wenyin Liu
Neural Process. Lett.2
2022 Semantic-Adversarial Graph Convolutional Network for Zero-Shot Cross-Modal Retrieval
Lunke Fei, Peipei Kang, Xiaozhao Fang, Shaohua Teng
PRICAI (2)3
2022 Intra-class low-rank regularization for supervised and semi-supervised cross-modal retrieval
Peipei Kang, Zehang Lin, Zhenguo Yang, Xiaozhao Fang, Alexander M. Bronstein, Qing Li 0001, Wenyin Liu
Appl. Intell.1
2022 Deep fused two-step cross-modal hashing with multiple semantic supervision
Peipei Kang, Zehang Lin, Zhenguo Yang, Alexander M. Bronstein, Qing Li 0001, Wenyin Liu
Multim. Tools Appl.1
2021 Dense Fusion Network with Multimodal Residual for Sentiment Classification
abstract
In this paper, we propose a deep dense fusion network with multimodal residual (DFMR) to integrate multimodal information including language, acoustic speeches, and visual images for sentiment analysis. DFMR exploits a dense fusion (DF) block to fuse the multimodal features obtained by modality-specific sequence networks, which is achieved by modelling their unimodal, bimodal and trimodal interactions jointly. Instead of concatenating the multimodal features directly, DF block conducts fusion for any two paired modalities firstly, and the fused information will be integrated with the other modalities subsequently. Furthermore, DFMR stacks multiple DF blocks to capture high-level semantic information conveyed by the multimodal representations. In particular, DFMR adopts a multimodal residual (MR) block to integrate the modality-specific features and fused features in each DF blocks, to avoid forgetting the multi-aspect information and alleviate gradient vanishing during stacking. Extensive experiments conducted on four public benchmark datasets show that DFMR outperforms eleven state-of-the-art baselines.
Peipei Kang, Zhenguo Yang, Tianyong Hao, Qing Li 0001, Wenyin Liu
ICME2
2020 Active Transfer Learning
abstract
A major assumption in data mining and machine learning is that the training set and test set come from the same domain. They share the same feature space and have the same distribution. However, in many real-world applications, the training set and test set usually come from different domains. Thus, there might be negative similarities between different domains so that the negative transfer problem caused by negative similarity may happen. In this paper, we propose a novel method named active transfer learning (ATL) to solve the above problem. Specifically, the orthogonal projection matrix and the weight coefficient vector are introduced to extend maximum mean discrepancy (MMD) so that it can minimize MMD and simultaneously eliminate the negative transfer. To find the informative and discriminative subsets from the source domain, we then propose an information diversity term by using the local geometric structure information of the source samples. Besides, by using the label information of source samples, our method can guarantee the selected subsets as discriminative as possible. Finally, to efficiently implement the proposed method, an alternating optimization approach, which is based on the alternating direction method of multipliers (ADMM), is designed to solve the optimization problem. To demonstrate the effectiveness of the proposed ATL model, experiments are conducted on five real-world data sets. The experimental results show the superiority of our method over the state-of-the-art methods.
Zhihao Peng 0002, Wei Zhang 0005, Na Han, Xiaozhao Fang, Peipei Kang, Luyao Teng
IEEE Trans. Circuits Syst. Video Technol.5
2020 Learning Shared Semantic Space with Correlation Alignment for Cross-Modal Event Retrieval
abstract
In this article, we propose to learn shared semantic space with correlation alignment ( S 3 CA ) for multimodal data representations, which aligns nonlinear correlations of multimodal data distributions in deep neural networks designed for heterogeneous data. In the context of cross-modal (event) retrieval, we design a neural network with convolutional layers and fully connected layers to extract features for images, including images on Flickr-like social media. Simultaneously, we exploit a fully connected neural network to extract semantic features for text documents, including news articles from news media. In particular, nonlinear correlations of layer activations in the two neural networks are aligned with correlation alignment during the joint training of the networks. Furthermore, we project the multimodal data into a shared semantic space for cross-modal (event) retrieval, where the distances between heterogeneous data samples can be measured directly. In addition, we contribute a Wiki-Flickr Event dataset, where the multimodal data samples are not describing each other in pairs like the existing paired datasets, but all of them are describing semantic events. Extensive experiments conducted on both paired and unpaired datasets manifest the effectiveness of S 3 CA , outperforming the state-of-the-art methods.
Zhenguo Yang, Zehang Lin, Peipei Kang, Jianming Lv, Qing Li 0001, Wenyin Liu
ACM Trans. Multim. Comput. Commun. Appl.3
2019 Deep Semantic Space with Intra-class Low-rank Constraint for Cross-modal Retrieval
abstract
In this paper, a novel Deep Semantic Space learning model with Intra-class Low-rank constraint (DSSIL) is proposed for cross-modal retrieval, which is composed of two subnetworks for modality-specific representation learning, followed by projection layers for common space mapping. In particular, DSSIL takes into account semantic consistency to fuse the cross-modal data in a high-level common space, and constrains the common representation matrix within the same class to be low-rank, in order to induce the intra-class representations more relevant. More formally, two regularization terms are devised for the two aspects, which have been incorporated into the objective of DSSIL. To optimize the modality-specific subnetworks and the projection layers simultaneously by exploiting the gradient decent directly, we approximate the nonconvex low-rank constraint by minimizing a few smallest singular values of the intra-class matrix with theoretical analysis. Extensive experiments conducted on three public datasets demonstrate the competitive superiority of DSSIL for cross-modal retrieval compared with the state-of-the-art methods.
Peipei Kang, Zehang Lin, Zhenguo Yang, Xiaozhao Fang, Qing Li 0001, Wenyin Liu
ICMR1
2019 Catboost-based Framework with Additional User Information for Social Media Popularity Prediction
abstract
In this paper, a Catboost-based framework is proposed to predict social media popularity. The framework is constituted by two components: feature representation and Catboost training. In the component of feature representation, numerical features are directly used, while categorical features are converted into numerical features by a method of order target statistics in Catboost. Besides, some additional user information is also tracked to enrich the feature space. In the other component, Catboost is adopted as the regression model which is trained by using post-related, user-related and additional user information. Moreover, to make full use of the dataset for model training, a dataset augmentation strategy based on pseudo labels is proposed. This strategy involves in two-stage training. In the first stage, it trains a first-stage model that is used to label the test set as pseudo labeled. In the next stage, a final model is trained based on the new training set that includes original validation set and the pseudo labeled test set. The proposed method achieves the 2nd place in the leader board of the Grand Challenge of Social Media Prediction.
Peipei Kang, Zehang Lin, Shaohua Teng, Guipeng Zhang, Lingni Guo, Wei Zhang 0005
ACM Multimedia1
2019 Cross-domain Beauty Item Retrieval via Unsupervised Embedding Learning
abstract
Cross-domain image retrieval is always encountering insufficient labelled data in real world. In this paper, we propose unsupervised embedding learning (UEL) for cross-domain beauty and personal care product retrieval to finetune the convolutional neural network (CNN). More specifically, UEL utilizes the non-parametric softmax to train the CNN model as instance-level classification, which reduces the influence of some inevitable problems (e.g., shape variations). In order to obtain better performance, we integrate a few existing retrieval methods trained on different datasets. Furthermore, a query expansion strategy (i.e., diffusion) is adopted to improve the performance. Extensive experiments conducted on a dataset including half million images of beauty and personal product items (Perfect-500K) manifest the effectiveness of our proposed method. Our approach achieves the 2nd place in the leader board of the Grand Challenge of AI Meets Beauty in ACM Multimedia 2019. Our code is available at: https://github.com/RetrainIt/Perfect-Half-Million-Beauty-Product-Image-Recognition-Challenge-2019.
Zehang Lin, Haoran Xie 0001, Peipei Kang, Zhenguo Yang, Wenyin Liu, Qing Li 0001
ACM Multimedia3
2019 Unsupervised feature selection with adaptive residual preserving
Luyao Teng, Zhenye Feng, Xiaozhao Fang, Shaohua Teng, Hua Wang 0002, Peipei Kang, Yanchun Zhang
Neurocomputing6
2019 Improving cross-dimensional weighting pooling with multi-scale feature fusion for image retrieval
Qi Wang 0079, Jinxiang Lai, Zhenguo Yang, Kai Xu 0010, Peipei Kang, Wenyin Liu, Liang Lei
Neurocomputing5
2018 Random Forest Exploiting Post-related and User-related Features for Social Media Popularity Prediction
abstract
Social media headline prediction (SMHP) is a thriving application scenario, which aims to predict the popularity of the post data shared on social media. In this paper, we propose to use multi-aspect features combined with the random forest (RF) model for popularity predictions. Firstly, we extract features by combining both metadata of the posts and users' features. More specifically, we adopt the binary coding strategy for dimensionality reduction and deal with the missing values by using some strategies, i.e., estimating the missing geographic information according to the information of users, and filling the missing features with median. Furthermore, regression models can be used directly to make predictions. In particular, a random forest (RF) model is adopted since it does not require much effort in tuning hyper-parameters and performs effectively. Extensive experiments conducted on the SMHP dataset consisting of 340K image posts shared by 80K users manifest the effectiveness of our method. Our approach achieves the 4nd place in the leader board of the Grand Challenge of SMHP in ACM Multimedia 2018.
Feitao Huang, Zehang Lin, Peipei Kang, Zhenguo Yang
ACM Multimedia4