Feifei Kou

dblp:223/2313 · also Fei-Fei Kou · DBLP profile ↗
← Back
35ranked-venue papers
7as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DiMA: Distinguishing Resident and Tourist Preferences via Multi-Modal LLM Alignment for Out-of-Town Cross-Domain Recommendation
abstract
Out-of-Town (OOT) recommendation aims to provide personalized suggestions for users in unfamiliar cities. However, OOT recommendation faces two fundamental challenges: the difficulty of reasoning across modalities, as preference signals in disparate formats such as images and text are hard to compare; and the preference deviation problem, since a user's resident and tourist preferences often diverge, rendering simple preference transfer ineffective. To address these challenges, we propose Distinguishing Resident and Tourist Preferences via Multi-Modal LLM Alignment for Out-of-Town Cross-Domain Recommendation (DiMA), a framework for re-ranking Points of Interest (POIs). To tackle the multimodal challenge, DiMA first leverages Multimodal Large Language Models and Large Language Models (LLMs) to transform heterogeneous POI data into unified semantic tags, enabling both cross-modal reasoning and efficient downstream processing. To address preference deviation, a ``teacher'' LLM executes a custom Chain-of-Thought (CoT) process to disentangle resident and tourist preferences from multi-city histories for re-ranking. Finally, a lightweight student model learns this CoT reasoning via Supervised Fine-Tuning and is then refined with Direct Preference Optimization to align with true user choices, with the potential to surpass the teacher. Extensive experiments on a real-world dataset demonstrate that DiMA significantly enhances the performance of baseline models in the OOT recommendation re-ranking task.
Jinpeng Chen 0001, Huan Li 0003, Senzhang Wang, Feifei Kou, Ye Ji 0002, Kaimin Wei, Zhenye Yang
AAAI6
2026 Collaborative Transformers with Multi-Level Forensic Attention for Image Manipulation Localization
abstract
The proliferation of the tampered images on social media can pose serious societal risks, influencing public opinion and causing panic. Image Manipulation Localization technique has advanced to address this, but some methods focus on microscopic traces, overlooking macroscopic semantics that deceive viewers. To address this problem, we propose a novel Image Manipulation Localization framework called Collaborative Transformers (Co-Transformers), designed to fully explore and utilize the collaborative information between macroscopic semantics and microscopic traces. This framework is based on two Vision Transformer variants. The first variant captures the semantic logic of the image. The second variant delves into microscopic tampering traces. By dynamically fusing these two complementary features, the framework enables interaction between macroscopic semantic inconsistencies and microscopic abnormal traces, effectively coordinating their relationship in the latent space. Furthermore, we introduce a new Multi-Level Forensic Attention (MLF-Attention) mechanism to enhance the model's ability to extract various tampered traces, this mechanism can be integrated into our framework. Compared with existing methods, our proposed framework achieves state-of-the-art results in localization accuracy and shows good robustness against various attacks.
Jiwei Zhang 0007, Wenbo Feng, Feifei Kou, Shaozhang Niu
AAAI4
2026 Adaptive Graph Attention Based Discrete Hashing for Incomplete Cross-modal Retrieval
abstract
Cross-modal hashing has emerged as a pivotal solution for efficient retrieval across diverse modalities, such as images and texts, by mapping them into compact binary hash spaces. However, in real-world scenarios, the modalities data is often missing or misaligned. Existing methods are most rely on fully paired training data and ignore missing or misaligned modalities data, resulting in the semantic inconsistencies. To address these challenges, we propose an Adaptive Graph Attention-Based Discrete Hashing (AGADH) method, which consists of three parts. First, to solve the problem of missing modalities, AGADH employs a masked completion strategy to reconstruct missing modalities. Second, to mitigate semantic misalignment, AGADH leverages a Graph Attention Network (GAT) encoder-decoder architecture with alignment module to construct features from different modalities. Additionally, to enhance the fusion performance, an adaptive fusion module dynamically adjusting the contributions of image and text modalities with learnable weighting coefficients is proposed. Extensive experiments on three benchmark datasets, MS-COCO, NUS-WIDE, and MIRFlickr-25K, demonstrating that AGADH outperforms state-of-the-art methods in both fully paired and incompletely paired scenarios, showing its robustness and effectiveness in cross-modal retrieval tasks.
Shuang Zhang 0009, Lei Shi 0030, Huilong Jin, Feifei Kou, Pengfei Zhang 0010, Mingying Xu, Pengtao Lv
AAAI5
2026 MusicRec: Multi-modal Semantic-Enhanced Identifier with Collaborative Signals for Generative Recommendation
abstract
Generative recommendation as a new paradigm is influencing the current development of recommender systems. It aims to assign identifiers that capture richer semantic and collaborative information to items, and subsequently predict item identifiers via autoregressive generation using Large Language Models (LLMs). Existing approaches primarily tokenize item text into codebooks with preserved semantic IDs through RQ-VAE, or separately tokenize different modality features of items. However, existing tokenization methods face two major challenges: (1) Learning decoupled multi-modal features limits the quality of the semantic representation. (2) Ignoring collaborative signals from interaction history limits the comprehensiveness of identifiers. To address these limitations, we propose a multi-modal semantic-enhanced identifier with collaborative signals for generative recommendation, named MusicRec. In MusicRec, we propose a tokenization approach based on shared-specific modal fusion, enabling the generated identifiers to preserve semantic information more comprehensively from all modalities. In addition, we incorporate collaborative signals from user interactions to guide identifier generation, preserving collaborative patterns in the semantic representation space. Extensive experiments on three public datasets demonstrate that MusicRec achieves state-of-the-art performance compared to existing baseline methods.
Yuqiu Zhao, Lei Shi 0030, Yan Zhong 0001, Feifei Kou, Pengfei Zhang 0010, Jiwei Zhang 0007, Mingying Xu
AAAI4
2026 Dual-perspective hypergraph learning network for multimodal entity and relation extraction
Jie Liu 0022, Mingying Xu, Baowen Wu, Linqi Song, Yinqiao Li, Lei Shi 0030, Feifei Kou
Expert Syst. Appl.8
2026 Wavelet transform-based versatile watermarking for facial manipulation source tracing and detection
Yibo Zhang 0002, Weiguo Lin, Lei Shi 0030, Wanshan Xu, Yikun Xu, Feifei Kou
Inf. Process. Manag.7
2026 Enhancing Explainable Sequential Recommendation With Disentangled Representations and Auxiliary Review Explanations
Jinpeng Chen 0001, Huachen Guan, Hongbo Gao 0001, Huan Li 0003, Zhenye Yang, Kaimin Wei, Feifei Kou, Xindong Wu 0001
IEEE Trans. Comput. Soc. Syst.8
2026 Dual Graph Network Hashing for Cross-Modal Retrieval
Shuang Zhang 0009, Lei Shi 0030, Feifei Kou, Huilong Jin, Pengfei Zhang 0010, Weiping Ding 0001, Mingying Xu, Muhammet Deveci
IEEE Trans. Knowl. Data Eng.4
2025 Leveraging the Dual Capabilities of LLM: LLM-Enhanced Text Mapping Model for Personality Detection
abstract
Personality detection aims to deduce a user’s personality from their published posts. The goal of this task is to map posts to specific personality types. Existing methods encode post information to obtain user vectors, which are then mapped to personality labels. However, existing methods face two main issues: first, only using small models makes it hard to accurately extract semantic features from multiple long documents. Second, the relationship between user vectors and personality labels is not fully considered. To address the issue of poor user representation, we utilize the text embedding capabilities of LLM. To solve the problem of insufficient consideration of the relationship between user vectors and personality labels, we leverage the text generation capabilities of LLM. Therefore, we propose the LLM-Enhanced Text Mapping Model (ETM) for Personality Detection. The model applies LLM’s text embedding capability to enhance user vector representations. Additionally, it uses LLM’s text generation capability to create multi-perspective interpretations of the labels, which are then used within a contrastive learning framework to strengthen the mapping of these vectors to personality labels. Experimental results show that our model achieves state-of-the-art performance on benchmark datasets.
Weihong Bi, Feifei Kou, Lei Shi 0030, Yawen Li 0001, Hai-Sheng Li 0002, Jinpeng Chen 0001, Mingying Xu
AAAI2
2025 IWRN: A Robust Blind Watermarking Method for Artwork Image Copyright Protection Against Noise Attack
abstract
Adding imperceptible watermarks to artwork images, such as paintings and photographs, can effectively safeguard the copyright of these images without compromising their usability. However, existing blind watermarking techniques encounter two major challenges in addressing this task: imperceptibility and robustness, particularly when subjected to various noise attacks. In this paper, we propose a blind watermarking method for artwork image copyright protection, IWRN, which can ensure both the Imperceptibility of the Watermark and Robustness against Noise attacks. For imperceptibility, we design a Learnable Wavelet Network (LWN) to adaptively embed the watermark into the high-frequency region where the watermark has better invisibility. For robustness, we establish a Deform-Attention based Invertible Neural Network (DA-INN) with a decoding optimization, which offers the advantage of computational reversion, and combines the deform-attention mechanism and decoding optimization to enhance the model's resistance against noises. Additionally, we design a Joint Contrast Learning (JCL) mechanism to improve imperceptibility and robustness simultaneously. Experiments show that our IWRN outperforms other state-of-the-art blind watermarking methods, achieves an average performance of 41.55 PSNR and 99.57% accuracy on the Coco2017, Wikiart, and Div2k datasets when facing 12 kinds of noise attacks.
Feifei Kou, Yuhan Yao 0001, Siyuan Yao, Lei Shi 0030, Yawen Li 0001, Xuejing Kang
AAAI1
2025 StrucFormer: Structural Prior Guided Transformer for Mobile Crowdsensing Data Inference
abstract
The inherent constraint of the "human-in-the-loop" sensing mechanism, imposes mobile crowdsensing with high dynamics and uncertainty, ultimately leading to the issue of incomplete data collection. Current data inference solutions in mobile crowdsensing can be broadly categorized as low-rank models and deep learning models. Low-rank models apply structural prior for data inference, but have limited model capacity, while deep learning models possess salient feature expressivity, but are prone to overfitting in sparse crowdsensing scenarios. In this paper, we try to absorb the strengths of both two paradigms, and propose a structural prior guided Transformer, StrucFormer, for crowd-sensing data inference. Specifically, we exploit structural prior of low-rankness to power canonical Transformer from the aspects of input embedding, attention forming and model regularization, which enables the model to precisely capture the spatiotemporal and multi-type data correlations for accurate inference with only sparse observations. Extensive empirical results demonstrate the superiority of StrucFormer in terms of accuracy and generality in heterogeneous urban sensing tasks. The code is available at: https://github.com/CUPK-K/StrucFormer.
Xu Kang 0001, Shouceng Tian, Feifei Kou, Lei Shi 0030, Jiadong Ren
ICASSP3
2025 CFPT: Empowering Time Series Forecasting through Cross-Frequency Interaction and Periodic-Aware Timestamp Modeling
abstract
Long-term time series forecasting has been widely studied, yet two aspects remain insufficiently explored: the interaction learning between different frequency components and the exploitation of periodic characteristics inherent in timestamps. To address the above issues, we propose CFPT, a novel method that empowering time series forecasting through Cross-Frequency Interaction (CFI) and Periodic-Aware Timestamp Modeling (PTM). To learn cross-frequency interactions, we design the CFI branch to process signals in frequency domain and captures their interactions through a feature fusion mechanism. Furthermore, to enhance prediction performance by leveraging timestamp periodicity, we develop the PTM branch which transforms timestamp sequences into 2D periodic tensors and utilizes 2D convolution to capture both intra-period dependencies and inter-period correlations of time series based on timestamp patterns. Extensive experiments on multiple real-world benchmarks demonstrate that CFPT achieves state-of-the-art performance in long-term forecasting tasks. The code is publicly available at this repository: https://github.com/BUPT-SN/CFPT.
Feifei Kou, Lei Shi 0030, Yuhan Yao 0001, Yawen Li 0001, Suguo Zhu, Zhongbao Zhang, Junping Du 0001
ICML1
2025 EVICheck: Evidence-Driven Independent Reasoning and Combined Verification Method for Fact-Checking
abstract
Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) have demonstrated significant potential in automated fact-checking. However, existing methods face limitations in insufficient evidence utilization and lack of explicit verification criteria. Specifically, these approaches aggregate evidence for collective reasoning without independently analyzing each piece, hindering their ability to leverage the available information thoroughly. Additionally, they rely on simple prompts or few-shot learning for verification, which makes truthfulness judgments less reliable, especially for complex claims. To address these limitations, we propose a novel method to enhance evidence utilization and introduce explicit verification criteria, named EVICheck. Our approach independently reasons each evidence piece and synthesizes the results to enable more thorough exploration and enhance interpretability. Additionally, by incorporating fine-grained truthfulness criteria, we make the model's verification process more structured and reliable, especially when handling complex claims. Experimental results on the public RAWFC dataset demonstrate that EVICheck achieves state-of-the-art performance across all evaluation metrics. Our method demonstrates strong potential in fake news verification, significantly improving the accuracy.
Lei Shi 0030, Feifei Kou, Ligu Zhu, Chen Ma 0003, Pengfei Zhang 0010, Mingying Xu
IJCAI3
2025 Leveraging Multimodal Data and Side Users for Diffusion Cross-Domain Recommendation
abstract
Cross-domain recommendation (CDR) aims to address the persistent cold-start problem in Recommender Systems. Current CDR research concentrates on transferring cold-start users' information from the auxiliary domain to the target domain. However, these systems face two main issues: the underutilization of multimodal data, which hinders effective cross-domain alignment, and the neglect of side users who interact solely within the target domain, leading to inadequate learning of the target domain's vector space distribution. To address these issues, we propose a model leveraging Multimodal data and Side users for diffusion Cross-domain recommendation (MuSiC). We first employ a multimodal large language model to extract item multimodal features and leverage a large language model to uncover user features. Secondly, we propose the cross-domain diffusion module to learn the generation of feature vectors in the target domain. This approach involves learning feature distribution from side users and understanding the patterns in cross-domain transformation through overlapping users. Subsequently, the trained diffusion module is used to generate feature vectors for cold-start users in the target domain, enabling the completion of cross-domain recommendation tasks. Finally, our experimental evaluation of the Amazon dataset confirms that MuSiC achieves state-of-the-art performance, significantly outperforming all selected baselines. Our code is available: https://github.com/zhangf16/MuSiC.
Jinpeng Chen 0001, Huan Li 0003, Senzhang Wang, Yuan Cao 0003, Kaimin Wei, Jianxiang He, Feifei Kou, Jinqing Wang
ACM Multimedia8
2025 OSTAR: Optimized Statistical Text-classifier with Adversarial Resistance
abstract
The advancements in generative models and the real-world attack of machine-generated text(MGT) create a demand for more robust detection methods. The existing MGT detection methods for adversarial environments primarily consist of manually designed statistical-based methods and fine-tuned classifier-based approaches. Statistical-based methods extract intrinsic features but suffer from rigid decision boundaries vulnerable to adaptive attacks, while fine-tuned classifiers achieve outstanding performance at the cost of overfitting to superficial textual feature. We argue that the key to detection in current adversarial environments lies in how to extract intrinsic invariant features and ensure that the classifier possesses dynamic adaptability. In that case, we propose OSTAR, a novel MGT detection framework designed for adversarial environments which composed of a statistical enhanced classifier and a Multi-Faceted Contrastive Learning(MFCL). In the classifier aspect, our Multi-Dimensional Statistical Profiling (MDSP) module extracts intrinsic difference between human and machine texts, complementing classifiers with useful stable features. In the model optimization aspect, the MFCL strategy enhances robustness by contrasting feature variations before and after text attacks, jointly optimizing statistical feature mapping and baseline pre-trained models. Experimental results on three public datasets under various adversarial scenarios demonstrate that our framework outperforms existing MGT detection methods, achieving state-of-the-art performance and robust against attacks.The code is available at https://github.com/BUPT-SN/OSTAR.
Yuhan Yao 0001, Feifei Kou, Lei Shi 0030, Zhongbao Zhang, Suguo Zhu, Jiwei Zhang 0007, Lirong Qiu, Hai-Sheng Li 0002
NeurIPS2
2025 Dynamic Masking and Auxiliary Hash Learning for Enhanced Cross-Modal Retrieval
abstract
The demand for multimodal data processing drives the development of information technology. Cross-modal hash retrieval has attracted much attention because it can overcome modal differences and achieve efficient retrieval, and has shown great application potential in many practical scenarios. Existing cross-modal hashing methods have difficulties in fully capturing the semantic information of different modal data, which leads to a significant semantic gap between modalities. Moreover, these methods often ignore the importance differences of channels, and due to the limitation of a single goal, the matching effect between hash codes is also affected to a certain extent, thus facing many challenges. To address these issues, we propose a Dynamic Masking and Auxiliary Hash Learning (AHLR) method for enhanced cross-modal retrieval. By jointly leveraging the dynamic masking and auxiliary hash learning mechanisms, our approach effectively resolves the problems of channel information imbalance and insufficient key information capture, thereby significantly improving the retrieval accuracy. Specifically, we introduce a dynamic masking mechanism that automatically screens and weights the key information in images and texts during the training process, enhancing the accuracy of feature matching. We further construct an auxiliary hash layer to adaptively balance the weights of features across each channel, compensating for the deficiencies of traditional methods in key information capture and channel processing. In addition, we design a contrastive loss function to optimize the generation of hash codes and enhance their discriminative power, further improving the performance of cross-modal retrieval. Comprehensive experimental results on NUS-WIDE, MIRFlickr-25K and MS-COCO benchmark datasets show that the proposed AHLR algorithm outperforms several existing algorithms.
Shuang Zhang 0009, Lei Shi 0030, Feifei Kou, Huilong Jin, Pengfei Zhang 0010, Meiyu Liang, Mingying Xu
NeurIPS5
2025 RFCSC: Communication efficient reinforcement federated learning with dynamic client selection and adaptive gradient compression
Zhenhui Pan, Yawen Li 0001, Zeli Guan, Meiyu Liang, Ang Li 0015, Jia Wang 0011, Feifei Kou
Neurocomputing7
2025 A language-guided cross-modal semantic fusion retrieval method
Ligu Zhu, Suping Wang, Lei Shi 0030, Feifei Kou, Pengpeng Zhou
Signal Process.5
2025 DualFocus GAN for Robust Watermarking in Transportation Cyber-Physical Systems
abstract
With the advancement of Transportation Cyber-Physical Systems (TCPS), information security has become increasingly critical. Invisible watermarking, which ensures reliable information traceability without compromising carrier quality, holds significant potential for TCPS. However, achieving high robustness in the real world while maintaining imperceptibility is a challenge. To address this, we propose DGWW (Dual-discriminator GAN-based WaveFusion Watermarking), a novel invisible watermarking method that balances robustness and imperceptibility. The GAN-based approach is well suited for TCPS, as it enables adaptive watermark embedding aligned with the dynamic and heterogeneous nature of transportation data, effectively handling diverse noise conditions and data types. DGWW integrates a WaveFusion Encoding Module, a Dual-Focus Discriminator, and a contrastive learning-based optimization strategy to enhance watermark embedding without degrading robustness. These components leverage multi-frequency information, assess local and global impacts on image quality, and guide model optimization. Experimental results show that DGWW outperforms state-of-the-art methods in visual quality and robustness under various noise conditions, offering a robust and scalable solution for image watermarking in TCPS environments. By maintaining data usability and strong resistance to noise attacks, DGWW advances digital watermarking in intelligent transportation systems.
Feifei Kou, Yuhan Yao 0001, Jideng Han, Hai-Sheng Li 0002, Jiwei Zhang 0007
IEEE Trans. Intell. Transp. Syst.1
2025 Potential Features Fusion Network for Multimodal Fake News Detection
abstract
With the popularization of social networks, fake news is also widely and rapidly spreading, which poses a great threat to the Internet. Therefore, how to detect fake news automatically and efficiently has become an urgent problem to be solved. However, the existing approaches mostly focus on the explicit features (images and text) and deep fusions, without considering potential features such as text emotion and image category. To find a solution to this issue, we propose a Potential Features Fusion Network (PFFN), which models the explicit and potential features at the same time. To exploit the potential image features, we introduce a mixture of experts structure to process the news image separately, which can best use the relationships between the news image category and fake news detection. Besides, we also extract emotion features as potential text features and fuse them with explicit text features. Finally, we establish an attention-based feature fusion network to fuse the potential features with the explicit features, which can obtain a multimodal fusion feature of a piece of news and thus further improve the performance. We make experiments on four public datasets (Weibo16, Weibo19, Twitter, and PolitiFact); the results compared with the baseline approaches demonstrate that our PFFN has a better performance. Our code is available at https://github.com/Wang-bupt/PFFN
Feifei Kou, Bingwei Wang, Hai-Sheng Li 0002, Chuangying Zhu, Lei Shi 0030, Jiwei Zhang 0007, Limei Qi
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Self-derived Knowledge Graph Contrastive Learning for Recommendation
abstract
Knowledge Graphs (KGs) serve as valuable auxiliary information to improve the accuracy of recommendation systems. Previous methods have leveraged the knowledge graph to enhance item representation and thus achieve excellent performance. However, these approaches heavily rely on high-quality knowledge graphs and learn enhanced representations with the assistance of carefully designed triplets. Furthermore, the emergence of knowledge graphs has led to models that ignore the inherent relationships between items and entities. To address these challenges, we propose a Self-Derived Knowledge Graph Contrastive Learning framework (CL-SDKG) to enhance recommendation systems. Specifically, we employ the variational graph reconstruction technique to estimate the Gaussian distribution of user-item nodes corresponding to the graph neural network aggregation layer. This process generates multiple KGs, referred to as self-derived KGs. The self-derived KG acquires more robust perceptual representations through the consistency of the estimated structure. Besides, the self-derived KG allows models to focus on user-item interactions and reduce the negative impact of miscellaneous dependencies introduced by conventional KGs. Finally, we apply contrastive learning to the self-derived KG to further improve the robustness of CL-SDKG through the traditional KG contrast-enhanced process. We conducted comprehensive experiments on three public datasets, and the results demonstrate that our CL-SDKG outperforms state-of-the-art baselines.
Lei Shi 0030, Pengtao Lv, Feifei Kou, Jia Luo 0001, Mingying Xu
ACM Multimedia5
2024 An End-To-End Graph Attention Network Hashing for Cross-Modal Retrieval
abstract
Due to its low storage cost and fast search speed, cross-modal retrieval based on hashing has attracted widespread attention and is widely used in real-world applications of social media search. However, most existing hashing methods are often limited by uncomprehensive feature representations and semantic associations, which greatly restricts their performance and applicability in practical applications. To deal with this challenge, in this paper, we propose an end-to-end graph attention network hashing (EGATH) for cross-modal retrieval, which can not only capture direct semantic associations between images and texts but also match semantic content between different modalities. We adopt the contrastive language image pretraining (CLIP) combined with the Transformer to improve understanding and generalization ability in semantic consistency across different data modalities. The classifier based on graph attention network is applied to obtain predicted labels to enhance cross-modal feature representation. We construct hash codes using an optimization strategy and loss function to preserve the semantic information and compactness of the hash code. Comprehensive experiments on the NUS-WIDE, MIRFlickr25K, and MS-COCO benchmark datasets show that our EGATH significantly outperforms against several state-of-the-art methods.
Huilong Jin, Lei Shi 0030, Shuang Zhang 0009, Feifei Kou, Chuangying Zhu, Jia Luo 0001
NeurIPS5
2023 Video Super-Resolution Reconstruction Based on Deep Learning and Spatio-Temporal Feature Self-similarity (Extended abstract)
abstract
Video super-resolution (SR) reconstruction technology aims at obtaining high quality reconstruction of high-resolution (HR) video sequences by inferring the lost detailed information from their low-resolution (LR) counterparts. However, this technology is an ill-posed problem because significant detailed information is lost in the process of video degrading. The existing learning-based SR reconstruction methods can be adapted to a larger super-resolution factor, but it cannot be guaranteed that any low-resolution image block can find its corresponding high-resolution block matching in a limited-scale training set. Some noise and over smooth phenomenon usually exist while dealing with some unique features that rarely appear in a given training data set. The self-similarity based SR methods do not rely on accurate sub-pixel motion estimation and thus can be adapted to complex motion patterns. However, under conditions of insufficient internal similar blocks, some visual flaws are usually produced due to the mismatched internal instances.
Meiyu Liang, Junping Du 0001, Zhe Xue, Xiaoxiao Wang 0006, Feifei Kou
ICDE6
2022 Few-shot node classification via local adaptive discriminant structure learning
Zhe Xue, Junping Du 0001, Xiangbin Liu, Junfu Wang, Feifei Kou
Frontiers Comput. Sci.6
2022 A scientific research topic trend prediction model based on multi-LSTM and graph convolutional network
abstract
Predicting the development trend of future scientific research not only provides a reference for researchers to understand the development of the discipline, but also provides support for decision-making and fund allocation for decision-makers. The continuous growth of scientific publications has brought challenges to track the development trends of scientific research topics. The existing topic trend prediction methods have proved that the research topic trend of a publication is influenced by other peer publications. However, they ignore the fact that the research topics of different publications belong to different research topic space. Moreover, the existing topic prediction methods do not fully consider the interactive influence among publications that the research topic of one publication affects the topics of other publications, it is also influenced by the research topics of other publications. In line with this, this paper proposes a scientific research topic trend prediction model based on multi-long short-term memory (multi-LSTM) and Graph Convolutional Network. Specifically, multiple LSTMs are employed to map research topics of different publications into their respective topic space. Then, the graph convolutional neural network is applied to learn the scientific influence context of each publication, so that the research topic of each publication not only integrates the influence of neighbor nodes, but also considers the influence of the neighbors of the neighbor node on the research topic of the publication, so as to more accurately fuse scientific influence context of research topic of peer publications. Experiments results on the data set of scientific research papers in the field of artificial intelligence and data mining demonstrate that the model improves the prediction precision and achieves the state-of-the-art research topic trend prediction effect compared with the other baseline models.
Mingying Xu, Junping Du 0001, Zhe Xue, Zeli Guan, Feifei Kou, Lei Shi 0030
Int. J. Intell. Syst.5
2022 Video Super-Resolution Reconstruction Based on Deep Learning and Spatio-Temporal Feature Self-Similarity
abstract
To address the problems in the existing video super-resolution methods, such as noise, over smooth and visual artifacts, which are caused by the reliance on limited external training or mismatch of internal similarity patch instances, this study proposes a novel video super-resolution reconstruction algorithm based on deep learning and spatio-temporal feature similarity (DLSS-VSR). The video super-resolution reconstruction mechanism with the joint internal and external constraints is established utilizing the complementary advantages of both external deep correlation mapping learning and internal spatio-temporal nonlocal self-similarity prior constraint. A deep learning model based on deep convolutional neural network is constructed to learn the nonlinear correlation mapping between low-resolution and high-resolution video frame patches. A novel spatio-temporal feature similarity calculation method is proposed, which considers both internal video spatio-temporal self-similarity and external clean nonlocal similarity. For the internal spatio-temporal feature self-similarity, we improve the accuracy and robustness of similarity matching by proposing a similarity measure strategy based on spatio-temporal moment feature similarity and structural similarity. The external nonlocal similarity prior constraint is learned by the patch group-based Gaussian mixture model. The time efficiency for spatio-temporal similarity matching is further improved based on saliency detection and region correlation judgment strategy, which achieves a better tradeoff between super-resolution accuracy and speed. Experimental results demonstrate that the DLSS-VSR algorithm achieves competitive super-resolution quality compared to other state-of-the-art algorithms in both subjective and objective evaluations.
Meiyu Liang, Junping Du 0001, Zhe Xue, Xiaoxiao Wang 0006, Feifei Kou
IEEE Trans. Knowl. Data Eng.6
2021 A semi-supervised semantic-enhanced framework for scientific literature retrieval
Mingying Xu, Junping Du 0001, Zhe Xue, Feifei Kou
Neurocomputing4
2021 MVGAN: Multi-View Graph Attention Network for Social Event Detection
abstract
Social networks are critical sources for event detection thanks to the characteristics of publicity and dissemination. Unfortunately, the randomness and semantic sparsity of the social network text bring significant challenges to the event detection task. In addition to text, time is another vital element in reflecting events since events are often followed for a while. Therefore, in this article, we propose a novel method named Multi-View Graph Attention Network (MVGAN) for event detection in social networks. It enriches event semantics through both neighbor aggregation and multi-view fusion in a heterogeneous social event graph. Specifically, we first construct a heterogeneous graph by adding the hashtag to associate the isolated short texts and describe events comprehensively. Then, we learn view-specific representations of events through graph convolutional networks from the perspectives of text semantics and time distribution, respectively. Finally, we design a hashtag-based multi-view graph attention mechanism to capture the intrinsic interaction across different views and integrate the feature representations to discover events. Extensive experiments on public benchmark datasets demonstrate that MVGAN performs favorably against many state-of-the-art social network event detection algorithms. It also proves that more meaningful signals can contribute to improving the event detection effect in social networks, such as published time and hashtags.
Wan-Qiu Cui, Junping Du 0001, Dawei Wang 0009, Feifei Kou, Zhe Xue
ACM Trans. Intell. Syst. Technol.4
2020 Cross-Media Semantic Correlation Learning Based on Deep Hash Network and Semantic Expansion for Social Network Cross-Media Search
abstract
Cross-media search from large-scale social network big data has become increasingly valuable in our daily life because it can support querying different data modalities. Deep hash networks have shown high potential in achieving efficient and effective cross-media search performance. However, due to the fact that social network data often exhibit text sparsity, diversity, and noise characteristics, the search performance of existing methods often degrades when dealing with this data. In order to address this problem, this article proposes a novel end-to-end cross-media semantic correlation learning model based on a deep hash network and semantic expansion for social network cross-media search (DHNS). The approach combines deep network feature learning and hash-code quantization learning for multimodal data into a unified optimization architecture, which successfully preserves both intramedia similarity and intermedia correlation, by minimizing both cross-media correlation loss and binary hash quantization loss. In addition, our approach realizes semantic relationship expansion by constructing the image-word relation graph and mining the potential semantic relationship between images and words, and obtaining the semantic embedding based on both internal graph deep walk and an external knowledge base. Experimental results demonstrate that DHNS yields better cross-media search performance on standard benchmarks.
Meiyu Liang, Junping Du 0001, Cong-Xian Yang, Zhe Xue, Hai-Sheng Li 0002, Feifei Kou, Yue Geng
IEEE Trans. Neural Networks Learn. Syst.6
2019 Interaction-Aware Arrangement for Event-Based Social Networks
abstract
The last decade has witnessed the emergence and popularity of event-based social networks (EBSNs), which extend online social networks to the physical world. Fundamental on EBSN platforms is to appropriately assign EBSN users to events they are interested to attend, known as event-participant arrangement. Previous event-participant arrangement studies either fail to avoid conflicts among events or ignore the social interactions among participants. In this work, we propose a new event-participant arrangement problem called Interaction-aware Global Event-Participant Arrangement (IGEPA). It globally optimizes arrangements between events and participants to avoid conflicts in events, and not only accounts for user interests, but also encourages socially active participants to join. To solve the IGEPA problem, we design an approximation algorithm which has an approximation ratio of at least 1\4. Experimental results validate the effectiveness of our solution.
Feifei Kou, Zimu Zhou, Junping Du 0001, Yexuan Shi, Pan Xu 0001
ICDE1
2019 A multi-feature probabilistic graphical model for social network semantic search
Feifei Kou, Junping Du 0001, Cong-Xian Yang, Yan-Song Shi, Meiyu Liang, Zhe Xue, Hai-Sheng Li 0002
Neurocomputing1
2019 Dynamic topic modeling via self-aggregation for short text streams
Lei Shi 0030, Junping Du 0001, Meiyu Liang, Feifei Kou
Peer-to-Peer Netw. Appl.4
2019 Short Text Analysis Based on Dual Semantic Extension and Deep Hashing in Microblog
abstract
Short text analysis is a challenging task as far as the sparsity and limitation of semantics. The semantic extension approach learns the meaning of a short text by introducing external knowledge. However, for the randomness of short text descriptions in microblogs, traditional extension methods cannot accurately mine the semantics suitable for the microblog theme. Therefore, we use the prominent and refined hashtag information in microblogs as well as complex social relationships to provide implicit guidance for semantic extension of short text. Specifically, we design a deep hash model based on social and conceptual semantic extension, which consists of dual semantic extension and deep hashing representation. In the extension method, the short text is first conceptualized to achieve the construction of hashtag graph under conceptual space. Then, the associated hashtags are generated by correlation calculation based on the integration of social relationships and concepts to extend the short text. In the deep hash model, we use the semantic hashing model to encode the abundant semantic features and form a compact and meaningful binary encoding. Finally, extensive experiments demonstrate that our method can learn and represent the short texts well by using more meaningful semantic signal. It can effectively enhance and guide the semantic analysis and understanding of short text in microblogs.
Wan-Qiu Cui, Junping Du 0001, Dawei Wang 0009, Xunpu Yuan, Feifei Kou, Liyan Zhou
ACM Trans. Intell. Syst. Technol.5
2019 Extended search method based on a semantic hashtag graph combining social and conceptual information
Wan-Qiu Cui, Junping Du 0001, Dawei Wang 0009, Feifei Kou, Meiyu Liang, Zhe Xue
World Wide Web4
2018 Hashtag Recommendation Based on Multi-Features of Microblogs
Feifei Kou, Junping Du 0001, Cong-Xian Yang, Yan-Song Shi, Wan-Qiu Cui, Meiyu Liang, Yue Geng
J. Comput. Sci. Technol.1