Jianhua Yin 0001

dblp:68/2053-1 · DBLP profile ↗
← Back
32ranked-venue papers
4as first author
24since 2021 · last 2026
0000-0002-4611-2986ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 14 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 12 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Self-Enhanced Image Clustering with Cross-Modal Semantic Consistency
abstract
While large language-image pre-trained models like CLIP offer powerful generic features for image clustering, existing methods typically freeze the encoder. This creates a fundamental mismatch between the model's task-agnostic representations and the demands of a specific clustering task, imposing a ceiling on performance. To break this ceiling, we propose a self-enhanced framework based on cross-modal semantic consistency for efficient image clustering. Our framework first builds a strong foundation via Cross-Modal Semantic Consistency and then specializes the encoder through Self-Enhancement. In the first stage, we focus on Cross-Modal Semantic Consistency. By mining consistency between generated image-text pairs at the instance, cluster assignment, and cluster center levels, we train lightweight clustering heads to align with the rich semantics of the pre-trained model. This alignment process is bolstered by a novel method for generating higher-quality cluster centers and a dynamic balancing regularizer to ensure well-distributed assignments. In the second stage, we introduce a Self-Enhanced fine-tuning strategy. The well-aligned model from the first stage acts as a reliable pseudo-label generator. These self-generated supervisory signals are then used to feed back the efficient, joint optimization of the vision encoder and clustering heads, unlocking their full potential. Extensive experiments on six mainstream datasets show that our method outperforms existing deep clustering methods by significant margins. Notably, our ViT-B/32 model already matches or even surpasses the accuracy of state-of-the-art methods built upon the far larger ViT-L/14.
Jianhua Yin 0001, Erwei Yin, Jianlong Wu
AAAI4
2026 Consistency and Invariance Guided Multi-View Hypergraph Learning for Robust Hyperedge Prediction
abstract
Hypergraphs, by extending traditional graphs with hyperedges, enable the modeling and prediction of complex higher-order interactions that go beyond simple pairwise interactions. Hyperedge prediction, an evolution of link prediction, aims to identify potential higher-order interactions—such as those in social media group chats—by recognizing and predicting hyperedges. Recently, hypergraph neural networks (HGNNs) have advanced hyperedge prediction by structuring higher-order interactions into a hypergraph, enabling effective capture of higher-order relations through information propagation across the hypergraph. However, existing methods primarily focus on developing complex HGNNs, underestimating the inherent unreliability of the underlying hypergraph due to incompleteness and noise, leading to suboptimal and fragile performance. In this article, we propose Multi-HyperLinker, a novel multi-view hypergraph learning framework that leverages the consistency and invariance across multiple views to capture reliable higher-order interaction patterns from historical observational data for robust hyperedge prediction. Specifically, to facilitate effective information propagation on incomplete hypergraphs, Multi-HyperLinker first synthesizes a tightly structured hypergraph and designs a consistency-guided dual-view learning strategy. To capture reliable higher-order interaction patterns on noisy hypergraphs, Multi-HyperLinker augments the hypergraphs by perturbing hyperedges to simulate variations and noise, and introduces an invariant learning strategy. Extensive experiments conducted on four real-world datasets demonstrate the superiority of Multi-HyperLinker, achieving performance improvements of up to 19.80% in hit rate compared to existing HGNN-based methods. Additionally, it exhibits enhanced robustness on incomplete and noisy hypergraphs.
Changyuan Tian 0001, Li Jin 0001, Zequn Zhang, Zhicong Lu, Wen Shi 0001, Jianhua Yin 0001, Shiyao Yan, Zhi Guo
ACM Trans. Knowl. Discov. Data6
2025 Dual-Center Graph Clustering with Neighbor Distribution
abstract
Graph clustering is crucial for unraveling intricate data structures, yet it presents significant challenges due to its unsupervised nature. Recently, goal-directed clustering techniques have yielded impressive results, with contrastive learning methods leveraging pseudo-label garnering considerable attention. Nonetheless, pseudo-label as a supervision signal is unreliable and existing goal-directed approaches utilize only features to construct a single-target distribution for single-center optimization, which lead to incomplete and less dependable guidance. In our work, we propose a novel Dual-Center Graph Clustering (DCGC) approach based on neighbor distribution properties, which includes representation learning with neighbor distribution and dual-center optimization. Specifically, we utilize neighbor distribution as a supervision signal to mine hard negative samples in contrastive learning, which is reliable and enhances the effectiveness of representation learning. Furthermore, neighbor distribution center is introduced alongside feature center to jointly construct a dual-target distribution for dual-center optimization. Extensive experiments and analysis demonstrate superior performance and effectiveness of our proposed method. The source code is available at https://github.com/chehaoa/DCGC.
Enhao Cheng, Shoujia Zhang, Jianhua Yin 0001, Li Jin 0001, Liqiang Nie
ECAI3
2025 Generative Agents for Multimodal Controversy Detection
abstract
Multimodal controversy detection, which involves determining whether a given video and its associated comments are controversial, plays a pivotal role in risk management on social video platforms. Existing methods typically provide only classification results, failing to identify what aspects are controversial and why, thereby lacking detailed explanations. To address this limitation, we propose a novel Agent-based Multimodal Controversy Detection architecture, termed AgentMCD. This architecture leverages Large Language Models (LLMs) as generative agents to simulate human behavior and improve explainability. AgentMCD employs a multi-aspect reasoning process, where multiple judges conduct evaluations from diverse perspectives to derive a final decision. Furthermore, a multi-agent simulation process is incorporated, wherein agents act as audiences, offering opinions and engaging in free discussions after watching videos. This hybrid framework enables comprehensive controversy evaluation and significantly enhances explainability. Experiments conducted on the MMCD dataset demonstrate that our proposed architecture outperforms existing LLM-based baselines in both high-resource and low-resource comment scenarios, while maintaining superior explainability.
Tianjiao Xu, Jinfei Gao, Keyi Kong, Jianhua Yin 0001, Tian Gan 0002, Liqiang Nie
IJCAI4
2025 Enhancing Democratic Mediation through Norm-Awareness in Generative Agent Societies
abstract
Democratic mediation serves as a vital mechanism for resolving social conflicts; however, current practices encounter three critical limitations: (1) inefficient operations, wherein traditional laborintensive mediation processes are both time-consuming and inefficient; (2) theoretical gaps, as prevailing mediation theories fail to explore the underlying causes of conflicts; and (3) inadequate analysis, with existing digital tools lacking comprehensive conflict mediation capabilities and primarily focusing on singular data types. To address these limitations, we introduce the Normative Social Simulator for Democratic Mediation, referred to as Norm Mediat. This framework is specifically designed to simulate democratic mediation, incorporating social norms. Central to this framework is the integration of normative reasoning into the mediation process, which enhances the ability to understand individuals' intrinsic needs and identify the root causes of conflicts. The framework comprises two essential components: (1) Dynamic Multimodal Conflict Modeling (DMCM), which generates the initial dataset of conflict interactions; and (2) Norm-Aware Iterative Mediation (NAIM), which implements an iterative democratic mediation process through norm awareness. The results of our human evaluation underscore the effectiveness of our norm-driven mediation strategies. This research significantly contributes to computational social science by providing a comprehensive methodological framework for simulating democratic processes and offering a benchmark dataset for conflict resolution studies.
Tianjiao Xu, Hao Fu 0004, Suiyang Zhang, Jianhua Yin 0001, Tian Gan 0002, Liqiang Nie
ACM Multimedia4
2025 Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization
abstract
Direct Preference Optimization (DPO) has emerged as an effective approach for mitigating hallucination in Multimodal Large Language Models (MLLMs). Although existing methods have achieved significant progress by utilizing vision-oriented contrastive objectives for enhancing MLLMs' attention to visual inputs and hence reducing hallucination, they suffer from non-rigorous optimization objective function and indirect preference supervision. To address these limitations, we propose a Symmetric Multimodal Preference Optimization (SymMPO), which conducts symmetric preference learning with direct preference supervision (i.e., response pairs) for visual understanding enhancement, while maintaining rigorous theoretical alignment with standard DPO. In addition to conventional ordinal preference learning, SymMPO introduces a preference margin consistency loss to quantitatively regulate the preference gap between symmetric preference pairs. Comprehensive evaluation across five benchmarks demonstrate SymMPO's superior performance, validating its effectiveness in hallucination mitigation of MLLMs.
Xuemeng Song, Yinwei Wei, Jianhua Yin 0001, Liqiang Nie
NeurIPS6
2025 Social Context-Aware Community-Level Propagation Prediction
abstract
With the increasing prevalence of online communities, social networks have become pivotal platforms for information propagation. However, this rise is accompanied by issues such as the spread of misinformation and online rumors. Community Level Information Pathway Prediction (CLIPP) is proposed to effectively stop the propagation of harmful information within specific communities. While progress has been made in understanding user-level propagation, there is a significant gap in addressing the CLIPP problem at the community level, particularly with regard to social context interpretation and the cold start problem in niche communities. To bridge this gap, we propose a novel model, named Community-Level Propagation Prediction with LLM enhanced Social Context Interpretation and Community Coldstart (ComPaSC3), which integrates three primary modules. The video enhancement module leverages LLMs to enrich the interpretation of multimedia content by embedding world knowledge. The community portrait building module utilizes LLMs to generate detailed community portraits for community interpretation. To tackle the community cold start problem, the dynamic commLink module links non-popular communities to the popular ones based on their portrait similarity, and dynamically updates their relationship weights. Our experimental results demonstrate that ComPaSC3 significantly improves predictive accuracy in both popular and non-popular scenarios. Particularly in non-popular communities, our approach outperforms existing state-of-the-art methods, achieving improvements of 10.00% - 15.20% in Rec@5 and 7.31% - 12.32% in NDCG@10.
Jinfei Gao, Xiao Wang 0056, Tian Gan 0002, Jianhua Yin 0001, Chuanchen Luo, Liqiang Nie
SIGIR4
2025 Hierarchical fine-grained multi-behavior recommendation with behavior-aware contrastive learning
Kaiyao Zhu, Jinhuan Liu, Xuemeng Song, Jianhua Yin 0001, Shuhan Qi, Junwei Du
Neural Networks4
2025 ClusMatch: Improving Deep Clustering by Unified Positive and Negative Pseudo-Label Learning
abstract
Recently, deep clustering methods have achieved remarkable results compared to traditional clustering approaches. However, its performance remains constrained by the absence of annotations. A thought-provoking observation is that there is still a significant gap between deep clustering and semi-supervised classification methods. Even with only a few labeled samples, the accuracy of semi-supervised learning is much higher than that of clustering. Given that we can annotate a small number of samples in a certain unsupervised way, the clustering task can be naturally transformed into a semi-supervised setting, thereby achieving comparable performance. Based on this intuition, we propose ClusMatch, a unified positive and negative pseudo-label learning based semi-supervised learning framework, which is pluggable and can be applied to existing deep clustering methods. Specifically, we first leverage the pre-trained deep clustering network to compute predictions for all samples, and then design specialized selection strategies to pick out a few high-quality samples as labeled samples for supervised learning. For the unselected samples, the novel unified positive and negative pseudo-label learning is introduced to provide additional supervised signals for semi-supervised fine-tuning. We also propose an adaptive positive-negative threshold learning strategy to further enhance the confidence of generated pseudo-labels. Extensive experiments on six widely-used datasets and one large-scale dataset demonstrate the superiority of our proposed ClusMatch. For example, ClusMatch achieves a significant accuracy improvement of 5.4% over the state-of-the-art method ProPos on an average of these six datasets.
Jianlong Wu, Jianhua Yin 0001, Liqiang Nie, Zhouchen Lin
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Video Event Extraction with Multi-View Interaction Knowledge Distillation
abstract
Video event extraction (VEE) aims to extract key events and generate the event arguments for their semantic roles from the video. Despite promising results have been achieved by existing methods, they still lack an elaborate learning strategy to adequately consider: (1) inter-object interaction, which reflects the relation between objects; (2) inter-modality interaction, which aligns the features from text and video modality. In this paper, we propose a Multi-view Interaction with knowledge Distillation (MID) framework to solve the above problems with the Knowledge Distillation (KD) mechanism. Specifically, we propose the self-Relational KD (self-RKD) to enhance the inter-object interaction, where the relation between objects is measured by distance metric, and the high-level relational knowledge from the deeper layer is taken as the guidance for boosting the shallow layer in the video encoder. Meanwhile, to improve the inter-modality interaction, the Layer-to-layer KD (LKD) is proposed, which integrates additional cross-modal supervisions (i.e., the results of cross-attention) with the textual supervising signal for training each transformer decoder layer. Extensive experiments show that without any additional parameters, MID achieves the state-of-the-art performance compared to other strong methods in VEE.
Kaiwen Wei, Runyan Du, Li Jin 0001, Jian Liu 0032, Jianhua Yin 0001, Linhao Zhang, Nayu Liu, Zhi Guo
AAAI5
2024 A Multi-View Clustering Algorithm for Short Text
abstract
The objective of the short text clustering task is to group semantically similar short texts into one class and segregate semantically different short texts. Despite the commendable performance achieved by existing topic model based short text clustering algorithms and deep clustering models, a fundamental limitation persists. Both of them are based on one view of the text, which inevitably constrains their clustering performance. Specifically, the topic model based short text clustering algorithms represent short texts as bag-of-words, while the deep clustering models represent short texts as document embeddings. To address these issues, we propose a Multi-View Clustering (MVC) model that considers both views of the text. We modeled the bag-of-words view using the Dirichlet Multinomial Mixture (DMM) model and the document embedding view using the Gaussian Mixture Model (GMM). A Bernoulli random variable is used to control these two models, enabling our proposed model to utilize the semantic information of short text embeddings while obtaining the bag-of-words information. Extensive experiments on four datasets demonstrate MVC's effectiveness. The code for MVC is available at https://github.com/jhyin12/MVC.
Minkuan Lu, Jianhua Yin 0001, Kaijun Wang, Liqiang Nie
ICDE2
2024 Self-Training Boosted Multi-Factor Matching Network for Composed Image Retrieval
abstract
The composed image retrieval (CIR) task aims to retrieve the desired target image for a given multimodal query, i.e., a reference image with its corresponding modification text. The key limitations encountered by existing efforts are two aspects: 1) ignoring the multiple query-target matching factors; 2) ignoring the potential unlabeled reference-target image pairs in existing benchmark datasets. To address these two limitations is non-trivial due to the following challenges: 1) how to effectively model the multiple matching factors in a latent way without direct supervision signals; 2) how to fully utilize the potential unlabeled reference-target image pairs to improve the generalization ability of the CIR model. To address these challenges, in this work, we first propose a CLIP-Transformer based muLtI-factor Matching Network (LIMN), which consists of three key modules: disentanglement-based latent factor tokens mining, dual aggregation-based matching token learning, and dual query-target matching modeling. Thereafter, we design an iterative dual self-training paradigm to further enhance the performance of LIMN by fully utilizing the potential unlabeled reference-target image pairs in a weakly-supervised manner. Specifically, we denote the iterative dual self-training paradigm enhanced LIMN as LIMN+. Extensive experiments on four datasets, including FashionIQ, Shoes, CIRR, and Fashion200 K, show that our proposed LIMN and LIMN+ significantly surpass the state-of-the-art baselines.
Haokun Wen, Xuemeng Song, Jianhua Yin 0001, Jianlong Wu, Weili Guan, Liqiang Nie
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Audio-Driven Talking Video Frame Restoration
abstract
Talking video frames occasionally drop while streaming for reasons like network errors, which greatly hurts the online team collaboration and user experiences. Directly generating the dropped frames from the remaining ones is unfavorable since a person’s lip motion is usually non-linear and thus hard to be restored when consecutive frames are missing. Nevertheless, the audio content provides strong signals for lip motion and is less likely to drop during transmitting. Inspired by this, as an initial attempt, we present the task of audio-driven talking video frame restoration in this paper, i.e., restoring dropped video frames by jointly leveraging the audio and remaining video frames. Towards the high-quality frame generation, we devise a cross-modal frame restoration network. This network aligns the complete audio content with video frames, precisely identifies and sequentially generates the dropped frames. To justify our model, we construct a new dataset, Talking Video Frames Drop, TVFD for short, consisting of 2.5K video and 144K frames in total. We conduct extensive experiments over TVFD and another publicly accessible dataset - Voxceleb2. Our model obtains significantly improved performance as compared to other state-of-the-art competitors.
Harry Cheng 0002, Jianhua Yin 0001, Jiafang Wang, Liqiang Nie
IEEE Trans. Multim.3
2023 Data-free Knowledge Distillation for Fine-grained Visual Categorization
abstract
Data-free knowledge distillation (DFKD) is a promising approach for addressing issues related to model compression, security privacy, and transmission restrictions. Although the existing methods exploiting DFKD have achieved inspiring achievements in coarse-grained classification, in practical applications involving fine-grained classification tasks that require more detailed distinctions between similar categories, sub-optimal results are obtained. To address this issue, we propose an approach called DFKD-FGVC that extends DFKD to fine-grained visual categorization (FGVC) tasks. Our approach utilizes an adversarial distillation framework with attention generator, mixed high-order attention distillation, and semantic feature contrast learning. Specifically, we introduce a spatial-wise attention mechanism to the generator to synthesize fine-grained images with more details of discriminative parts. We also utilize the mixed high-order attention mechanism to capture complex interactions among parts and the subtle differences among discriminative features of the fine-grained categories, paying attention to both local features and semantic context relationships. Moreover, we leverage the teacher and student models of the distillation framework to contrast high-level semantic feature maps in the hyperspace, comparing variances of different categories. We evaluate our approach on three widely-used FGVC benchmarks (Aircraft, Cars196, and CUB200) and demonstrate its superior performance. Code is available at https://github.com/RoryShao/DFKD-FGVC.git
Renrong Shao, Wei Zhang 0056, Jianhua Yin 0001, Jun Wang 0006
ICCV3
2023 Semantic-Aware Modular Capsule Routing for Visual Question Answering
abstract
Visual Question Answering (VQA) is fundamentally compositional in nature, and many questions are simply answered by decomposing them into modular sub-problems. The recent proposed Neural Module Network (NMN) employ this strategy to question answering, whereas heavily rest with off-the-shelf layout parser or additional expert policy regarding the network architecture design instead of learning from the data. These strategies result in the unsatisfactory adaptability to the semantically-complicated variance of the inputs, thereby hindering the representational capacity and generalizability of the model. To tackle this problem, we propose a Semantic-aware modUlar caPsulE Routing framework, termed as SUPER, to better capture the instance-specific vision-semantic characteristics and refine the discriminative representations for prediction. Particularly, five powerful specialized modules as well as dynamic routers are tailored in each layer of the SUPER network, and the compact routing spaces are constructed such that a variety of customizable routes can be sufficiently exploited and the vision-semantic representations can be explicitly calibrated. We comparatively justify the effectiveness and generalization ability of our proposed SUPER scheme over five benchmark datasets, as well as the parametric-efficient advantage. It is worth emphasizing that this work is not to pursue the state-of-the-art results in VQA. Instead, we expect that our model is responsible to provide a novel perspective towards architecture learning and representation calibration for VQA.
Jianhua Yin 0001, Jianlong Wu, Yinwei Wei, Liqiang Nie
IEEE Trans. Image Process.2
2023 Causal Inference for Leveraging Image-Text Matching Bias in Multi-Modal Fake News Detection
abstract
Multi-modal fake news detection has drawn considerable attention with the development of online social media. Existing methods primarily conduct direct cross-modal fusion, while ignoring the image-text matching degree which may introduce unexpected bias. This work studies an unexplored problem in multi-modal fake news detection – how to deconfound and leverage the image-text matching bias to improve the performance of fake news detection. The key lies in two aspects: how to remove the confounding effect of the image-text matching bias during training, and how to utilize the bias in the inference stage since the news with mismatched image and text is more likely to be fake. To achieve our goal, we formulate the fake news detection task as a causal graph that reflects the cause-effect factors, and propose a novel framework –Causal Inference forLeveragingImage-textMatchingBias (CLIMB) in multi-modal fake news detection. To our best knowledge, this is the first work that considers the image-text matching degree into the fake news detection task with the approach of causal inference. CLIMB can be applied to any fake news detection models with visual and textual features as inputs. Extensive experiments on two real-world datasets validate the effectiveness of CLIMB.
Linmei Hu, Ziwang Zhao, Jianhua Yin 0001, Liqiang Nie
IEEE Trans. Knowl. Data Eng.4
2023 HS-GCN: Hamming Spatial Graph Convolutional Networks for Recommendation
abstract
An efficient solution to the large-scale recommender system is to represent users and items as binary hash codes in the Hamming space. Towards this end, existing methods tend to code users by modeling their Hamming similarities with the items they historically interact with, which are termed as the first-order similarities in this work. Despite of their efficiency, these methods suffer from the suboptimal representative capacity, since they forgo the correlation established by connecting multiple first-order similarities, i.e., the relation among the indirect instances, which could be defined as the high-order similarity. To tackle this drawback, we propose to model both the first- and the high-order similarities in the Hamming space through the user-item bipartite graph. Therefore, we develop a novel learning to hash framework, namely Hamming Spatial Graph Convolutional Networks (HS-GCN), which explicitly models the Hamming similarity and embeds it into the codes of users and items. Extensive experiments on three public benchmark datasets demonstrate that our proposed model significantly outperforms several state-of-the-art hashing models, and obtains performance comparable with the real-valued recommendation models.
Yinwei Wei, Jianhua Yin 0001, Liqiang Nie
IEEE Trans. Knowl. Data Eng.3
2023 Self-Supervised Correlation Learning for Cross-Modal Retrieval
abstract
Cross-modal retrieval aims to retrieve relevant data from another modality when given a query of one modality. Although most existing methods that rely on the label information of multimedia data have achieved promising results, the performance benefiting from labeled data comes at a high cost since labeling data often requires enormous labor resources, especially on large-scale multimedia datasets. Therefore, unsupervised cross-modal learning is of crucial importance in real-world applications. In this paper, we propose a novel unsupervised cross-modal retrieval method, named Self-supervised Correlation Learning (SCL), which takes full advantage of large amounts of unlabeled data to learn discriminative and modality-invariant representations. Since unsupervised learning lacks the supervision of category labels, we incorporate the knowledge from the input as a supervisory signal by maximizing the mutual information between the input and the output of different modality-specific projectors. Besides, for the purpose of learning discriminative representations, we exploit unsupervised contrastive learning to model the relationship among intra- and inter-modality instances, which makes similar samples closer and pushes dissimilar samples apart. Moreover, to further eliminate the modality gap, we use a weight-sharing scheme and minimize the modality-invariant loss in the joint representation space. Beyond that, we also extend the proposed method to the semi-supervised setting. Extensive experiments conducted on three widely-used benchmark datasets demonstrate that our method achieves competitive results compared with current state-of-the-art cross-modal retrieval approaches.
Jianlong Wu, Leigang Qu, Tian Gan 0002, Jianhua Yin 0001, Liqiang Nie
IEEE Trans. Multim.5
2023 DualGNN: Dual Graph Neural Network for Multimedia Recommendation
abstract
One of the important factors affecting micro-video recommender systems is to model the multi-modal user preference on the micro-video. Despite the remarkable performance of prior arts, they are still limited by fusing the user preference derived from different modalities in a unified manner, ignoring the users tend to place different emphasis on different modalities. Furthermore, modality-missing is ubiquity and unavoidable in the micro-video recommendation, some modalities information of micro-videos are lacked in many cases, which negatively affects the multi-modal fusion operations. To overcome these disadvantages, we propose a novel framework for the micro-video recommendation, dubbed Dual Graph Neural Network (DualGNN), upon the user-microvideo bipartite and user co-occurrence graphs, which leverages the correlation between users to collaboratively mine the particular fusion pattern for each user. Specifically, we first introduce a single-modal representation learning module, which performs graph operations on the user-microvideo graph in each modality to capture single-modal user preferences on different modalities. And then, we devise a multi-modal representation learning module to explicitly model the user’s attentions over different modalities and inductively learn the multi-modal user preference. Finally, we propose a prediction module to rank the potential micro-videos for users. Extensive experiments on two public datasets demonstrate the significant superiority of our DualGNN over state-of-the-arts methods.
Qifan Wang 0001, Yinwei Wei, Jianhua Yin 0001, Jianlong Wu, Xuemeng Song, Liqiang Nie
IEEE Trans. Multim.3
2023 Review Polarity-Wise Recommender
abstract
The de facto review-involved recommender systems, using review information to enhance recommendation, have received increasing interest over the past years. Thereinto, one advanced branch is to extract salient aspects from textual reviews (i.e., the item attributes that users express) and combine them with the matrix factorization (MF) technique. However, the existing approaches all ignore the fact that semantically different reviews often include opposite aspect information. In particular, positive reviews usually express aspects that users prefer, while the negative ones describe aspects that users dislike. As a result, it may mislead the recommender systems into making incorrect decisions pertaining to user preference modeling. Toward this end, in this article, we present a review polarity-wise recommender model, dubbed as RPR, to discriminately treat reviews with different polarities. To be specific, in this model, positive and negative reviews are separately gathered and used to model the user-preferred and user-rejected aspects, respectively. Besides, to overcome the imbalance of semantically different reviews, we further develop an aspect-aware importance weighting strategy to align the aspect importance for these two kinds of reviews. Extensive experiments conducted on eight benchmark datasets have demonstrated the superiority of our model when compared with several state-of-the-art review-involved baselines. Moreover, our method can provide certain explanations to real-world rating prediction scenarios.
Jianhua Yin 0001, Zan Gao 0001, Liqiang Nie
IEEE Trans. Neural Networks Learn. Syst.3
2022 Attentive Representation Learning With Adversarial Training for Short Text Clustering
abstract
Short text clustering has far-reaching effects on semantic analysis, showing its importance for multiple applications such as corpus summarization and information retrieval. However, it inevitably encounters the severe sparsity of short text representations, making the previous clustering approaches still far from satisfactory. In this paper, we present a novel attentive representation learning model for shot text clustering, wherein cluster-level attention is proposed to capture the correlations between text representations and cluster representations. Relying on this, the representation learning and clustering for short texts are seamlessly integrated into a unified model. To further ensure robust model training for short texts, we apply adversarial training to the unsupervised clustering setting, by injecting perturbations into the cluster representations. The model parameters and perturbations are optimized alternately through a minimax game. Extensive experiments on four real-world short text datasets demonstrate the superiority of the proposed model over several strong competitors, verifying that robust adversarial training yields substantial performance gains.
Wei Zhang 0056, Jianhua Yin 0001, Jianyong Wang 0001
IEEE Trans. Knowl. Data Eng.3
2022 Question Tagging via Graph-guided Ranking
abstract
With the increasing prevalence of portable devices and the popularity of community Question Answering (cQA) sites, users can seamlessly post and answer many questions. To effectively organize the information for precise recommendation and easy searching, these platforms require users to select topics for their raised questions. However, due to the limited experience, certain users fail to select appropriate topics for their questions. Thereby, automatic question tagging becomes an urgent and vital problem for the cQA sites, yet it is non-trivial due to the following challenges. On the one hand, vast and meaningful topics are available yet not utilized in the cQA sites; how to model and tag them to relevant questions is a highly challenging problem. On the other hand, related topics in the cQA sites may be organized into a directed acyclic graph. In light of this, how to exploit relations among topics to enhance their representations is critical. To settle these challenges, we devise a graph-guided topic ranking model to tag questions in the cQA sites appropriately. In particular, we first design a topic information fusion module to learn the topic representation by jointly considering the name and description of the topic. Afterwards, regarding the special structure of topics, we propose an information propagation module to enhance the topic representation. As the comprehension of questions plays a vital role in question tagging, we design a multi-level context-modeling-based question encoder to obtain the enhanced question representation. Moreover, we introduce an interaction module to extract topic-aware question information and capture the interactive information between questions and topics. Finally, we utilize the interactive information to estimate the ranking scores for topics. Extensive experiments on three Chinese cQA datasets have demonstrated that our proposed model outperforms several state-of-the-art competitors.
Xiao Zhang 0015, Meng Liu 0006, Jianhua Yin 0001, Zhaochun Ren, Liqiang Nie
ACM Trans. Inf. Syst.3
2022 Answer Questions with Right Image Regions: A Visual Attention Regularization Approach
abstract
Visual attention in Visual Question Answering (VQA) targets at locating the right image regions regarding the answer prediction, offering a powerful technique to promote multi-modal understanding. However, recent studies have pointed out that the highlighted image regions from the visual attention are often irrelevant to the given question and answer, leading to model confusion for correct visual reasoning. To tackle this problem, existing methods mostly resort to aligning the visual attention weights with human attentions. Nevertheless, gathering such human data is laborious and expensive, making it burdensome to adapt well-developed models across datasets. To address this issue, in this article, we devise a novel visual attention regularization approach, namely, AttReg, for better visual grounding in VQA. Specifically, AttReg first identifies the image regions that are essential for question answering yet unexpectedly ignored (i.e., assigned with low attention weights) by the backbone model. And then a mask-guided learning scheme is leveraged to regularize the visual attention to focus more on these ignored key regions. The proposed method is very flexible and model-agnostic, which can be integrated into most visual attention-based VQA models and require no human attention supervision. Extensive experiments over three benchmark datasets, i.e., VQA-CP v2, VQA-CP v1, and VQA v2, have been conducted to evaluate the effectiveness of AttReg. As a by-product, when incorporating AttReg into the strong baseline LMH, our approach can achieve a new state-of-the-art accuracy of 60.00% with an absolute performance gain of 7.01% on the VQA-CP v2 benchmark dataset. In addition to the effectiveness validation, we recognize that the faithfulness of the visual attention in VQA has not been well explored in literature. In the light of this, we propose to empirically validate such property of visual attention and compare it with the prevalent gradient-based approaches.
Yibing Liu, Jianhua Yin 0001, Xuemeng Song, Weifeng Liu 0001, Liqiang Nie, Min Zhang 0005
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Focal and Composed Vision-semantic Modeling for Visual Question Answering
abstract
Visual Question Answering (VQA) is a vital yet challenging task in the field of multimedia comprehension. In order to correctly answer questions about an image, a VQA model requires to sufficiently understand the visual scene, especially the vision-semantic reasonings between the two modalities. Traditional relation-based methods allow to encode the pairwise relations of objects to boost the VQA model performance. However, this simple strategy is deficient to exploit the abundant concepts expressed by the composition of diverse image objects, leading to sub-optimal performance. In this paper, we propose a focal and composed vision-semantic modeling method, which is a trainable end-to-end model, for better vision-semantic redundancy removal and compositionality modeling. Concretely, we first introduce the LENA cell, a plug-and-play reasoning module, which removes redundant semantic by a focal mechanism in the first step, followed by the vision-semantic compositionality modeling for better visual reasoning. We then incorporate the cell into a full LENA network, which progressively refines multimodal composed representations, and can be leveraged to infer the high-order vision-semantic in a multi-step learning way. Extensive experiments on two benchmark datasets, i.e., VQA v2 and VQA-CP v2, verify the superiority of our model as compared with several state-of-the-art baselines.
Jianhua Yin 0001, Meng Liu 0006, Yupeng Hu 0003, Liqiang Nie
ACM Multimedia3
2020 Editorial of Special Issue of ICDM 2019
Wei Shen 0004, Wei Zhang 0056, Jianhua Yin 0001, Jianyong Wang 0001
Data Sci. Eng.3
2019 Long-tail Hashtag Recommendation for Micro-videos with Graph Convolutional Network
abstract
Hashtags, a user provides to a micro-video, are the ones which can well describe the semantics of the micro-video's content in his/her mind. At the same time, hashtags have been widely used to facilitate various micro-video retrieval scenarios (e.g., search, browse, and categorization). Despite their importance, numerous micro-videos lack hashtags or contain inaccurate or incomplete hashtags. In light of this, hashtag recommendation, which suggests a list of hashtags to a user when he/she wants to annotate a post, becomes a crucial research problem. However, little attention has been paid to micro-video hashtag recommendation, mainly due to the following three reasons: 1) lack of benchmark dataset; 2) the temporal and multi-modality characteristics of micro-videos; and 3) hashtag sparsity and long-tail distributions. In this paper, we recommend hashtags for micro-videos by presenting a novel multi-view representation interactive embedding model with graph-based information propagation. It is capable of boosting the performance of micro-videos hashtag recommendation by jointly considering the sequential feature learning, the video-user-hashtag interaction, and the hashtag correlations. Extensive experiments on a constructed dataset demonstrate our proposed method outperforms state-of-the-art baselines. As a side research contribution, we have released our dataset and codes to facilitate the research in this community.
Tian Gan 0002, Meng Liu 0006, Zhiyong Cheng 0001, Jianhua Yin 0001, Liqiang Nie
CIKM5
2019 Routing Micro-videos via A Temporal Graph-guided Recommendation System
abstract
In the past few years, micro-videos have become the dominant trend in the social media era. Meanwhile, as the number of microvideos increases, users are frequently overwhelmed by their uninterested ones. Despite the success of existing recommendation systems developed for various communities, they cannot be applied to routing micro-videos, since users in micro-video platforms have their unique characteristics: diverse and dynamic interest, multilevel interest, as well as true negative samples. To address these problems, we present a temporal graph-guided recommendation system. In particular, we first design a novel graph-based sequential network to simultaneously model users' dynamic and diverse interest.Similarly, uninterested information can be captured from users'true negative samples. Beyond that, we introduce users' multi-level interest into our recommendation model via a user matrix that is able to learn the enhanced representation of users' interest. Finally, the system can make accurate recommendation by considering the above characteristics. Experimental results on two public datasets verify the effectiveness of our proposed model.
Yongqi Li 0001, Meng Liu 0006, Jianhua Yin 0001, Chaoran Cui, Xin-Shun Xu, Liqiang Nie
ACM Multimedia3
2019 Prototype-guided Attribute-wise Interpretable Scheme for Clothing Matching
abstract
Recently, as an essential part of people's daily life, clothing matching has gained increasing research attention. Most existing efforts focus on the numerical compatibility modeling between fashion items with advanced neural networks, and hence suffer from the poor interpretation, which makes them less applicable in real world applications. In fact, people prefer to know not only whether the given fashion items are compatible, but also the reasonable interpretations as well as suggestions regarding how to make the incompatible outfit harmonious. Considering that the research line of the comprehensively interpretable clothing matching is largely untapped, in this work, we propose a prototype-guided attribute-wise interpretable compatibility modeling (PAICM) scheme, which seamlessly integrates the latent compatible/incompatible prototype learning and compatibility modeling with the Bayesian personalized ranking (BPR) framework. In particular, the latent attribute interaction prototypes, learned by the non-negative matrix factorization (NMF), are treated as templates to interpret the discordant attribute and suggest the alternative item for each fashion item pair. Extensive experiments on the real-world dataset have demonstrated the effectiveness of our scheme.
Xianjing Han, Xuemeng Song, Jianhua Yin 0001, Yinglong Wang 0001, Liqiang Nie
SIGIR3
2018 Model-based Clustering of Short Text Streams
abstract
Short text stream clustering has become an increasingly important problem due to the explosive growth of short text in diverse social medias. In this paper, we propose a model-based short text stream clustering algorithm (MStream) which can deal with the concept drift problem and sparsity problem naturally. The MStream algorithm can achieve state-of-the-art performance with only one pass of the stream, and can have even better performance when we allow multiple iterations of each batch. We further propose an improved algorithm of MStream with forgetting rules called MStreamF, which can efficiently delete outdated documents by deleting clusters of outdated batches. Our extensive experimental study shows that MStream and MStreamF can achieve better performance than three baselines on several real datasets.
Jianhua Yin 0001, Daren Chao, Zhongkun Liu, Wei Zhang 0056, Xiaohui Yu 0001, Jianyong Wang 0001
KDD1
2016 A model-based approach for text clustering with outlier detection
abstract
Text clustering is a challenging problem due to the high-dimensional and large-volume characteristics of text datasets. In this paper, we propose a collapsed Gibbs Sampling algorithm for the Dirichlet Process Multinomial Mixture model for text clustering (abbr. to GSDPMM) which does not need to specify the number of clusters in advance and can cope with the high-dimensional problem of text clustering. Our extensive experimental study shows that GSDPMM can achieve significantly better performance than three other clustering methods and can achieve high consistency on both long and short text datasets. We found that GSDPMM has low time and space complexity and can scale well with huge text datasets. We also propose some novel and effective methods to detect the outliers in the dataset and obtain the representative words of each cluster.
Jianhua Yin 0001, Jianyong Wang 0001
ICDE1
2016 A Text Clustering Algorithm Using an Online Clustering Scheme for Initialization
abstract
In this paper, we propose a text clustering algorithm using an online clustering scheme for initialization called FGSDMM+. FGSDMM+ assumes that there are at most Kmax clusters in the corpus, and regards these Kmax potential clusters as one large potential cluster at the beginning. During initialization, FGSDMM+ processes the documents one by one in an online clustering scheme. The first document will choose the potential cluster, and FGSDMM+ will create a new cluster to store this document. Later documents will choose one of the non-empty clusters or the potential cluster with probabilities derived from the Dirichlet multinomial mixture model. Each time a document chooses the potential cluster, FGSDMM+ will create a new cluster to store that document and decrease the probability of later documents choosing the potential cluster. After initialization, FGSDMM+ will run a collapsed Gibbs sampling algorithm several times to obtain the final clustering result. Our extensive experimental study shows that FGSDMM+ can achieve better performance than three other clustering methods on both short and long text datasets.
Jianhua Yin 0001, Jianyong Wang 0001
KDD1
2014 A dirichlet multinomial mixture model-based approach for short text clustering
abstract
Short text clustering has become an increasingly important task with the popularity of social media like Twitter, Google+, and Facebook. It is a challenging problem due to its sparse, high-dimensional, and large-volume characteristics. In this paper, we proposed a collapsed Gibbs Sampling algorithm for the Dirichlet Multinomial Mixture model for short text clustering (abbr. to GSDMM). We found that GSDMM can infer the number of clusters automatically with a good balance between the completeness and homogeneity of the clustering results, and is fast to converge. GSDMM can also cope with the sparse and high-dimensional problem of short texts, and can obtain the representative words of each cluster. Our extensive experimental study shows that GSDMM can achieve significantly better performance than three other clustering models.
Jianhua Yin 0001, Jianyong Wang 0001
KDD1