Xiaowen Huang 0001

dblp:166/0337 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0001-9590-3285ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Membership Inference Attack Against Large Language Model-Based Recommendation Systems: A New Distillation-Based Paradigm
abstract
Membership Inference Attack (MIA) aims to determine whether a specific data sample was included in the training dataset of a target model. Traditional MIA approaches rely on shadow models to mimic target model behavior, but their effectiveness diminishes for Large Language Model (LLM)-based recommendation systems due to the scale and complexity of training data. This paper introduces a novel knowledge distillation-based MIA paradigm tailored for LLM-based recommendation systems. Our method constructs a reference model via distillation, applying distinct strategies for member and non-member data to enhance discriminative capabilities. The paradigm extracts fused features (e.g., confidence, entropy, loss, and hidden layer vectors) from the reference model to train an attack model, overcoming limitations of individual features. Extensive experiments on extended datasets (Last.FM, MovieLens, Book-Crossing, Delicious) and diverse LLMs (T5, GPT-2, LLaMA3) demonstrate that our approach significantly outperforms shadow model-based MIAs and individual-feature baselines. The results show its practicality for privacy attacks in LLM-driven recommender systems.
Cuihong Li, Xiaowen Huang 0001, Chuanhuan Yin, Jitao Sang 0001
AAAI2
2026 ITDR: An Instruction Tuning Dataset for Enhancing Large Language Models in Recommendations
abstract
Large language models (LLMs) have demonstrated outstanding performance in natural language processing tasks. However, in the field of recommender systems, due to the inherent structural discrepancy between user behavior data and natural language, LLMs struggle to effectively model the associations between user preferences and items. Although prompt-based methods can generate recommendation results, their inadequate understanding of recommendation tasks leads to constrained performance. To address this gap, we construct a comprehensive instruction tuning dataset, ITDR, which encompasses seven subtasks across two root tasks: user-item interaction and user-item understanding. The dataset integrates data from 13 public recommendation datasets and is built using manually crafted standardized templates, comprising approximately 200,000 instances. Experimental results demonstrate that ITDR significantly enhances the performance of mainstream open-source LLMs such as GLM-4, Qwen2.5, Qwen2.5-Instruct and LLaMA-3.2 on recommendation tasks. Furthermore, we analyze the correlations between tasks and explore the impact of task descriptions and data scale on instruction tuning effectiveness. Finally, we perform comparative experiments against closed-source LLMs with massive parameters. Our tuning dataset ITDR, the fine-tuned large recommendation models, all LoRA modules, and the complete experimental results are available at https://github.com/hellolzk/ITDR.
Xiaowen Huang 0001, Jitao Sang 0001
KDD (1)2
2026 Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper
abstract
Recent video-to-audio (V2A) generation relies on extracting semantic and temporal features from video to condition generative models. Training these models from scratch is resource-intensive. Consequently, leveraging foundation models (FMs) has gained traction due to their cross-modal knowledge transfer and generalization capabilities. One prior work has explored fine-tuning a lightweight mapper network to connect a pre-trained visual encoder with a text-to-audio generation model for V2A. Inspired by this, we introduce the Multiple Foundation Model Mapper (MFM-Mapper). Compared to the previous mapper approach, MFM-Mapper benefits from richer semantic and temporal information by fusing features from dual visual encoders. Furthermore, by replacing a linear mapper with GPT-2, MFM-Mapper improves feature alignment, drawing parallels between cross-modal features mapping and autoregressive translation tasks. Our MFM-Mapper demonstrates high efficiency in terms of both data requirements and training costs. Compared to prior linear mapper-based work, it achieves superior performance in semantic and temporal consistency with only about 40% of the training data and 40% of the training epochs.
Gehui Chen, Guan'an Wang, Xiaowen Huang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2025 SILLM4Rec: Self-Improving with Chain of Thought Enhanced Preference Optimization for Multimodal Recommendation
abstract
Recent explorations into the potential of large language models (LLMs) within recommendation systems have demonstrated promising performance. However, LLM-based recommenders still face significant challenges in complex scenarios. Most existing alignment approaches primarily focus on direct item generation based on user-interaction histories, frequently neglecting to fully exploit valuable feedback such as reviews and ratings. Furthermore, these methods rarely consider the benefits of incorporating reasoning mechanisms to enhance recommendation accuracy. To overcome these limitations and improve the reliability of LLM-based recommenders under data-scarce conditions, we propose SILLM4Rec, a framework specifically designed to strengthen both reasoning ability and recommendation performance. Our approach begins by extracting high-quality Chain-of-Thought (CoT) reasoning samples from a more capable teacher model. These samples are then used to fine-tune a smaller student model through a self-learning and iterative optimization process, enabling it to refine its outputs and adapt to user preferences. Comprehensive experiments on three publicly available benchmark datasets demonstrate that SILLM4Rec not only achieves superior performance metrics, but also enhances transparency and robustness in recommendation tasks. Our code is available at https://github.com/MKC-Lab/SILLM4Rec.
Fang Quan, Xiaowen Huang 0001, Jitao Sang 0001
MMAsia4
2025 A LLM-based Controllable, Scalable, Human-Involved User Simulator Framework for Conversational Recommender Systems
abstract
Conversational Recommender System (CRS) leverages real-time feedback from users to dynamically model their preferences, thereby enhancing the system's ability to provide personalized recommendations and improving the overall user experience. CRS has demonstrated significant promise, prompting researchers to concentrate their efforts on developing user simulators that are both more realistic and trustworthy. The advent of Large Language Models (LLMs) has demonstrated capabilities that approach human-level intelligence across a diverse range of tasks. Research efforts have been made to utilize LLMs for building user simulators to evaluate the performance of CRS. Although these efforts showcase innovation, they are accompanied by certain limitations. In this work, we introduce a Controllable, Scalable, and Human-Involved (CSHI) simulator framework that manages the behavior of user simulators across various stages via a plugin manager. CSHI tailors behavioral simulations and interaction patterns to deliver authentic user-system engagement experiences. Through experiments and case studies in two conversational recommendation scenarios, we show that our framework can adapt to a variety of conversational recommendation settings and effectively simulate users' personalized preferences. Consequently, our simulator is able to generate feedback that closely mirrors that of real users. This facilitates a reliable assessment of existing CRS studies and promotes the creation of high-quality conversational recommendation datasets.
Lixi Zhu, Xiaowen Huang 0001, Jitao Sang 0001
WWW2
2025 Towards Robust Recommendation: A Review and an Adversarial Robustness Evaluation Library
abstract
Recently, recommender system has achieved significant success. However, due to the openness of recommender systems, they remain vulnerable to malicious attacks. Additionally, natural noise in training data and issues such as data sparsity can also degrade the performance of recommender systems. Therefore, enhancing the robustness of recommender systems has become an increasingly important research topic. In this survey, we provide a comprehensive overview of the robustness of recommender systems. Based on our investigation, we categorize the robustness of recommender systems into adversarial robustness and non-adversarial robustness. In the adversarial robustness, we introduce the fundamental principles and classical methods of recommender system adversarial attacks and defenses. In the non-adversarial robustness, we analyze nonadversarial robustness from the perspectives of data sparsity, natural noise, and data imbalance. Additionally, we summarize commonly used datasets and evaluation metrics for evaluating the robustness of recommender systems. Finally, we also discuss the current challenges in the field of recommender system robustness and potential future research directions. Additionally, to facilitate fair and efficient evaluation of attack and defense methods in adversarial robustness, we propose an adversarial robustness evaluation library–ShillingREC, and we conduct evaluations of basic attack models and recommendation models. ShillingREC project is released at https://github.com/ chengleileilei/ShillingREC.
Xiaowen Huang 0001, Jitao Sang 0001, Jian Yu 0001
IEEE Trans. Knowl. Data Eng.2
2024 Euclidean-Distance-Preserved Feature Reduction for efficient person re-identification
Guan'an Wang, Xiaowen Huang 0001, Yang Yang 0062, Prayag Tiwari, Jian Zhang 0018
Neural Networks2
2024 Faster Person Re-Identification: One-Shot-Filter and Coarse-to-Fine Search
abstract
Fast person re-identification (ReID) aims to search person images quickly and accurately. The main idea of recent fast ReID methods is the hashing algorithm, which learns compact binary codes and performs fast Hamming distance and counting sort. However, a very long code is needed for high accuracy (e.g.2048), which compromises search speed. In this work, we introduce a new solution for fast ReID by formulating a novel Coarse-to-Fine (CtF) hashing code search strategy, which complementarily uses short and long codes, achieving both faster speed and better accuracy. It uses shorter codes to coarsely rank broad matching similarities and longer codes to refine only a few top candidates for more accurate instance ReID. Specifically, we design an All-in-One (AiO) module together with a Distance Threshold Optimization (DTO) algorithm. In AiO, we simultaneously learn and enhance multiple codes of different lengths in a single model. It learns multiple codes in a pyramid structure, and encourage shorter codes to mimic longer codes by self-distillation. DTO solves a complex threshold search problem by a simple optimization process, and the balance between accuracy and speed is easily controlled by a single parameter. It formulates the optimization target as a$F_{\beta }$score that can be optimised by Gaussian cumulative distribution functions. Besides, we find even short code (e.g.32) still takes a long time under large-scale gallery due to the$O(n)$time complexity. To solve the problem, we propose a gallery-size-free latent-attributes-based One-Shot-Filter (OSF) strategy, that is always$O(1)$time complexity, to quickly filter major easy negative gallery images, Specifically, we design a Latent-Attribute-Learning (LAL) module supervised a Single-Direction-Metric (SDM) Loss. LAL is derived from principal component analysis (PCA) that keeps largest variance using shortest feature vector, meanwhile enabling batch and end-to-end learning. Every logit of a feature vector represents a meaningful attribute. SDM is carefully designed for fine-grained attribute supervision, outperforming common metrics such as Euclidean and Cosine metrics. Experimental results on 2 datasets show that CtF+OSF is not only$2\%$more accurate but also$5\times$faster than contemporary hashing ReID methods. Compared with non-hashing ReID methods, CtF is$50\times$faster with comparable accuracy. OSF further speeds CtF by$2\times$again and upto$10\times$in total with almost no accuracy drop.
Guan'an Wang, Xiaowen Huang 0001, Shaogang Gong, Jian Zhang 0018, Wen Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Knowledge Graph-Enhanced Sampling for Conversational Recommendation System
abstract
The traditional recommendation systems mainly use offline user data to train offline models, and then recommend items for online users, thus suffering from the unreliable estimation of user preferences based on sparse and noisy historical data. Conversational Recommendation System(CRS) uses the interactive form of the dialogue systems to solve the intrinsic problems of traditional recommendation systems. However, due to the lack of contextual information modeling, the existing CRS models are unable to deal with the exploitation and exploration(E&E) problem well, resulting in the heavy burden on users. To address the aforementioned issue, this work proposes a contextual information enhancement model tailored for CRS, called Knowledge Graph-enhanced Sampling(KGenSam). KGenSam integrates the dynamic graph of user interaction data with the external knowledge into one heterogeneous Knowledge Graph(KG) as the contextual information environment. Then, two samplers are designed to enhance knowledge by sampling fuzzy samples with high uncertainty for obtaining user preferences and reliable negative samples for updating recommender to achieve efficient acquisition of user preferences and model updating, and thus provide a powerful solution for CRS to deal with E&E problem. Experimental results on two real-world datasets demonstrate the superiority of KGenSam with significant improvements over state-of-the-art methods.
Xiaowen Huang 0001, Lixi Zhu, Jitao Sang 0001, Jian Yu 0001
IEEE Trans. Knowl. Data Eng.2
2022 Non-generative Generalized Zero-shot Learning via Task-correlated Disentanglement and Controllable Samples Synthesis
abstract
Synthesizing pseudo samples is currently the most effective way to solve the Generalized Zero Shot Learning (GZSL) problem. Most models achieve competitive performance but still suffer from two problems: (1) Feature confounding, the overall representations confound task-correlated and task-independent features, and existing models disentangle them in a generative way, but they are unreasonable to synthesize reliable pseudo samples with limited samples; (2) Distribution uncertainty, that massive data is needed when existing models synthesize samples from the uncertain distribution, which causes poor performance in limited samples of seen classes. In this paper, we propose a non-generative model to address these problems correspondingly in two modules: (1) Task-correlated feature disentanglement, to exclude the task-correlated features from task-independent ones by adversarial learning of domain adaption towards reasonable synthesis; (2) Controllable pseudo sample synthesis, to synthesize edge-pseudo and center-pseudo samples with certain characteristics towards more diversity generated and intuitive transfer. In addation, to describe the new scene that is the limit seen class samples in the training process, we further formulate a new ZSL task named the ‘Few-shot Seen class and Zero-shot Unseen class learning’ (FSZU). Extensive experiments on four benchmarks verify that the proposed method is competitive in the GZSL and the FSZU tasks.
Yaogong Feng, Xiaowen Huang 0001, Pengbo Yang, Jian Yu 0001, Jitao Sang 0001
CVPR2
2022 Learning to Learn a Cold-start Sequential Recommender
abstract
The cold-start recommendation is an urgent problem in contemporary online applications. It aims to provide users whose behaviors are literally sparse with as accurate recommendations as possible. Many data-driven algorithms, such as the widely used matrix factorization, underperform because of data sparseness. This work adopts the idea of meta-learning to solve the user’s cold-start recommendation problem. We propose a meta-learning-based cold-start sequential recommendation framework called metaCSR, including three main components: Diffusion Representer for learning better user/item embedding through information diffusion on the interaction graph; Sequential Recommender for capturing temporal dependencies of behavior sequences; and Meta Learner for extracting and propagating transferable knowledge of prior users and learning a good initialization for new users. metaCSR holds the ability to learn the common patterns from regular users’ behaviors and optimize the initialization so that the model can quickly adapt to new users after one or a few gradient updates to achieve optimal performance. The extensive quantitative experiments on three widely used datasets show the remarkable performance of metaCSR in dealing with the user cold-start problem. Meanwhile, a series of qualitative analysis demonstrates that the proposed metaCSR has good generalization.
Xiaowen Huang 0001, Jitao Sang 0001, Jian Yu 0001, Changsheng Xu
ACM Trans. Inf. Syst.1
2022 Image-Based Personality Questionnaire Design
abstract
This article explores the problem of image-based personality questionnaire design. Compared with the traditional text-based personality questionnaire, the image-based personality questionnaire is more natural, truthful, and language insensitive. Instead of responding to textual questions, the subjects are provided a set of “choose-your-favorite-image” visual questions. With each question, consisting of image options describing the same semantic concept, the subjects are requested to choose their favorite image. Based on responses to typically 15 to 25 questions, we can accurately estimate the subjects’ personality traits in five dimensions. The solution to design such an image-based personality questionnaire consists of concept-question identification and image-option selection. We have presented a preliminary framework to regularize these two steps in this exploratory study. A demo automatically adapting between desktop and mobile devices is available at http://120.27.209.14/vbfi . Subjective and objective evaluations have demonstrated the feasibility of accurately estimating a subject’s personality in a limited round of questions.
Xiaowen Huang 0001, Jitao Sang 0001, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.1
2021 Trustworthy Multimedia Analysis
abstract
This tutorial discusses the trustworthiness issue in multimedia analysis. Starting from introducing two types of spurious correlations learned from distilling human knowledge, we partition the (visual) feature space along two dimensions of task-relevance and semantic-orientation. Trustworthy multimedia analysis ideally relies on the task-relevant semantic features and consists of three modules as trainer, interpreter and tester. These three modules essentially form a closed loop, which respectively address goals of extracting task-relevant features, extracting task-relevant semantic features, and detecting spurious correlations to be corrected by the trainer and interpreter.
Xiaowen Huang 0001, Jiaming Zhang 0006, Yi Zhang 0101, Jitao Sang 0001
ACM Multimedia1
2021 APF: An Adversarial Privacy-preserving Filter to Protect Portrait Information
abstract
While widely adopted in practical applications, face recognition has been disputed on the malicious use of face images and potential privacy issues. Online photo sharing services accidentally act as the main approach for the malicious crawlers to exploit face recognition to access portrait privacy. In this demo, we propose an adversarial privacy-preserving filter, which can preserve face image from malicious face recognition algorithms. This filter is generated by an end-cloud collaborated adversarial attack framework consisting of three modules: (1) Image-specific gradient generation module, to extract image-specific gradient in the user end; (2) Adversarial gradient transfer module, to fine-tune the image-specific gradient in the server; and (3) Universal adversarial perturbation enhancement module, to append image-independent perturbation to derive the final adversarial perturbation. A short video about our system is available at https://github.com/Anonymity-for-submission/3247.
Jiaming Zhang 0006, Xiaowen Huang 0001
ACM Multimedia3
2020 Adversarial Privacy-preserving Filter
abstract
While widely adopted in practical applications, face recognition has been critically discussed regarding the malicious use of face images and the potential privacy problems, e.g., deceiving payment system and causing personal sabotage. Online photo sharing services unintentionally act as the main repository for malicious crawler and face recognition applications. This work aims to develop a privacy-preserving solution, called Adversarial Privacy-preserving Filter (APF), to protect the online shared face images from being maliciously used. We propose an end-cloud collaborated adversarial attack solution to satisfy requirements of privacy, utility and non-accessibility. Specifically, the solutions consist of three modules: (1) image-specific gradient generation, to extract image-specific gradient in the user end with a compressed probe model; (2) adversarial gradient transfer, to fine-tune the image-specific gradient in the server cloud; and (3) universal adversarial perturbation enhancement, to append image-independent perturbation to derive the final adversarial noise. Extensive experiments on three datasets validate the effectiveness and efficiency of the proposed solution. A prototype application is also released for further evaluation. We hope the end-cloud collaborated attack framework could shed light on addressing the issue of online multimedia sharing privacy-preserving issues from user side.
Jiaming Zhang 0006, Jitao Sang 0001, Xiaowen Huang 0001, Yongli Hu
ACM Multimedia4
2020 Meta-path Augmented Sequential Recommendation with Contextual Co-attention Network
abstract
It is critical to comprehensively and efficiently learn user preferences for an effective sequential recommender system. Existing sequential recommendation methods mainly focus on modeling local preference from users’ historical behaviors, which largely ignore the global context information from the heterogeneous information network. This prevents a comprehensive user preference representation. To address these issues, we propose a joint learning approach to incorporate global context with local preferences efficiently. The proposed approach introduces meta-paths from a heterogeneous information network to capture the global context information, and the position-based self-attention mechanism is adopted to model the local preference representation efficiently. Compared with the methods that only consider the local preference, our proposed method takes the advantages of incorporating global context information, which extracts structural features that captures relevant semantics to construct users’ global preference representation for the sequential recommendation. We further adopt a co-attention mechanism to model complex interactions between global context and users’ historical behaviors for better user representations. Quantitative and qualitative experimental evaluations are conducted on nine large-scale Amazon datasets and a multi-modal Zhihu dataset. The promising results demonstrate the effectiveness of the proposed model.
Xiaowen Huang 0001, Shengsheng Qian, Quan Fang, Jitao Sang 0001, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.1
2019 Explainable Interaction-driven User Modeling over Knowledge Graph for Sequential Recommendation
abstract
Compared with the traditional recommendation system, sequential recommendation holds the ability of capturing the evolution of users' dynamic interests. Many previous studies in sequential recommendation focus on the accuracy of predicting the next item that a user might interact with, while generally ignore providing explanations why the item is recommended to the user. Appropriate explanations are critical to help users adopt the recommended item, and thus improve the transparency and trustworthiness of the recommendation system. In this paper, we propose a novel Explainable Interaction-driven User Modeling (EIUM) algorithm to exploit Knowledge Graph (KG) for constructing an effective and explainable sequential recommender. Qualified semantic paths between specific user-item pair are extracted from KG. Encoding those semantic paths and learning the importance scores for each path provides the path-wise explanation for the recommendation system. Different from traditional item- level sequential modeling methods, we capture the interaction-level user dynamic preferences by modeling the sequential interactions. It is a high- level representation which contains auxiliary semantic information from KG. Furthermore, we adopt a joint learning manner for better representation learning by employing multi-modal fusion, which benefits from the structural constraints in KG and involves three kinds of modalities. Extensive experiments on the large-scale dataset show the better performance of our approach in making sequential recommendations in terms of both accuracy and explainability.
Xiaowen Huang 0001, Quan Fang, Shengsheng Qian, Jitao Sang 0001, Yan Li 0068, Changsheng Xu
ACM Multimedia1
2019 Biomedia ACM MM Grand Challenge 2019: Using Data Enhancement to Solve Sample Unbalance
abstract
The Biomedia ACM MM Grand Challenge focuses on medical applications with a task to detect and classify abnormalities within gastrointestinal (GI) tract. As a part of the submission for this challenge, several methods we applied are reported in this paper. The data we used is from the KVASIR dataset and the NEETHUS dataset. It contains training and test data in form of image or video. The main challenge of this task is the data's insufficiency and unbalance, which will significantly decrease the performance. To solve this problem and achieve better result, we conduct multiple data enhancement operations on the data. The method is proved to be efficient. Apart from the operations applied on the data, we also test several classification structures include SCNN (Shallow Convolutional Neural Network), SCNN-SVM (Support Vector Machine), ResNet32-SVM, SVM, ResNet16, ResNet32, ResNet50, ResNet101 and Residual Attention Network. We finally chose ResNet50 as the main structure considered with the balance of accuracy and efficiency. We obtain 91.51% precision, 87.45% sensitivity, 99.48% specificity, 87.93% F1-score (the harmonic mean of precision and sensitivity) and 91.40% MCC (Matthews correlation coefficient) on the test dataset.
Wenhua Meng, Xiaoshan Yang, Changsheng Xu, Xiaowen Huang 0001
ACM Multimedia6
2019 Multi-source User Attribute Inference based on Hierarchical Auto-encoder
abstract
With the rapid development of Online Social Networks (OSNs), it is crucial to construct users' portraits from their dynamic behaviors to address the increasing needs for customized information services. Previous work on user attribute inference mainly concentrated on developing advanced features/models or exploiting external information and knowledge but ignored the contradiction between dynamic behaviors and stable demographic attributes, which results in deviation of user understanding
Xiangguo Ding, Xiaowen Huang 0001, Jitao Sang 0001, Jian Yu 0001
MMAsia3
2018 CSAN: Contextual Self-Attention Network for User Sequential Recommendation
abstract
The sequential recommendation is an important task for online user-oriented services, such as purchasing products, watching videos, and social media consumption. Recent work usually used RNN-based methods to derive an overall embedding of the whole behavior sequence, which fails to discriminate the significance of individual user behaviors and thus decreases the recommendation performance. Besides, RNN-based encoding has fixed size and makes further recommendation application inefficient and inflexible. The online sequential behaviors of a user are generally heterogeneous, polysemous, and dynamically context-dependent. In this paper, we propose a unified Contextual Self-Attention Network (CSAN) to address the three properties. Heterogeneous user behaviors are considered in our model that are projected into a common latent semantic space. Then the output is fed into the feature-wise self-attention network to capture the polysemy of user behaviors. In addition, the forward and backward position encoding matrices are proposed to model dynamic contextual dependency. Through extensive experiments on two real-world datasets, we demonstrate the superior performance of the proposed model compared with other state-of-the-art algorithms.
Xiaowen Huang 0001, Shengsheng Qian, Quan Fang, Jitao Sang 0001, Changsheng Xu
ACM Multimedia1
2017 Towards SMP Challenge: Stacking of Diverse Models for Social Image Popularity Prediction
abstract
Popularity prediction on social media has attracted extensive attention nowadays due to its widespread applications, such as online marketing and economical trends. In this paper, we describe a solution of our team CASIA-NLPR-MMC for Social Media Prediction (SMP) challenge. This challenge is designed to predict the popularity of social media posts. We present a stacking framework by combining a diverse set of models to predict the popularity of images on Flickr using user-centered, image content and image context features. Several individual models are employed for scoring popularity of an image at earlier stage, and then a stacking model of Support Vector Regression (SVR) is utilized to train a meta model of different individual models trained beforehand. The Spearman's Rho of this Stacking model is 0.88 and the mean absolute error is about 0.75 on our test set. On the official final-released test set, the Spearman's Rho is 0.7927 and mean absolute error is about 1.1783. The results on provided dataset demonstrate the effectiveness of our proposed approach for image popularity prediction.
Xiaowen Huang 0001, Yuqi Gao, Quan Fang, Jitao Sang 0001, Changsheng Xu
ACM Multimedia1