VLDB 2026 Research / reviewers in the wild / expert
Jitao Sang 0001
dblp:84/286-1
· DBLP profile ↗
77ranked-venue papers
5as first author
45since 2021 · last 2026
0000-0002-0699-3205ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 59 · 3 first-author · 30 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 15 since 2021Databases, data management, data science and information retrieval · 15 · 2 first-author · 9 since 2021Computer networks · 8 · 5 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Membership Inference Attack Against Large Language Model-Based Recommendation Systems: A New Distillation-Based ParadigmabstractMembership Inference Attack (MIA) aims to determine whether a specific data sample was included in the training dataset of a target model. Traditional MIA approaches rely on shadow models to mimic target model behavior, but their effectiveness diminishes for Large Language Model (LLM)-based recommendation systems due to the scale and complexity of training data. This paper introduces a novel knowledge distillation-based MIA paradigm tailored for LLM-based recommendation systems. Our method constructs a reference model via distillation, applying distinct strategies for member and non-member data to enhance discriminative capabilities. The paradigm extracts fused features (e.g., confidence, entropy, loss, and hidden layer vectors) from the reference model to train an attack model, overcoming limitations of individual features. Extensive experiments on extended datasets (Last.FM, MovieLens, Book-Crossing, Delicious) and diverse LLMs (T5, GPT-2, LLaMA3) demonstrate that our approach significantly outperforms shadow model-based MIAs and individual-feature baselines. The results show its practicality for privacy attacks in LLM-driven recommender systems. Cuihong Li, Xiaowen Huang 0001, Chuanhuan Yin, Jitao Sang 0001 |
AAAI | 4 |
| 2026 | WebSynthesis: World Model-Guided Monte Carlo Tree Search for Efficient WebAgent Trajectory SynthesisabstractYifei Gao, Junhong Ye, Yifan Yang, Jiaqi Wang, Yi Zhang, Zhang Ruichen, Jitao Sang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junhong Ye, Jiaqi Wang 0003, Yi Zhang 0101, Jitao Sang 0001 |
ACL (1) | 7 |
| 2026 | VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language ModelsabstractOmni-modal large language models (OLLMs) offer a promising end-to-end solution for slideenhanced speech recognition due to their inherent multimodal capabilities.However, we found a fundamental issue faced by OLLMs: Visual Interference, where models show a bias towards visible text over auditory signals, causing them to hallucinate slide content that was never spoken.To address this, we propose Visually-Anchored Policy Optimization (VAPO), which aims to reshape models' inference process to follow the human-like "Look-then-Listen" inference chain.Specifically, we design a temporally decoupled policy: the model first extracts visual priors in a block to serve as semantic anchors, then generates the transcription in an block.The policy is optimized via multi-objective reinforcement learning.Furthermore, we introduce SlideASR-Bench, a comprehensive benchmark designed to address the scarcity of entity-rich data, comprising a large-scale synthetic corpus for training and a challenging real-world test set for evaluation.We conduct extensive evaluations demonstrating that VAPO effectively eliminates visual interference and achieves state-of-the-art performance on SlideASR-Bench and public datasets, significantly reducing entity recognition errors in specialized domains. Rui Hu 0011, Delai Qiu, Shengping Liu, Jitao Sang 0001 |
ACL (1) | 5 |
| 2026 | ITDR: An Instruction Tuning Dataset for Enhancing Large Language Models in RecommendationsabstractLarge language models (LLMs) have demonstrated outstanding performance in natural language processing tasks. However, in the field of recommender systems, due to the inherent structural discrepancy between user behavior data and natural language, LLMs struggle to effectively model the associations between user preferences and items. Although prompt-based methods can generate recommendation results, their inadequate understanding of recommendation tasks leads to constrained performance. To address this gap, we construct a comprehensive instruction tuning dataset, ITDR, which encompasses seven subtasks across two root tasks: user-item interaction and user-item understanding. The dataset integrates data from 13 public recommendation datasets and is built using manually crafted standardized templates, comprising approximately 200,000 instances. Experimental results demonstrate that ITDR significantly enhances the performance of mainstream open-source LLMs such as GLM-4, Qwen2.5, Qwen2.5-Instruct and LLaMA-3.2 on recommendation tasks. Furthermore, we analyze the correlations between tasks and explore the impact of task descriptions and data scale on instruction tuning effectiveness. Finally, we perform comparative experiments against closed-source LLMs with massive parameters. Our tuning dataset ITDR, the fine-tuned large recommendation models, all LoRA modules, and the complete experimental results are available at https://github.com/hellolzk/ITDR. Xiaowen Huang 0001, Jitao Sang 0001 |
KDD (1) | 3 |
| 2026 | NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language ModelsabstractVision-Language Models (VLMs) such as CLIP have demonstrated remarkable capabilities in understanding relationships between visual and textual data through joint embedding spaces. Despite their effectiveness, these models remain vulnerable to adversarial attacks, particularly in the image modality, posing significant security concerns. Building upon our previous work on Adversarial Prompt Tuning (AdvPT), which introduced learnable text prompts to enhance adversarial robustness in VLMs without extensive parameter training, we present a significant extension by introducing the Neural Augmentor framework for Multi-modal Adversarial Prompt Tuning (NAP-Tuning). As a significant extension, NAP-Tuning first establishes a comprehensive multi-modal (text and visual) and multi-layer prompting framework. The core of this framework is a targeted structural augmentation for feature-level purification, implemented through our Neural Augmentor approach. This framework implements feature purification by incorporating TokenRefiners-lightweight neural modules that learn to reconstruct purified features via residual connections-to directly address distortions in the feature space. This structural intervention is what enables the multi-modal and multi-layer system to effectively perform modality-specific and layer-specific feature rectification. Comprehensive experiments demonstrate that NAP-Tuning significantly outperforms existing methods across various datasets and attack types. Notably, our approach shows significant improvements over the strongest baselines under the challenging AutoAttack benchmark, outperforming them by 32.3% on ViT-B16 and 31.3% on ViT-B32 architectures while maintaining competitive clean accuracy. This work highlights the efficacy of internal feature-level intervention in prompt tuning for adversarial robustness, moving beyond input-side alignment approaches to create an adaptive defense mechanism that can identify and rectify adversarial perturbations across embedding spaces. Jiaming Zhang 0006, Xin Wang 0119, Xingjun Ma, Lingyu Qiu, Yu-Gang Jiang 0001, Jitao Sang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Inference-Time Rule Eraser: Fair Recognition via Distilling and Removing Biased RulesabstractMachine learning models often make predictions based on biased features such as gender, race, and other social attributes, posing significant fairness risks, especially in societal applications, such as hiring, banking, and criminal justice. Traditional approaches to addressing this issue involve retraining or fine-tuning neural networks with fairness-aware optimization objectives. However, these methods can be impractical due to significant computational resources, complex industrial tests, and the associated CO2 footprint. Additionally, regular users often fail to fine-tune models because they lack access to model parameters. In this paper, we introduce the Inference-Time Rule Eraser (Eraser), a novel method designed to address fairness concerns by removing biased decision-making rules from deployed models during inference without altering model weights. We begin by establishing a theoretical foundation for modifying model outputs to eliminate biased rules through Bayesian analysis. Next, we present a specific implementation of Eraser that involves two stages: (1) distilling the biased rules from the deployed model into an additional patch model, and (2) removing these biased rules from the output of the deployed model during inference. Extensive experiments validate the effectiveness of our approach, showcasing its superior performance in addressing fairness concerns in AI systems. Yi Zhang 0101, Dongyuan Lu, Jitao Sang 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | GAROD: Delve into Gradient-Based Attribution Reliability for Out-of-Distribution DetectionabstractThe deployment scenarios often include conditions not anticipated during training. Therefore, Out-of-Distribution (OOD) detection is essential for ensuring the reliability and security of neural networks. However, many existing OOD detectors suffer from instability, with performance degrading significantly when the dataset or model changes. This challenge highlights the need to approach OOD detection by examining intrinsic differences between In-Distribution (ID) and OOD samples in terms of model capacities, rather than relying on their observable characteristics. In this article, we propose Gradient-based Attribution Reliability for OOD Detection (GAROD), a novel method grounded in the capacity of invariance to irrelevant inputs, an important property linked to model generalization. We hypothesize that models exhibit such properties with ID samples, and samples for which the model lacks this invariance are classified as OOD. Specifically, GAROD leverages gradient-based attribution to separate relevant and irrelevant pixels in the input samples and observes how a model’s decisions change after removing irrelevant pixels. The approach most closely related to ours is attribution reliability evaluation (e.g., Insertion or Deletion metrics). However, these methods have never been applied to OOD detection. Moreover, directly using classical reliability metrics does not yield effective results. We identify two key issues: (1) model outputs are insufficient to capture decision changes effectively, and (2) using Insertion or Deletion metrics individually lacks comprehensiveness. In GAROD, we address these by observing final features instead, fusing both metrics to achieve robust OOD detection. Extensive experiments on CIFAR and ImageNet benchmarks demonstrate GAROD’s superiority over state-of-the-art post hoc methods, as well as its resilience to performance degradation under dataset/model variations. Code: https://github.com/iceshade000/GAROD . Guanhua Zheng, Jitao Sang 0001, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | ODE: Open-Set Evaluation of Hallucinations in Multimodal Large Language ModelsabstractHallucination poses a persistent challenge for multimodal large language models (MLLMs). However, existing benchmarks for evaluating hallucinations are generally static, which may overlook the potential risk of data contamination. To address this issue, we propose ODE, an openset, dynamic protocol designed to evaluate object hallucinations in MLLMs at both the existence and attribute levels. ODE employs a graph-based structure to represent real-world object concepts, their attributes, and the distributional associations between them. This structure facilitates the extraction of concept combinations based on diverse distributional criteria, generating varied samples for structured queries that evaluate hallucinations in both generative and discriminative tasks. Through the generation of new samples, dynamic concept combinations, and varied distribution frequencies, ODE mitigates the risk of data contamination and broadens the scope of evaluation. This protocol is applicable to both general and specialized scenarios, including those with limited data. Experimental results demonstrate the effectiveness of our protocol, revealing that MLLMs exhibit higher hallucination rates when evaluated with ODE-generated samples, which indicates potential data contamination. Furthermore, these generated samples aid in analyzing hallucination patterns and fine-tuning models, offering an effective approach to mitigating hallucinations in MLLMs. Our code are available at https://github.com/Iridescent-y/ODE. Yahan Tu, Rui Hu 0011, Jitao Sang 0001 |
CVPR | 3 |
| 2025 | Anyattack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language ModelsabstractDue to their multimodal capabilities, Vision-Language Models (VLMs) have found numerous impactful applications in real-world scenarios. However, recent studies have revealed that VLMs are vulnerable to image-based adversarial attacks. Traditional targeted adversarial attacks require specific targets and labels, limiting their real-world impact. We present AnyAttack, a self-supervised framework that transcends the limitations of conventional attacks through a novel foundation model approach. By pretraining on the massive LAION-400M dataset without label supervision, AnyAttack achieves unprecedented flexibility - enabling any image to be transformed into an attack vector targeting any desired output across different VLMs. This approach fundamentally changes the threat landscape, making adversarial capabilities accessible at an unprecedented scale. Our extensive validation across five open-source VLMs (CLIP, BLIP, BLIP2, InstructBLIP, and MiniGPT-4) demonstrates AnyAttack’s effectiveness across diverse multimodal tasks. Most concerning, Any-Attack seamlessly transfers to commercial systems including Google Gemini, Claude Sonnet, Microsoft Copilot and OpenAI GPT, revealing a systemic vulnerability requiring immediate attention. Jiaming Zhang 0006, Junhong Ye, Xingjun Ma, Yige Li, Yunfan Yang, Jitao Sang 0001, Dit-Yan Yeung |
CVPR | 7 |
| 2025 | Debiased Prompt Tuning for Vision-Language Models without AnnotationsabstractPrompt tuning of Vision-Language Models (VLMs) such as CLIP, has demonstrated the ability to rapidly adapt to various downstream tasks. However, recent studies indicate that tuned VLMs may suffer from the problem of spurious correlations, where the model relies on spurious features (e.g. background and gender) in the data. This may lead to the model having worse robustness in out-of-distribution data. Standard methods for eliminating spurious correlation typically require us to know the spurious attribute labels of each sample, which is hard in the real world. In this work, we explore improving the group robustness of prompt tuning in VLMs without relying on manual annotation of spurious features. We leverage the zero-shot image recognition ability of VLMs to identify spurious features, thus avoiding the cost of manual annotation. By leveraging pseudo-spurious attribute annotations, we further propose a method to automatically adjust the training weights of different groups. Extensive experiments show that our approach efficiently improves the worst-group accuracy on CelebA, Waterbirds, and MetaShift datasets, achieving the best robustness gap between the worst-group accuracy and the overall accuracy. Chaoquan Jiang, Yunfan Yang, Rui Hu 0011, Jitao Sang 0001 |
IJCNN | 4 |
| 2025 | SILLM4Rec: Self-Improving with Chain of Thought Enhanced Preference Optimization for Multimodal RecommendationabstractRecent explorations into the potential of large language models (LLMs) within recommendation systems have demonstrated promising performance. However, LLM-based recommenders still face significant challenges in complex scenarios. Most existing alignment approaches primarily focus on direct item generation based on user-interaction histories, frequently neglecting to fully exploit valuable feedback such as reviews and ratings. Furthermore, these methods rarely consider the benefits of incorporating reasoning mechanisms to enhance recommendation accuracy. To overcome these limitations and improve the reliability of LLM-based recommenders under data-scarce conditions, we propose SILLM4Rec, a framework specifically designed to strengthen both reasoning ability and recommendation performance. Our approach begins by extracting high-quality Chain-of-Thought (CoT) reasoning samples from a more capable teacher model. These samples are then used to fine-tune a smaller student model through a self-learning and iterative optimization process, enabling it to refine its outputs and adapt to user preferences. Comprehensive experiments on three publicly available benchmark datasets demonstrate that SILLM4Rec not only achieves superior performance metrics, but also enhances transparency and robustness in recommendation tasks. Our code is available at https://github.com/MKC-Lab/SILLM4Rec. Fang Quan, Xiaowen Huang 0001, Jitao Sang 0001 |
MMAsia | 5 |
| 2025 | A LLM-based Controllable, Scalable, Human-Involved User Simulator Framework for Conversational Recommender SystemsabstractConversational Recommender System (CRS) leverages real-time feedback from users to dynamically model their preferences, thereby enhancing the system's ability to provide personalized recommendations and improving the overall user experience. CRS has demonstrated significant promise, prompting researchers to concentrate their efforts on developing user simulators that are both more realistic and trustworthy. The advent of Large Language Models (LLMs) has demonstrated capabilities that approach human-level intelligence across a diverse range of tasks. Research efforts have been made to utilize LLMs for building user simulators to evaluate the performance of CRS. Although these efforts showcase innovation, they are accompanied by certain limitations. In this work, we introduce a Controllable, Scalable, and Human-Involved (CSHI) simulator framework that manages the behavior of user simulators across various stages via a plugin manager. CSHI tailors behavioral simulations and interaction patterns to deliver authentic user-system engagement experiences. Through experiments and case studies in two conversational recommendation scenarios, we show that our framework can adapt to a variety of conversational recommendation settings and effectively simulate users' personalized preferences. Consequently, our simulator is able to generate feedback that closely mirrors that of real users. This facilitates a reliable assessment of existing CRS studies and promotes the creation of high-quality conversational recommendation datasets. Lixi Zhu, Xiaowen Huang 0001, Jitao Sang 0001 |
WWW | 3 |
| 2025 | Prescribing the right remedy: Mitigating hallucinations in large vision-language models via targeted instruction tuning
Rui Hu 0011, Yahan Tu, Shuyu Wei, Dongyuan Lu, Jitao Sang 0001 |
Inf. Sci. | 5 |
| 2025 | Backdoor for Debias: Mitigating Model Bias With Backdoor Attack-Based Artificial BiasabstractWith the swift advancement of deep learning, state-of-the-art algorithms have been utilized in various social situations. Nonetheless, some algorithms have been discovered to exhibit biases and provide unequal results. The current debiasing methods face challenges such as poor utilization of data or intricate training requirements. In this work, we found that the backdoor attack can construct an artificial bias similar to the model bias derived in standard training. Considering the strong adjustability of backdoor triggers, we are motivated to mitigate the model bias by carefully designing reverse artificial bias created from backdoor attack. Based on this, we propose a backdoor debiasing framework based on knowledge distillation, which effectively reduces the model bias from original data and minimizes security risks from the backdoor attack. The proposed solution is validated on both image and structured datasets, showing promising results. This work advances the understanding of backdoor attacks and highlights its potential for beneficial applications. The code for the study can be found athttps://github.com/KirinNg/DwB. Shangxi Wu, Qiuyang He, Jian Yu 0001, Jitao Sang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | A Disguised Wolf Is More Harmful Than a Toothless Tiger: Adaptive and Malicious Code Injection Backdoor Attack Leveraging User Behavior as TriggersabstractIn recent years, large language models (LLMs) have made significant progress in code generation. However, as these models are increasingly adopted for software development, their associated security risks have become more pronounced. Studies have shown that traditional deep learning robustness issues also adversely affect the reliability of code generation. In this paper, we use game theory to systematically examine security vulnerabilities in code generation and illustrate how attackers can propagate malicious models to create genuine threats. We also demonstrate, for the first time, that attackers can leverage user behavior as a trigger for backdoor attacks—dynamically controlling when malicious code is injected—and calibrate these attacks to a user’s skill level, leading to varying degrees of impact. Through extensive experiments on leading code generation models, we verify that these security threats are both feasible and dangerous. Our research code will be available at https://github.com/KirinNg/Adaptive_Malicious_Code_Injection_Backdoor_Attack. Shangxi Wu, Jinlin Xiao, Jitao Sang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | MF-CLIP: Leveraging CLIP as Surrogate Models for No-Box Adversarial Attacks
Jiaming Zhang 0006, Lingyu Qiu, Qi Yi, Yige Li, Jitao Sang 0001, Changsheng Xu, Dit-Yan Yeung |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Towards Robust Recommendation: A Review and an Adversarial Robustness Evaluation LibraryabstractRecently, recommender system has achieved significant success. However, due to the openness of recommender systems, they remain vulnerable to malicious attacks. Additionally, natural noise in training data and issues such as data sparsity can also degrade the performance of recommender systems. Therefore, enhancing the robustness of recommender systems has become an increasingly important research topic. In this survey, we provide a comprehensive overview of the robustness of recommender systems. Based on our investigation, we categorize the robustness of recommender systems into adversarial robustness and non-adversarial robustness. In the adversarial robustness, we introduce the fundamental principles and classical methods of recommender system adversarial attacks and defenses. In the non-adversarial robustness, we analyze nonadversarial robustness from the perspectives of data sparsity, natural noise, and data imbalance. Additionally, we summarize commonly used datasets and evaluation metrics for evaluating the robustness of recommender systems. Finally, we also discuss the current challenges in the field of recommender system robustness and potential future research directions. Additionally, to facilitate fair and efficient evaluation of attack and defense methods in adversarial robustness, we propose an adversarial robustness evaluation library–ShillingREC, and we conduct evaluations of basic attack models and recommendation models. ShillingREC project is released at https://github.com/ chengleileilei/ShillingREC. Xiaowen Huang 0001, Jitao Sang 0001, Jian Yu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Adversarial Prompt Tuning for Vision-Language Models
Jiaming Zhang 0006, Xingjun Ma, Xin Wang 0119, Lingyu Qiu, Jiaqi Wang 0003, Yu-Gang Jiang 0001, Jitao Sang 0001 |
ECCV (45) | 7 |
| 2024 | AIGCs Confuse AI Too: Investigating and Explaining Synthetic Image-induced Hallucinations in Large Vision-Language ModelsabstractThe evolution of Artificial Intelligence Generated Contents (AIGCs) is advancing towards higher quality. The growing interactions with AIGCs present a new challenge to the data-driven AI community: While AI-generated contents have played a crucial role in a wide range of AI models, the potential hidden risks they introduce have not been thoroughly examined. Beyond human-oriented forgery detection, AI-generated content poses potential issues for AI models originally designed to process natural data. In this study, we underscore the exacerbated hallucination phenomena in Large Vision-Language Models (LVLMs) caused by AI-synthetic images. Remarkably, our findings shed light on a consistent AIGC hallucination bias: the object hallucinations induced by synthetic images are characterized by a greater quantity and a more uniform position distribution, even these synthetic images do not manifest unrealistic or additional relevant visual features compared to natural images. Moreover, our investigations on Q-former and Linear projector reveal that synthetic images may present token deviations after visual projection, thereby amplifying the hallucination bias. Jiaqi Wang 0003, Jitao Sang 0001 |
ACM Multimedia | 4 |
| 2024 | Poisoning for Debiasing: Fair Recognition via Eliminating Bias Uncovered in Data PoisoningabstractNeural networks often tend to rely on bias features that have strong but spurious correlations with the target labels for decision-making, leading to poor performance on data that does not adhere to these correlations. Early debiasing methods typically construct an unbiased optimization objective based on the labels of bias features. Recent work assumes that bias label is unavailable and usually trains two models: a biased model to deliberately learn bias features for exposing data bias, and a target model to eliminate bias captured by the bias model. In this paper, we first reveal that previous biased models fit target labels, which resulted in failing to expose data bias. To tackle this issue, we propose poisoner, which utilizes data poisoning to embed the biases learned by biased models into the poisoned training data, thereby encouraging the models to learn more biases. Specifically, we couple data poisoning and model training to continuously prompt the biased model to learn more bias. By utilizing the biased model, we can identify samples in the data that contradict these biased correlations. Subsequently, we amplify the influence of these samples in the training of the target model to prevent the model from learning such biased correlations. Experiments show the superior debiasing performance of our method. Yi Zhang 0101, Zhefeng Wang 0001, Rui Hu 0011, Xinyu Duan, Yi Zheng 0007, Baoxing Huai, Jiarun Han, Jitao Sang 0001 |
ACM Multimedia | 8 |
| 2024 | Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationabstractMobile device operation tasks are increasingly becoming a popular multi-modal AI application scenario. Current Multi-modal Large Language Models (MLLMs), constrained by their training data, lack the capability to function effectively as operation assistants. Instead, MLLM-based agents, which enhance capabilities through tool invocation, are gradually being applied to this scenario. However, the two major navigation challenges in mobile device operation tasks — task progress navigation and focus content navigation — are difficult to effectively solve under the single-agent architecture of existing work. This is due to the overly long token sequences and the interleaved text-image data format, which limit performance. To address these navigation challenges effectively, we propose Mobile-Agent-v2, a multi-agent architecture for mobile device operation assistance. The architecture comprises three agents: planning agent, decision agent, and reflection agent. The planning agent condenses lengthy, interleaved image-text history operations and screens summaries into a pure-text task progress, which is then passed on to the decision agent. This reduction in context length makes it easier for decision agent to navigate the task progress. To retain focus content, we design a memory unit that updates with task progress by decision agent. Additionally, to correct erroneous operations, the reflection agent observes the outcomes of each operation and handles any mistake accordingly. Experimental results indicate that Mobile-Agent-v2 achieves over a 30% improvement in task completion compared to the single-agent architecture of Mobile-Agent. The code is open-sourced at https://github.com/X-PLUG/MobileAgent. Junyang Wang 0001, Haiyang Xu 0001, Haitao Jia, Ming Yan 0008, Weizhou Shen, Ji Zhang 0011, Fei Huang 0002, Jitao Sang 0001 |
NeurIPS | 9 |
| 2024 | TIF: Threshold Interception and Fusion for Compact and Fine-Grained Visual AttributionabstractThe blackbox nature of deep models has prompted a growing interest in explaining their inner workings and decision-making processes. Although backpropagation (BP)-based attribution methods are popular for visual interpretation, existing methods frequently yield implausible outcomes. For example, gradient-based attribution methods tend to highlight irrelevant regions and generate noise, while CAM-based attributions suffer from low resolution and blurry results. These limitations undermine their ability to correctly identify the target objects and fail to provide the desired justification to guarantee the credibility of the model's decision-making. In this article, we analyze plausibility issues in the frequency domain and point out that the plausibility issues correspond to frequency-domain incompleteness, i.e., the frequency-domain representation of explanations lacks low- or high-frequency components. Then, we propose a straightforward yet effective approach, threshold interception and fusion (TIF), to address this issue by fusing multilayer attributions. Our strategy involves collecting attribution results for all neurons and dividing the attribution map into a concept region that represents the current neuron and background regions based on a given threshold value$\alpha$. We then fuse these concept regions with the neuron weights in each layer and upsample the layer attributions to match the input size. Finally, we obtain the overall attribution by summing the layer attributions pixelwise. Our experiments demonstrate TIF efficacy by consistently enhancing visual performance across a variety of gradient-based attributions. To further demonstrate the ability to provide compact and fine-grained target objects, we directly employ TIF for the weakly supervised semantic segmentation task. Our results illustrate that TIF significantly outperforms existing methods without additional supervision or architectural modifications. We also observe an overall TIF improvement in the fidelity metric, suggesting that compactness and fine-graininess are not only plausibility issues but also fidelity issues. Guanhua Zheng, Jitao Sang 0001, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2024 | Adaptive Adversarial Logits PairingabstractAdversarial examples provide an opportunity as well as impose a challenge for understanding image classification systems. Based on the analysis of the adversarial training solution—Adversarial Logits Pairing (ALP), we observed in this work that: (1) The inference of adversarially robust model tends to rely on fewer high-contribution features compared with vulnerable ones. (2) The training target of ALP does not fit well to a noticeable part of samples, where the logits pairing loss is overemphasized and obstructs minimizing the classification loss. Motivated by these observations, we design an Adaptive Adversarial Logits Pairing (AALP) solution by modifying the training process and training target of ALP. Specifically, AALP consists of an adaptive feature optimization module with Guided Dropout to systematically pursue fewer high-contribution features, and an adaptive sample weighting module by setting sample-specific training weights to balance between logits pairing loss and classification loss. The proposed AALP solution demonstrates superior defense performance on multiple datasets with extensive experiments. Shangxi Wu, Jitao Sang 0001, Kaiyan Xu, Guanhua Zheng, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | ImageNet Pre-training Also Transfers Non-robustnessabstractImageNet pre-training has enabled state-of-the-art results on many tasks. In spite of its recognized contribution to generalization, we observed in this study that ImageNet pre-training also transfers adversarial non-robustness from pre-trained model into fine-tuned model in the downstream classification tasks. We first conducted experiments on various datasets and network backbones to uncover the adversarial non-robustness in fine-tuned model. Further analysis was conducted on examining the learned knowledge of fine-tuned model and standard model, and revealed that the reason leading to the non-robustness is the non-robust features transferred from ImageNet pre-trained model. Finally, we analyzed the preference for feature learning of the pre-trained model, explored the factors influencing robustness, and introduced a simple robust ImageNet pre-training solution. Our code is available at https://github.com/jiamingzhang94/ImageNet-Pretraining-transfers-non-robustness. Jiaming Zhang 0006, Jitao Sang 0001, Qi Yi, Yunfan Yang, Huiwen Dong, Jian Yu 0001 |
AAAI | 2 |
| 2023 | Unlearnable Clusters: Towards Label-Agnostic Unlearnable ExamplesabstractThere is a growing interest in developing unlearnable examples (UEs) against visual privacy leaks on the Internet. UEs are training samples added with invisible but unlearnable noise, which have been found can prevent unauthorized training of machine learning models. UEs typically are generated via a bilevel optimization framework with a surrogate model to remove (minimize) errors from the original samples, and then applied to protect the data against unknown target models. However, existing UE generation methods all rely on an ideal assumption called label-consistency, where the hackers and protectors are assumed to hold the same label for a given sample. In this work, we propose and promote a more practical label-agnostic setting, where the hackers may exploit the protected data quite differently from the protectors. E.g., amclass unlearnable dataset held by the protector may be exploited by the hacker as a n-class dataset. Existing UE generation methods are rendered ineffective in this challenging setting. To tackle this challenge, we present a novel technique called Unlearnable Clusters (UCs) to generate label-agnostic unlearnable examples with cluster-wise perturbations. Furthermore, we propose to leverage Vision-and-Language Pre-trained Models (VLPMs) like CLIP as the surrogate model to improve the transferability of the crafted UCs to diverse domains. We empirically verify the effectiveness of our proposed approach under a variety of settings with different datasets, target models, and even commercial platforms Microsoft Azure and Baidu PaddlePaddle. Code is available at https://github.com/jiamingzhang94/Unlearnable-Clusters. Jiaming Zhang 0006, Xingjun Ma, Qi Yi, Jitao Sang 0001, Yu-Gang Jiang 0001, Yaowei Wang 0001, Changsheng Xu |
CVPR | 4 |
| 2023 | Improved Visual Fine-tuning with Natural Language SupervisionabstractFine-tuning a visual pre-trained model can leverage the semantic information from large-scale pre-training data and mitigate the over-fitting problem on downstream vision tasks with limited training examples. While the problem of catastrophic forgetting in pre-trained backbone has been extensively studied for fine-tuning, its potential bias from the corresponding pre-training task and data, attracts less attention. In this work, we investigate this problem by demonstrating that the obtained classifier after fine-tuning will be close to that induced by the pre-trained model. To reduce the bias in the classifier effectively, we introduce a reference distribution obtained from a fixed text classifier, which can help regularize the learned vision classifier. The proposed method, Text Supervised fine-tuning (TeS), is evaluated with diverse pre-trained vision models including ResNet and ViT, and text encoders including BERT and CLIP, on 11 downstream tasks. The consistent improvement with a clear margin over distinct scenarios confirms the effectiveness of our proposal. Code is available at https://github.com/idstcv/TeS. Junyang Wang 0001, Yuanhong Xu, Juhua Hu, Ming Yan 0008, Jitao Sang 0001, Qi Qian 0001 |
ICCV | 5 |
| 2023 | From Association to Generation: Text-only Captioning by Unsupervised Cross-modal MappingabstractWith the development of Vision-Language Pre-training Models (VLPMs) represented by CLIP and ALIGN, significant breakthroughs have been achieved for association-based visual tasks such as image classification and image-text retrieval by the zero-shot capability of CLIP without fine-tuning. However, CLIP is hard to apply to generation-based tasks. This is due to the lack of decoder architecture and pre-training tasks for generation. Although previous works have created generation capacity for CLIP through additional language models, a modality gap between the CLIP representations of different modalities and the inability of CLIP to model the offset of this gap, which results in the failure of the concept to transfer across modes. To solve the problem, we try to map images/videos to the language modality and generate captions from the language modality. In this paper, we propose the K-nearest-neighbor Cross-modality Mapping (Knight), a zero-shot method from association to generation. With vision-free unsupervised training, Knight achieves state-of-the-art performance in zero-shot methods for image captioning and video captioning. Junyang Wang 0001, Ming Yan 0008, Yi Zhang 0101, Jitao Sang 0001 |
IJCAI | 4 |
| 2023 | Echoes: Unsupervised Debiasing via Pseudo-bias Labeling in an Echo ChamberabstractNeural networks often learn spurious correlations when exposed to biased training data, leading to poor performance on out-of-distribution data. A biased dataset can be divided, according to biased features, into bias-aligned samples (i.e., with biased features) and bias-conflicting samples (i.e., without biased features). Recent debiasing works typically assume that no bias label is available during the training phase, as obtaining such information is challenging and labor-intensive. Following this unsupervised assumption, existing methods usually train two models: a biased model specialized to learn biased features and a target model that uses information from the biased model for debiasing. This paper first presents experimental analyses revealing that the existing biased models overfit to bias-conflicting samples in the training data, which negatively impacts the debiasing performance of the target models. To address this issue, we propose a straightforward and effective method called Echoes, which trains a biased model and a target model with a different strategy. We construct an "echo chamber" environment by reducing the weights of samples which are misclassified by the biased model, to ensure the biased model fully learns the biased features without overfitting to the bias-conflicting samples. The biased model then assigns lower weights on the bias-conflicting samples. Subsequently, we use the inverse of the sample weights of the biased model for training the target model. Experiments show that our approach achieves superior debiasing results compared to the existing baselines on both synthetic and real-world datasets. Our code is available at https://github.com/isruihu/Echoes. Rui Hu 0011, Yahan Tu, Jitao Sang 0001 |
ACM Multimedia | 3 |
| 2023 | mPLUG-Octopus: The Versatile Assistant Empowered by A Modularized End-to-End Multimodal LLMabstractInspired by the recent developments of large language models (LLMs), we propose mPLUG-Octopus, a versatile conversational assistant designed to provide users with coherent, engaging, and helpful interaction experiences in both text-only and multi-modal scenarios. Unlike traditional pipeline chatting systems, mPLUG-Octopus offers a diverse range of creative capabilities including open-domain QA, multi-turn chatting, and multi-modal creation, all built with a unified multimodal LLM without relying on any external API. With the modularized end-to-end multimodal LLM technology, mPLUG-Octopus efficiently facilitates engaging and open-domain conversation experience. It exhibits a wide range of uni/multi-modal elemental capabilities, enabling it to seamlessly communicate with users on open-domain topics and engage in multi-turn conversations. It also assists users in accomplishing various content creation and application tasks. Our conversational assistant can also be deployed on smart hardware to drive advanced AIGC applications. Qinghao Ye, Haiyang Xu 0001, Ming Yan 0008, Chenlin Zhao, Junyang Wang 0001, Xiaoshan Yang, Ji Zhang 0011, Fei Huang 0002, Jitao Sang 0001, Changsheng Xu |
ACM Multimedia | 9 |
| 2023 | Benign Shortcut for Debiasing: Fair Visual Recognition via Intervention with Shortcut FeaturesabstractMachine learning models often learn to make predictions that rely on sensitive social attributes like gender and race, which poses significant fairness risks, especially in societal applications, such as hiring, banking, and criminal justice. Existing work tackles this issue by minimizing the employed information about social attributes in models for debiasing. However, the high correlation between target task and these social attributes makes learning on the target task incompatible with debiasing. Given that model bias arises due to the learning of bias features (i.e., gender) that help target task optimization, we explore the following research question: Can we leverage shortcut features to replace the role of bias feature in target task optimization for debiasing? To this end, we propose Shortcut Debiasing, to first transfer the target task's learning of bias attributes from bias features to shortcut features, and then employ causal intervention to eliminate shortcut features during inference. The key idea of Shortcut Debiasing is to design controllable shortcut features to on one hand replace bias features in contributing to the target task during the training stage, and on the other hand be easily removed by intervention during the inference stage. This guarantees the learning of the target task does not hinder the elimination of bias features. We apply Shortcut Debiasing to several benchmark datasets, and achieve significant improvements over the state-of-the-art debiasing methods in both accuracy and fairness. Yi Zhang 0101, Jitao Sang 0001, Junyang Wang 0001, Dongmei Jiang, Yaowei Wang 0001 |
ACM Multimedia | 2 |
| 2023 | Debiasing backdoor attack: A benign application of backdoor attack in eliminating data bias
Shangxi Wu, Qiuyang He, Yi Zhang 0101, Dongyuan Lu, Jitao Sang 0001 |
Inf. Sci. | 5 |
| 2023 | Low-mid adversarial perturbation against unauthorized face recognition system
Jiaming Zhang 0006, Qi Yi, Dongyuan Lu, Jitao Sang 0001 |
Inf. Sci. | 4 |
| 2023 | Towards a multimodal human activity dataset for healthcare
Menghao Hu, Mingxuan Luo, Menghua Huang, Wenhua Meng, Baochen Xiong, Xiaoshan Yang, Jitao Sang 0001 |
Multim. Syst. | 7 |
| 2023 | Knowledge Graph-Enhanced Sampling for Conversational Recommendation SystemabstractThe traditional recommendation systems mainly use offline user data to train offline models, and then recommend items for online users, thus suffering from the unreliable estimation of user preferences based on sparse and noisy historical data. Conversational Recommendation System(CRS) uses the interactive form of the dialogue systems to solve the intrinsic problems of traditional recommendation systems. However, due to the lack of contextual information modeling, the existing CRS models are unable to deal with the exploitation and exploration(E&E) problem well, resulting in the heavy burden on users. To address the aforementioned issue, this work proposes a contextual information enhancement model tailored for CRS, called Knowledge Graph-enhanced Sampling(KGenSam). KGenSam integrates the dynamic graph of user interaction data with the external knowledge into one heterogeneous Knowledge Graph(KG) as the contextual information environment. Then, two samplers are designed to enhance knowledge by sampling fuzzy samples with high uncertainty for obtaining user preferences and reliable negative samples for updating recommender to achieve efficient acquisition of user preferences and model updating, and thus provide a powerful solution for CRS to deal with E&E problem. Experimental results on two real-world datasets demonstrate the superiority of KGenSam with significant improvements over state-of-the-art methods. Xiaowen Huang 0001, Lixi Zhu, Jitao Sang 0001, Jian Yu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Attention, Please! Adversarial Defense via Activation Rectification and PreservationabstractThis study provides a new understanding of the adversarial attack problem by examining the correlation between adversarial attack and visual attention change. In particular, we observed that: (1) images with incomplete attention regions are more vulnerable to adversarial attacks; and (2) successful adversarial attacks lead to deviated and scattered activation map. Therefore, we use the mask method to design an attention-preserving loss and a contrast method to design a loss that makes the model’s attention rectification. Accordingly, an attention-based adversarial defense framework is designed, under which better adversarial training or stronger adversarial attacks can be performed through the above constraints. We hope the attention-related data analysis and defense solution in this study will shed some light on the mechanism behind the adversarial attack and also facilitate future adversarial defense/attack model design. Shangxi Wu, Jitao Sang 0001, Jiaming Zhang 0006, Jian Yu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Non-generative Generalized Zero-shot Learning via Task-correlated Disentanglement and Controllable Samples SynthesisabstractSynthesizing pseudo samples is currently the most effective way to solve the Generalized Zero Shot Learning (GZSL) problem. Most models achieve competitive performance but still suffer from two problems: (1) Feature confounding, the overall representations confound task-correlated and task-independent features, and existing models disentangle them in a generative way, but they are unreasonable to synthesize reliable pseudo samples with limited samples; (2) Distribution uncertainty, that massive data is needed when existing models synthesize samples from the uncertain distribution, which causes poor performance in limited samples of seen classes. In this paper, we propose a non-generative model to address these problems correspondingly in two modules: (1) Task-correlated feature disentanglement, to exclude the task-correlated features from task-independent ones by adversarial learning of domain adaption towards reasonable synthesis; (2) Controllable pseudo sample synthesis, to synthesize edge-pseudo and center-pseudo samples with certain characteristics towards more diversity generated and intuitive transfer. In addation, to describe the new scene that is the limit seen class samples in the training process, we further formulate a new ZSL task named the ‘Few-shot Seen class and Zero-shot Unseen class learning’ (FSZU). Extensive experiments on four benchmarks verify that the proposed method is competitive in the GZSL and the FSZU tasks. Yaogong Feng, Xiaowen Huang 0001, Pengbo Yang, Jian Yu 0001, Jitao Sang 0001 |
CVPR | 5 |
| 2022 | Benign Adversarial Attack: Tricking Models for GoodnessabstractIn spite of the successful application in many fields, machine learning models today suffer from notorious problems like vulnerability to adversarial examples. Beyond falling into the cat-and-mouse game between adversarial attack and defense, this paper provides alternative perspective to consider adversarial example and explore whether we can exploit it in benign applications. We first attribute adversarial example to the human-model disparity on employing non-semantic features. While largely ignored in classical machine learning mechanisms, non-semantic feature enjoys three interesting characteristics as (1) exclusive to model, (2) critical to affect inference, and (3) utilizable as features. Inspired by this, we present brave new idea of benign adversarial attack to exploit adversarial examples for goodness in three directions: (1) adversarial Turing test, (2) rejecting malicious model application, and (3) adversarial data augmentation. Each direction is positioned with motivation elaboration, justification analysis and prototype applications to showcase its potential. Jitao Sang 0001, Jiaming Zhang 0006 |
ACM Multimedia | 1 |
| 2022 | Counterfactually Measuring and Eliminating Social Bias in Vision-Language Pre-training ModelsabstractVision-Language Pre-training (VLP) models have achieved state-of-the-art performance in numerous cross-modal tasks. Since they are optimized to capture the statistical properties of intra- and inter-modality, there remains risk to learn social biases presented in the data as well. In this work, we (1) introduce a counterfactual-based bias measurement CounterBias to quantify the social bias in VLP models by comparing the [MASK]ed prediction probabilities of factual and counterfactual samples; (2) construct a novel VL-Bias dataset including 24K image-text pairs for measuring gender bias in VLP models, from which we observed that significant gender bias is prevalent in VLP models; and (3) propose a VLP debiasing method FairVLP to minimize the difference in the [MASK]ed prediction probabilities between factual and counterfactual image-text pairs for VLP debiasing. Although CounterBias and FairVLP focus on social bias, they are generalizable to serve as tools and provide new insights to probe and regularize more knowledge in VLP models. Yi Zhang 0101, Junyang Wang 0001, Jitao Sang 0001 |
ACM Multimedia | 3 |
| 2022 | Towards Adversarial Attack on Vision-Language Pre-training ModelsabstractWhile vision-language pre-training model (VLP) has shown revolutionary improvements on various vision-language (V+L) tasks, the studies regarding its adversarial robustness remain largely unexplored. This paper studied the adversarial attack on popular VLP models and V+L tasks. First, we analyzed the performance of adversarial attacks under different settings. By examining the influence of different perturbed objects and attack targets, we concluded some key observations as guidance on both designing strong multimodal adversarial attack and constructing robust VLP models. Second, we proposed a novel multimodal attack method on the VLP models called Collaborative Multimodal Adversarial Attack (Co-Attack), which collectively carries out the attacks on the image modality and the text modality. Experimental results demonstrated that the proposed method achieves improved attack performances on different V+L downstream tasks and VLP models. The analysis observations and novel attack method hopefully provide new understanding into the adversarial robustness of VLP models, so as to contribute their safe and reliable deployment in more real-world scenarios. Jiaming Zhang 0006, Qi Yi, Jitao Sang 0001 |
ACM Multimedia | 3 |
| 2022 | Learning to Learn a Cold-start Sequential RecommenderabstractThe cold-start recommendation is an urgent problem in contemporary online applications. It aims to provide users whose behaviors are literally sparse with as accurate recommendations as possible. Many data-driven algorithms, such as the widely used matrix factorization, underperform because of data sparseness. This work adopts the idea of meta-learning to solve the user’s cold-start recommendation problem. We propose a meta-learning-based cold-start sequential recommendation framework called metaCSR, including three main components: Diffusion Representer for learning better user/item embedding through information diffusion on the interaction graph; Sequential Recommender for capturing temporal dependencies of behavior sequences; and Meta Learner for extracting and propagating transferable knowledge of prior users and learning a good initialization for new users. metaCSR holds the ability to learn the common patterns from regular users’ behaviors and optimize the initialization so that the model can quickly adapt to new users after one or a few gradient updates to achieve optimal performance. The extensive quantitative experiments on three widely used datasets show the remarkable performance of metaCSR in dealing with the user cold-start problem. Meanwhile, a series of qualitative analysis demonstrates that the proposed metaCSR has good generalization. Xiaowen Huang 0001, Jitao Sang 0001, Jian Yu 0001, Changsheng Xu |
ACM Trans. Inf. Syst. | 2 |
| 2022 | Image-Based Personality Questionnaire DesignabstractThis article explores the problem of image-based personality questionnaire design. Compared with the traditional text-based personality questionnaire, the image-based personality questionnaire is more natural, truthful, and language insensitive. Instead of responding to textual questions, the subjects are provided a set of “choose-your-favorite-image” visual questions. With each question, consisting of image options describing the same semantic concept, the subjects are requested to choose their favorite image. Based on responses to typically 15 to 25 questions, we can accurately estimate the subjects’ personality traits in five dimensions. The solution to design such an image-based personality questionnaire consists of concept-question identification and image-option selection. We have presented a preliminary framework to regularize these two steps in this exploratory study. A demo automatically adapting between desktop and mobile devices is available at http://120.27.209.14/vbfi . Subjective and objective evaluations have demonstrated the feasibility of accurately estimating a subject’s personality in a limited round of questions. Xiaowen Huang 0001, Jitao Sang 0001, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Trustworthy Multimedia AnalysisabstractThis tutorial discusses the trustworthiness issue in multimedia analysis. Starting from introducing two types of spurious correlations learned from distilling human knowledge, we partition the (visual) feature space along two dimensions of task-relevance and semantic-orientation. Trustworthy multimedia analysis ideally relies on the task-relevant semantic features and consists of three modules as trainer, interpreter and tester. These three modules essentially form a closed loop, which respectively address goals of extracting task-relevant features, extracting task-relevant semantic features, and detecting spurious correlations to be corrected by the trainer and interpreter. Xiaowen Huang 0001, Jiaming Zhang 0006, Yi Zhang 0101, Jitao Sang 0001 |
ACM Multimedia | 5 |
| 2021 | Metadata Connector: Exploiting Hashtag and Tag for Cross-OSN Event SearchabstractSocial media has revolutionized the way people understand and keep track of real-world events. Various related multimedia information in different modalities such as texts, images and videos is updated on social media and reflects the events. These quantities of information distributes on different Online Social Networks (OSNs), which provides rich, wide coverage, comprehensive information about the trending events. Faced with such large amounts of data, searching has become a handy tool for event understanding and tracking on social media. However, existing single-OSN search mainly involves with single modality on single platform. Moreover, most OSNs usually focus on biased perspective of events, which significantly limits the coverage and diversity of single-OSN based event search. In this paper, we introduce a novel cross-OSN framework to help integrate these cross-OSN information regarding the same event and provide an immersive experience for information retrieval. Since social media information is widely distributed in different OSNs where semantic gap exists among these heterogeneous spaces, we propose to utilize hashtag and tag, which are user-generated metadata for organizing and labeling in many OSNs, as bridges to connect between different OSNs. In our four-stage solution framework, various methods are adopted for hashtag and tag filtering, search results representation, clustering and demonstration. Given an event query, in the first stage we generate related items with corresponding tags and hashtags from OSNs and filter the hashtags and tags we need. Then, topical representation is generated for hashtag and tag. The third stage leverages the derived representation for cross-OSN hashtag and tag clustering. Finally, demonstration for each query is produced and the results are organized hierarchically. Experiments on a dataset containing hundreds of search queries and related items demonstrate the effectiveness of our cross-OSN event search framework. Yuqi Gao, Jitao Sang 0001, Chengpeng Fu, Zhengjia Wang 0001, Tongwei Ren, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2021 | Robust CAPTCHAs Towards Malicious OCRabstractTuring test was originally proposed to examine whether machine's behavior is indistinguishable from a human. The most popular and practical Turing test is CAPTCHA, which is to discriminate algorithm from human by offering recognition-alike questions. The recent development of deep learning has significantly advanced the capability of algorithm in solving CAPTCHA questions, forcing CAPTCHA designers to increase question complexity. Instead of designing questions difficult for both algorithm and human, this study attempts to employ the limitations of algorithm to design robust CAPTCHA questions easily solvable to human. Specifically, our data analysis observes that human and algorithm demonstrates different vulnerability to visual distortions: adversarial perturbation is significantly annoying to algorithm yet friendly to human. We are motivated to employ adversarially perturbed images for robust CAPTCHA design in the context of character-based questions. Four modules of multi-target attack, ensemble adversarial training, image preprocessing differentiable approximation, and expectation are proposed to address the characteristics of character-based CAPTCHA cracking. Qualitative and quantitative experimental results demonstrate the effectiveness of the proposed solution. We hope this study can lead to the discussions around adversarial attack/defense in CAPTCHA design and also inspire the future attempts in employing algorithm limitation for practical usage. Jiaming Zhang 0006, Jitao Sang 0001, Shangxi Wu, Yongli Hu, Jian Yu 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Knowledge-driven Egocentric Multimodal Activity RecognitionabstractRecognizing activities from egocentric multimodal data collected by wearable cameras and sensors, is gaining interest, as multimodal methods always benefit from the complementarity of different modalities. However, since high-dimensional videos contain rich high-level semantic information while low-dimensional sensor signals describe simple motion patterns of the wearer, the large modality gap between the videos and the sensor signals raises a challenge for fusing the raw data. Moreover, the lack of large-scale egocentric multimodal datasets due to the cost of data collection and annotation processes makes another challenge for employing complex deep learning models. To jointly deal with the above two challenges, we propose a knowledge-driven multimodal activity recognition framework that exploits external knowledge to fuse multimodal data and reduce the dependence on large-scale training samples. Specifically, we design a dual-GCLSTM (Graph Convolutional LSTM) and a multi-layer GCN (Graph Convolutional Network) to collectively model the relations among activities and intermediate objects. The dual-GCLSTM is designed to fuse temporal multimodal features with top-down relation-aware guidance. In addition, we apply a co-attention mechanism to adaptively attend to the features of different modalities at different timesteps. The multi-layer GCN aims to learn relation-aware classifiers of activity categories. Experimental results on three publicly available egocentric multimodal datasets show the effectiveness of the proposed model. Yi Huang 0037, Xiaoshan Yang, Junyu Gao 0002, Jitao Sang 0001, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | Towards Accuracy-Fairness Paradox: Adversarial Example-based Data Augmentation for Visual DebiasingabstractMachine learning fairness concerns about the biases towards certain protected or sensitive group of people when addressing the target tasks. This paper studies the debiasing problem in the context of image classification tasks. Our data analysis on facial attribute recognition demonstrates (1) the attribution of model bias from imbalanced training data distribution and (2) the potential of adversarial examples in balancing data distribution. We are thus motivated to employ adversarial example to augment the training data for visual debiasing. Specifically, to ensure the adversarial generalization as well as cross-task transferability, we propose to couple the operations of target task classifier training, bias task classifier training, and adversarial example generation. The generated adversarial examples supplement the target task training dataset via balancing the distribution over bias variables in an online fashion. Results on simulated and real-world debiasing experiments demonstrate the effectiveness of the proposed solution in simultaneously improving model accuracy and fairness. Preliminary experiment on few-shot learning further shows the potential of adversarial attack-based pseudo sample generation as alternative solution to make up for the training data lackage. Yi Zhang 0101, Jitao Sang 0001 |
ACM Multimedia | 2 |
| 2020 | Adversarial Privacy-preserving FilterabstractWhile widely adopted in practical applications, face recognition has been critically discussed regarding the malicious use of face images and the potential privacy problems, e.g., deceiving payment system and causing personal sabotage. Online photo sharing services unintentionally act as the main repository for malicious crawler and face recognition applications. This work aims to develop a privacy-preserving solution, called Adversarial Privacy-preserving Filter (APF), to protect the online shared face images from being maliciously used. We propose an end-cloud collaborated adversarial attack solution to satisfy requirements of privacy, utility and non-accessibility. Specifically, the solutions consist of three modules: (1) image-specific gradient generation, to extract image-specific gradient in the user end with a compressed probe model; (2) adversarial gradient transfer, to fine-tune the image-specific gradient in the server cloud; and (3) universal adversarial perturbation enhancement, to append image-independent perturbation to derive the final adversarial noise. Extensive experiments on three datasets validate the effectiveness and efficiency of the proposed solution. A prototype application is also released for further evaluation. We hope the end-cloud collaborated attack framework could shed light on addressing the issue of online multimedia sharing privacy-preserving issues from user side. Jiaming Zhang 0006, Jitao Sang 0001, Xiaowen Huang 0001, Yongli Hu |
ACM Multimedia | 2 |
| 2020 | Beyond Literal Visual Modeling: Understanding Image Metaphor Based on Literal-Implied Concept Mapping
Chengpeng Fu, Jitao Sang 0001, Jian Yu 0001, Changsheng Xu |
MMM (1) | 3 |
| 2020 | Locality-constrained discrete graph hashing
Wenjie Ying, Jitao Sang 0001, Jian Yu 0001 |
Neurocomputing | 2 |
| 2020 | Meta-path Augmented Sequential Recommendation with Contextual Co-attention NetworkabstractIt is critical to comprehensively and efficiently learn user preferences for an effective sequential recommender system. Existing sequential recommendation methods mainly focus on modeling local preference from users’ historical behaviors, which largely ignore the global context information from the heterogeneous information network. This prevents a comprehensive user preference representation. To address these issues, we propose a joint learning approach to incorporate global context with local preferences efficiently. The proposed approach introduces meta-paths from a heterogeneous information network to capture the global context information, and the position-based self-attention mechanism is adopted to model the local preference representation efficiently. Compared with the methods that only consider the local preference, our proposed method takes the advantages of incorporating global context information, which extracts structural features that captures relevant semantics to construct users’ global preference representation for the sequential recommendation. We further adopt a co-attention mechanism to model complex interactions between global context and users’ historical behaviors for better user representations. Quantitative and qualitative experimental evaluations are conducted on nine large-scale Amazon datasets and a multi-modal Zhihu dataset. The promising results demonstrate the effectiveness of the proposed model. Xiaowen Huang 0001, Shengsheng Qian, Quan Fang, Jitao Sang 0001, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2019 | Explainable Interaction-driven User Modeling over Knowledge Graph for Sequential RecommendationabstractCompared with the traditional recommendation system, sequential recommendation holds the ability of capturing the evolution of users' dynamic interests. Many previous studies in sequential recommendation focus on the accuracy of predicting the next item that a user might interact with, while generally ignore providing explanations why the item is recommended to the user. Appropriate explanations are critical to help users adopt the recommended item, and thus improve the transparency and trustworthiness of the recommendation system. In this paper, we propose a novel Explainable Interaction-driven User Modeling (EIUM) algorithm to exploit Knowledge Graph (KG) for constructing an effective and explainable sequential recommender. Qualified semantic paths between specific user-item pair are extracted from KG. Encoding those semantic paths and learning the importance scores for each path provides the path-wise explanation for the recommendation system. Different from traditional item- level sequential modeling methods, we capture the interaction-level user dynamic preferences by modeling the sequential interactions. It is a high- level representation which contains auxiliary semantic information from KG. Furthermore, we adopt a joint learning manner for better representation learning by employing multi-modal fusion, which benefits from the structural constraints in KG and involves three kinds of modalities. Extensive experiments on the large-scale dataset show the better performance of our approach in making sequential recommendations in terms of both accuracy and explainability. Xiaowen Huang 0001, Quan Fang, Shengsheng Qian, Jitao Sang 0001, Yan Li 0068, Changsheng Xu |
ACM Multimedia | 4 |
| 2019 | Comprehensive Event Storyline Generation from MicroblogsabstractMicroblogging data contains a wealth of information of trending events and has gained increased attention among users, organizations, and research scholars for social media mining in different disciplines. Event storyline generation is one typical task of social media mining, whose goal is to extract the development stages with associated description of events. Existing storyline generation methods either generate storyline with less integrity or fail to guarantee the coherence between the discovered stages. Secondly, there are no scientific method to evaluate the quality of the storyline. In this paper, we propose a comprehensive storyline generation framework to address the above disadvantages. Given Microblogging data related to the specified event, we first propose Hot-Word-Based stage detection algorithm to identify the potential stages of event, which can effectively avoid ignoring important stages and preventing inconsistent sequence between stages. Community detection algorithm is applied then to select representative data for each stage. Finally, we conduct graph optimization algorithm to generate the logically coherent storylines of the event. We also introduce a new evaluation metric, SLEU, to emphasize the importance of the integrity and coherence of the generated storyline. Extensive experiments on real-world Chinese microblogging data demonstrate the effectiveness of the proposed methods in each module and the overall framework. Wenjin Sun, Yuhang Wang 0007, Yuqi Gao, Zesong Li, Jitao Sang 0001, Jian Yu 0001 |
MMAsia | 5 |
| 2019 | Multi-source User Attribute Inference based on Hierarchical Auto-encoderabstractWith the rapid development of Online Social Networks (OSNs), it is crucial to construct users' portraits from their dynamic behaviors to address the increasing needs for customized information services. Previous work on user attribute inference mainly concentrated on developing advanced features/models or exploiting external information and knowledge but ignored the contradiction between dynamic behaviors and stable demographic attributes, which results in deviation of user understanding Xiangguo Ding, Xiaowen Huang 0001, Jitao Sang 0001, Jian Yu 0001 |
MMAsia | 5 |
| 2019 | Multimodal Attribute and Feature Embedding for Activity RecognitionabstractHuman Activity Recognition (HAR) automatically recognizes human activities such as daily life and work based on digital records, which is of great significance to medical and health fields. Egocentric video and human acceleration data comprehensively describe human activity patterns from different aspects, which have laid a foundation for activity recognition based on multimodal behavior data. However, on the one hand, the low-level multimodal signal structures differ greatly and the mapping to high-level activities is complicated. On the other hand, the activity labeling based on multimodal behavior data has high cost and limited data amount, which limits the technical development in this field. In this paper, an activity recognition model MAFE based on multimodal attribute feature embedding is proposed. Before the activity recognition, the middle-level attribute features are extracted from the low-level signals of different modes. On the one hand, the mapping complexity from the low-level signals to the high-level activities is reduced, and on the other hand, a large number of middle-level attribute labeling data can be used to reduce the dependency on the activity labeling data. We conducted experiments on Stanford-ECM datasets to verify the effectiveness of the proposed MAFE method. Yi Huang 0037, Wanting Yu, Xiaoshan Yang, Wei Wang 0354, Jitao Sang 0001 |
MMAsia | 6 |
| 2018 | CSAN: Contextual Self-Attention Network for User Sequential RecommendationabstractThe sequential recommendation is an important task for online user-oriented services, such as purchasing products, watching videos, and social media consumption. Recent work usually used RNN-based methods to derive an overall embedding of the whole behavior sequence, which fails to discriminate the significance of individual user behaviors and thus decreases the recommendation performance. Besides, RNN-based encoding has fixed size and makes further recommendation application inefficient and inflexible. The online sequential behaviors of a user are generally heterogeneous, polysemous, and dynamically context-dependent. In this paper, we propose a unified Contextual Self-Attention Network (CSAN) to address the three properties. Heterogeneous user behaviors are considered in our model that are projected into a common latent semantic space. Then the output is fed into the feature-wise self-attention network to capture the polysemy of user behaviors. In addition, the forward and backward position encoding matrices are proposed to model dynamic contextual dependency. Through extensive experiments on two real-world datasets, we demonstrate the superior performance of the proposed model compared with other state-of-the-art algorithms. Xiaowen Huang 0001, Shengsheng Qian, Quan Fang, Jitao Sang 0001, Changsheng Xu |
ACM Multimedia | 4 |
| 2018 | Bundled Local Features for Image RepresentationabstractLocal features have been widely used for image representation. Traditional methods often treat each local feature independently or simply model the correlations of local features with spatial partition. However, local features are correlated and should be jointly modeled. Besides, due to the variety of images, predefined partition rules will probably introduce noisy information. To solve these problems, in this paper we propose a novel bundled local features method for efficient image representation and apply it for classification. Specially, we first extract local features and bundle them together with over-complete spatial shapes by viewing each local feature as the central point. Then, the most discriminatively bundling features are selected by reconstruction error minimization. The encoding parameters are then used for image representations in a matrix form. Finally, we train bi-linear classifiers with quadratic hinge loss to predict the classes of images. The proposed method can combine local features appropriately and efficiently for discriminative representations. Experimental results on three image data sets show the effectiveness of the proposed method compared with other local features combination strategies. Chunjie Zhang 0001, Jitao Sang 0001, Guibo Zhu, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Understanding Dynamic Cross-OSN Associations for Cold-Start RecommendationabstractOnline social networks (OSNs) have become an essential part of people's daily life, and an increasing number of users are now using multiple OSNs for different social media services simultaneously. As a result, user's interests and preferences usually distribute in different OSNs. While most of the existing work mainly aggregates the distributed user behaviors or features directly, recently very few efforts have been focused on understanding the cross-OSN association from collective user behaviors. In this paper, we go one step further to consider the dynamic characteristic of user behaviors and propose a dynamic cross-OSN association mining framework. In this framework, dynamic user modeling is first conducted to capture the drift of user interest in each OSN. A session-based factorization method is then proposed to establish the cross-OSN association in a dynamic manner, by incrementally updating the derived association each time a new session of data arrives. Based on the derived dynamic association, we finally design a cold-start YouTube video recommendation application, by only utilizing users' behaviors in Twitter. Experiments are conducted using real-world user data from Twitter and YouTube. The results demonstrate the effectiveness of this proposed framework in capturing the underlying association between different OSNs and achieving superior cold-start recommendation performance. Jitao Sang 0001, Ming Yan 0008, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2017 | Hashtag-centric Immersive Search on Social MediaabstractSocial media information distributes in different Online Social Networks (OSNs). This paper addresses the problem integrating the cross-OSN information to facilitate an immersive social media search experience. We exploit hashtag, which is widely used to annotate and organize multi-modal items in different OSNs, as the bridge for information aggregation and organization. A three-stage solution framework is proposed for hashtag representation, clustering and demonstration. Given an event query, the related items from three OSNs, Twitter, Flickr and YouTube, are organized in cluster-hashtag-item hierarchy for display. The effectiveness of the proposed solution is validated by qualitative and quantitative experiments on hundreds of trending event queries. Yuqi Gao, Jitao Sang 0001, Tongwei Ren, Changsheng Xu |
ACM Multimedia | 2 |
| 2017 | Towards SMP Challenge: Stacking of Diverse Models for Social Image Popularity PredictionabstractPopularity prediction on social media has attracted extensive attention nowadays due to its widespread applications, such as online marketing and economical trends. In this paper, we describe a solution of our team CASIA-NLPR-MMC for Social Media Prediction (SMP) challenge. This challenge is designed to predict the popularity of social media posts. We present a stacking framework by combining a diverse set of models to predict the popularity of images on Flickr using user-centered, image content and image context features. Several individual models are employed for scoring popularity of an image at earlier stage, and then a stacking model of Support Vector Regression (SVR) is utilized to train a meta model of different individual models trained beforehand. The Spearman's Rho of this Stacking model is 0.88 and the mean absolute error is about 0.75 on our test set. On the official final-released test set, the Spearman's Rho is 0.7927 and mean absolute error is about 1.1783. The results on provided dataset demonstrate the effectiveness of our proposed approach for image popularity prediction. Xiaowen Huang 0001, Yuqi Gao, Quan Fang, Jitao Sang 0001, Changsheng Xu |
ACM Multimedia | 4 |
| 2017 | A Demo for Image-Based Personality Test
Huaiwen Zhang, Jiaming Zhang 0006, Jitao Sang 0001, Changsheng Xu |
MMM (2) | 3 |
| 2017 | Who Are Your "Real" Friends: Analyzing and Distinguishing Between Offline and Online Friendships From Social Multimedia DataabstractThe Internet has extended the physical boundary of people's social circles to manage an inordinate number of online friends. It is recognized that only a fraction of these online friends are also known with each other in offline circumstances, i.e., the offline friends. An important type of offline friend, onsite offline friend, is defined and addressed in this paper. We explores the possibility of utilizing users' online photo sharing-related behaviors and network topologies to analyze and distinguish between online and onsite offline friendships. Different from traditional social science studies which rely on survey-based data, we employ users' tagged people on the shared Instagram photos as the ground-truth for onsite offline friends. This enables a large-scale and objective analysis and experimental evaluation, which compares between different factors and identifies the features that are key to onsite offline friend identification. Dongyuan Lu, Jitao Sang 0001, Zhineng Chen, Min Xu 0001, Tao Mei 0001 |
IEEE Trans. Multim. | 2 |
| 2016 | Folksonomy-Based Visual Ontology Construction and Its ApplicationsabstractAn ontology hierarchically encodes concepts and concept relationships, and has a variety of applications such as semantic understanding and information retrieval. Previous work for building ontologies has primarily relied on labor-intensive human contributions or focused on text-based extraction. In this paper, we consider the problem of automatically constructing a folksonomy-based visual ontology (FBVO) from the user-generated annotated images. A systematic framework is proposed consisting of three stages as concept discovery, concept relationship extraction, and concept hierarchy construction. The noisy issues of the user-generated tags are carefully addressed to guarantee the quality of derived FBVO. The constructed FBVO finally consists of 139 825 concept nodes and millions of concept relationships by mining more than 2.4 million Flickr images. Experimental evaluations show that the derived FBVO is of high quality and consistent with human perception. We further demonstrate the utility of the derived FBVO in applications of complex visual recognition and exploratory image search. Quan Fang, Changsheng Xu, Jitao Sang 0001, M. Shamim Hossain, Ahmed Ghoneim |
IEEE Trans. Multim. | 3 |
| 2016 | A Unified Video Recommendation by Cross-Network User ModelingabstractOnline video sharing sites are increasingly encouraging their users to connect to the social network venues such as Facebook and Twitter, with goals to boost user interaction and better disseminate the high-quality video content. This in turn provides huge possibilities to conduct cross-network collaboration for personalized video recommendation. However, very few efforts have been devoted to leveraging users’ social media profiles in the auxiliary network to capture and personalize their video preferences, so as to recommend videos of interest. In this article, we propose a unified YouTube video recommendation solution by transferring and integrating users’ rich social and content information in Twitter network. While general recommender systems often suffer from typical problems like cold-start and data sparsity, our proposed recommendation solution is able to effectively learn from users’ abundant auxiliary information on Twitter for enhanced user modeling and well address the typical problems in a unified framework. In this framework, two stages are mainly involved: (1) auxiliary-network data transfer, where user preferences are transferred from an auxiliary network by learning cross-network knowledge associations; and (2) cross-network data integration, where transferred user preferences are integrated with the observed behaviors on a target network in an adaptive fashion. Experimental results show that the proposed cross-network collaborative solution achieves superior performance not only in terms of accuracy, but also in improving the diversity and novelty of the recommended videos. Ming Yan 0008, Jitao Sang 0001, Changsheng Xu, M. Shamim Hossain |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2015 | Unified YouTube Video Recommendation via Cross-network CollaborationabstractThe ever growing number of videos on YouTube makes recommendation an important way to help users explore interesting videos. Similar to general recommender systems, YouTube video recommendation suffers from typical problems like new user, cold-start, data sparsity, etc. In this paper, we propose a unified YouTube video recommendation solution via cross-network collaboration: users' auxiliary information on Twitter are exploited to address the typical problems in single network-based recommendation solutions. The proposed two-stage solution first transfers user preferences from auxiliary network by learning cross-network behavior correlations, and then integrates the transferred preferences with the observed behaviors on target network in an adaptive fashion. Experimental results show that the proposed cross-network collaborative solution achieves superior performance not only in term of accuracy, but also in improving the diversity and novelty of the recommended videos. Ming Yan 0008, Jitao Sang 0001, Changsheng Xu |
ICMR | 2 |
| 2015 | Activity Sensor: Check-In Usage Mining for Local RecommendationabstractWhile on the go, people are using their phones as a personal concierge discovering what is around and deciding what to do. Mobile phone has become a recommendation terminal customized for individuals—capable of recommending activities and simplifying the accomplishment of related tasks. In this article, we conduct usage mining on the check-in data, with summarized statistics identifying the local recommendation challenges of huge solution space, sparse available data, and complicated user intent, and discovered observations to motivate the hierarchical, contextual, and sequential solution. We present a point-of-interest (POI) category-transition--based approach, with a goal of estimating the visiting probability of a series of successive POIs conditioned on current user context and sensor context. A mobile local recommendation demo application is deployed. The objective and subjective evaluations validate the effectiveness in providing mobile users both accurate recommendation and favorable user experience. Jitao Sang 0001, Tao Mei 0001, Changsheng Xu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2015 | Relational User Attribute Inference in Social MediaabstractNowadays, more and more people are engaged in social media to generate multimedia information, i.e., creating text and photo profiles and posting multimedia messages . Such multimodal social networking activities reveal multiple user attributes such as age, gender, and personal interest. Inferring user attributes is important for user profiling, retrieval , and personalization . Existing work is devoted to inferring user attributes independently and ignores the dependency relations between attributes. In this work, we investigate the problem of relational user attribute inference by exploring the relations between user attributes and extracting both lexical and visual features from online user-generated content. We systematically study six types of user attributes: gender, age, relationship , occupation , interest, and emotional orientation. In view of methodology , we propose a relational latent SVM (LSVM) model to combine a rich set of user features, attribute inference, and attribute relations in a unified framework. In the model, one attribute is selected as the target attribute and others are selected as the auxiliary attributes to assist the target attribute inference. The model infers user attributes and attribute relations simultaneously . Extensive experiments conducted on a collected dataset from Google+ with full attribute annotations demonstrate the effectiveness of the proposed approach in user attribute inference and attribute-based user retrieval. Quan Fang, Jitao Sang 0001, Changsheng Xu, M. Shamim Hossain |
IEEE Trans. Multim. | 2 |
| 2015 | Word-of-Mouth Understanding: Entity-Centric Multimodal Aspect-Opinion Mining in Social MediaabstractMost existing approaches on aspect-opinion mining focus on the text domain and cannot be applied to social media where the aspects are essentially multimodal and the opinions depend on the specific aspects. To address the problem of multimodal aspect-opinion mining for entities by leveraging multiple cross-collection sources in social media, in this paper we propose a multimodal aspect-opinion model (mmAOM) considering both user-generated photos and textual documents to simultaneously capture correlations between textual and visual modalities, as well as associations between aspects and opinions . By identifying the aspects and the corresponding opinions related to entities, we apply the mmAOM to entity association visualization and multimodal aspect-opinion retrieval. We have conducted extensive experiments on real-world datasets of entities including Flickr photos, Tripadvisor reviews, and news articles. Qualitative and quantitative evaluation results have validated the effectiveness of the multimodal aspect-opinion mining model, and demonstrated the utility of the derived aspects and opinions from mmAOM in applications of entity association visualization and aspect-opinion retrieval. Quan Fang, Changsheng Xu, Jitao Sang 0001, M. Shamim Hossain, Muhammad Ghulam |
IEEE Trans. Multim. | 3 |
| 2015 | YouTube Video Promotion by Cross-Network Association: @Britney to Advertise Gangnam StyleabstractThe emergence and rapid proliferation of various social media networks have reshaped the way how video contents are generated, distributed, and consumed in traditional video sharing portals. Nowadays, online videos can be accessed from far beyond the internal mechanisms of the video sharing portals, such as internal search and front page highlight. Recent studies have found that external referrers, such as external search engines and other social media websites, arise to be the new and important portals to lead users to online videos. In this paper, we introduce a novel cross-network collaborative application to help drive the online traffic for given videos in the traditional video portal YouTube by leveraging the high propagation efficiency of the popular Twitter followees. Since YouTube videos and Twitter followees distribute on heterogeneous spaces, we present a cross-network association-based solution framework. In this framework, we first represent YouTube videos and Twitter followees in the corresponding topic spaces separately by employing generative topic models. Then, the cross-network topic spaces are associated from both semantic-based and network-based perspectives through the collective intelligence of the observed overlapped users. Based on the derived cross-network association, we finally match the query YouTube videos and candidate Twitter followees in the same topic space with a unified ranking method. The experiments on a real-world large-scale dataset of more than 2.2 million YouTube videos and 31.8 million tweets from 38,540 YouTube users and 39,400 Twitter users demonstrate the effectiveness and superiority of our solution in which network-based and semantic-based association are integrated. Ming Yan 0008, Jitao Sang 0001, Changsheng Xu, M. Shamim Hossain |
IEEE Trans. Multim. | 2 |
| 2015 | Learning Feature Hierarchies: A Layer-Wise Tag-Embedded ApproachabstractFeature representation learning is an important and fundamental task in multimedia and pattern recognition research. In this paper, we propose a novel framework to explore the hierarchical structure inside the images from the perspective of feature representation learning, which is applied to hierarchical image annotation. Different from the current trend in multimedia analysis of using pre-defined features or focusing on the end-task “flat” representation, we propose a novel layer-wise tag- embedded deep learning (LTDL) model to learn hierarchical features which correspond to hierarchical semantic structures in the tag hierarchy . Unlike most existing deep learning models, LTDL utilizes both the visual content of the image and the hierarchical information of associated social tags. In the training stage, the two kinds of information are fused in a bottom-up way. Supervised training and multi-modal fusion alternate in a layer-wise way to learn feature hierarchies. To validate the effectiveness of LTDL, we conduct extensive experiments for hierarchical image annotation on a large-scale public dataset. Experimental results show that the proposed LTDL can learn representative features with improved performances. Zhaoquan Yuan, Changsheng Xu, Jitao Sang 0001, Shuicheng Yan, M. Shamim Hossain |
IEEE Trans. Multim. | 3 |
| 2014 | Mining Cross-network Association for YouTube Video PromotionabstractWe introduce a novel cross-network collaborative problem in this work: given YouTube videos, to find optimal Twitter followees that can maximize the video promotion on Twitter. Since YouTube videos and Twitter followees distribute on heterogeneous spaces, we present a cross-network association-based solution framework. Three stages are addressed: (1) heterogeneous topic modeling, where YouTube videos and Twitter followees are modeled in topic level; (2) cross-network topic association, where the overlapped users are exploited to conduct cross-network topic distribution transfer; and (3) referrer identification, where the query YouTube video and candidate Twitter followees are matched in the same topic space. Different methods in each stage are designed and compared by qualitative as well as quantitative experiments. Based on the proposed framework, we also discuss the potential applications, extensions, and suggest some principles for future heterogeneous social media utilization and cross-network collaborative applications. Ming Yan 0008, Jitao Sang 0001, Changsheng Xu |
ACM Multimedia | 2 |
| 2014 | A Unified Framework of Latent Feature Learning in Social MediaabstractThe current trend in social media analysis and application is to use the pre-defined features and devoted to the later model development modules to meet the end tasks. Representation learning has been a fundamental problem in machine learning, and widely recognized as critical to the performance of end tasks. In this paper, we provide evidence that specially learned features will addresses the diverse, heterogeneous, and collective characteristics of social media data. Therefore, we propose to transfer the focus from the model development to latent feature learning, and present a unified framework of latent feature learning on social media. To address the noisy, diverse, heterogeneous, and interconnected characteristics of social media data, the popular deep learning is employed due to its excellent abstract abilities. In particular, we instantiate the proposed framework by (1) designing a novel relational generative deep learning model to solve the social media link analysis task, and (2) developing a multimodal deep learning to lambda rank model towards the social image retrieval task. We show that the derived latent features lead to improvement in both of the social media tasks. Zhaoquan Yuan, Jitao Sang 0001, Changsheng Xu, Yan Liu 0004 |
IEEE Trans. Multim. | 2 |
| 2014 | Twitter is Faster: Personalized Time-Aware Video Recommendation from Twitter to YouTubeabstractTraditional personalized video recommendation methods focus on utilizing user profile or user history behaviors to model user interests, which follows a static strategy and fails to capture the swift shift of the short-term interests of users. According to our cross-platform data analysis, the information emergence and propagation is faster in social textual stream-based platforms than that in multimedia sharing platforms at micro user level. Inspired by this, we propose a dynamic user modeling strategy to tackle personalized video recommendation issues in the multimedia sharing platform YouTube, by transferring knowledge from the social textual stream-based platform Twitter. In particular, the cross-platform video recommendation strategy is divided into two steps. (1) Real-time hot topic detection: the hot topics that users are currently following are extracted from users' tweets, which are utilized to obtain the related videos in YouTube. (2) Time-aware video recommendation: for the target user in YouTube, the obtained videos are ranked by considering the user profile in YouTube, time factor, and quality factor to generate the final recommendation list. In this way, the short-term (hot topics) and long-term (user profile) interests of users are jointly considered. Carefully designed experiments have demonstrated the advantages of the proposed method. Zhengyu Deng, Ming Yan 0008, Jitao Sang 0001, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2013 | Friend transfer: Cold-start friend recommendation with cross-platform transfer learning of social knowledgeabstractThe emergence of various and disparate social media platforms has opened opportunities for the research on cross-platform media analysis. This provides huge potentials to solve many challenging problems which cannot be well explored in one single platform. In this paper, we investigate into cross-platform social relation and behavior information to address the cold-start friend recommendation problem. In particular, we conduct an in-depth data analysis to examine what information can better transfer from one platform to another and the result demonstrates a strong correlation for the bidirectional relation and common contact behavior between our test platforms. Inspired by the observations, we design a random walk-based method to employ and integrate these convinced social information to boost friend recommendation performance. To validate the effectiveness of our cross-platform social transfer learning, we have collected a cross-platform dataset including 3,000 users with recognized accounts in both Flickr and Twitter. We demonstrate the effectiveness of the proposed friend transfer methods by promising results. Ming Yan 0008, Jitao Sang 0001, Tao Mei 0001, Changsheng Xu |
ICME | 2 |
| 2013 | Tag-aware image classification via Nested Deep Belief netsabstractWith the rising of internet photos-sharing web sites, the rich aware text information surrounding images on the sites are proved helpful to improve the image classification. This paper presents a novel nested deep learning model called Nested Deep Belief Network(NDBN) for tag-aware image classification. A multi-layer structure of Deep Belief Network(DBN) is established to learn a unified representation of visual feature and tag feature for an image, and an additional Gaussian Restricted Boltzmann Machine is built to capture the tag-tag dependency. Compared with conventional methods, the proposed model can not only find correlations across modalities, but mine the importance for different tags, and also bring about low-rank tag feature representation. We conduct experiments over the MIR Flickr dataset and the results show that the proposed NDBN model outperforms the existing image classification techniques. Zhaoquan Yuan, Jitao Sang 0001, Changsheng Xu |
ICME | 2 |
| 2013 | Latent feature learning in social media networkabstractThe current trend in social media analysis and application is to use the pre-defined features and devoted to the later model development modules to meet the end tasks. In this work, we claim that representation is critical to the end tasks and contributes much to the model development module. We provide evidence that specially learned feature well addresses the diverse, heterogeneous and collective characteristics of social media data. Therefore, we propose to transfer the focus from the model development to latent feature learning, and present a general feature learning framework based on the popular deep architecture. In particular, following the proposed framework, we design a novel relational generative deep learning model to test the idea on link analysis tasks in the social media networks. We show that the derived latent features well embed both the media content and their observed links, leading to improvement in social media tasks of user recommendation and social image annotation. Zhaoquan Yuan, Jitao Sang 0001, Yan Liu 0004, Changsheng Xu |
ACM Multimedia | 2 |
| 2013 | Interaction Design for Mobile Visual SearchabstractMobile devices are becoming ubiquitous. People take pictures via their phone cameras to explore the world on the go. In many cases, they are concerned with the picture-related information. Understanding user intent conveyed by those pictures therefore becomes important. Existing mobile applications employ visual search to connect the captured picture with the physical world. However, they only achieve limited success due to the ambiguity nature of user intent in the picture-one picture usually contains multiple objects. By taking advantage of multitouch interactions on mobile devices, this paper presents a prototype of interactive mobile visual search, named TapTell, to help users formulate their visual intent more conveniently. This kind of search leverages limited yet natural user interactions on the phone to achieve more effective visual search while maintaining a satisfying user experience. We make three contributions in this work. First, we conduct a focus study on the usage patterns and concerned factors for mobile visual search, which in turn leads to the interactive design of expressing visual intent by gesture. Second, we introduce four modes of gesture-based interactions (crop, line, lasso, and tap) and develop a mobile prototype. Third, we perform an in-depth usability evaluation on these different modes, which demonstrates the advantage of interactions and shows that lasso is the most natural and effective interaction mode. We show that TapTell provides a natural user experience to use phone camera and gesture to explore the world. Based on the observation and conclusion, we also suggest some design principles for interactive mobile visual search in the future. Jitao Sang 0001, Tao Mei 0001, Ying-Qing Xu, Changsheng Xu, Shipeng Li 0001 |
IEEE Trans. Multim. | 1 |
| 2012 | Probabilistic sequential POIs recommendation via check-in dataabstractWhile on the go, people are using their phones as a personal concierge discovering what is around and deciding what to do. Mobile phone has become a recommendation terminal customized for individuals. While existing research predominantly focuses on one-step recommendation---recommending the next single activity according to current context, this work moves one step beyond by recommending a series of activities, which is a package of sequential Points of Interest (POIs). The recommended POIs are not only relevant to user context (i.e., current location, time, and check-in), but also personalized to his/her check-in history. We presents a probabilistic approach, which is highly motivated from a large-scale commercial mobile check-in data analysis, to ranking a list of sequential POI categories (e.g., "Japanese food" and "bar") and POIs (e.g., "I love sushi"). The approach enables users to plan consecutive activities on the move. Specifically, the probabilistic recommendation approach estimates the transition probability from one POI to another, conditioned on current context and check-in history in a Markov chain. To alleviate the discritization error and sparsity problem, we further introduce context collaboration and integrate prior information. Experiments on over 100k real-world check-in records and 20k POIs validate the effectiveness of the proposed approach. Jitao Sang 0001, Tao Mei 0001, Jian-Tao Sun, Changsheng Xu, Shipeng Li 0001 |
SIGSPATIAL/GIS | 1 |