VLDB 2026 Research / reviewers in the wild / expert
Bo Liu 0004
dblp:58/2670-4
· DBLP profile ↗
97ranked-venue papers
8as first author
66since 2021 · last 2026
0000-0001-5209-9063ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 1 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 18 since 2021Databases, data management, data science and information retrieval · 12 · 8 since 2021Computer networks · 10 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2Security and privacy · 2Software engineering, systems software and programming languages · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Don't Be Misled by Style: A Style-Adaptive Reranker for Capturing Effective Knowledge in Retrieval-Augmented GenerationabstractRuwen Zhang, Bo Liu, Zhang Sheng Xiang, Yida Chen, Hantao Zhao, Ding Ding, Jiahui Jin, Jiuxin Cao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ruwen Zhang, Bo Liu 0004, Zhang Sheng Xiang, Hantao Zhao, Ding Ding 0002, Jiahui Jin 0001, Jiuxin Cao |
ACL (1) | 2 |
| 2026 | Virtual Minds, Real Work: LLM-Powered Preference-Based Planning through Spatial Multi-Agent-Human CollaborationabstractPeople frequently face preference-based planning tasks requiring balancing goals with nuanced constraints, yet even advanced LLMs demand considerable effort to produce and adjust plans reflecting complex user preferences. We present MAVIS (Multi-Agent Virtual Interactive Synergy), a multi-agent system within a virtual workspace that introduces an incremental collaboration mechanism. This mechanism automatically decomposes tasks into guidelines and sequentially introduces expert agents. Each agent proactively engages users in focused dialog to uncover implicit preferences, while successive agents add perspectives and transparently negotiate trade-offs. To mitigate textual overload, MAVIS employs spatial visualizations that externalize agents’ reasoning through step-linked summaries and context-aware boards, with embodied avatars supporting natural interaction. Across studies, Study 1 showed our collaboration mechanism doubled expressed preferences and improved planning quality by 60.3% over a conventional LLM baseline. Study 2 affirmed visualization’s benefits over a non-spatial baseline, while Study 3 confirmed its versatility across VR and desktop modalities and diverse tasks. Xin Yi 0001, Zitong Dai, Xuewen Yu, Bo Liu 0004, Jiuxin Cao, Hantao Zhao |
CHI | 6 |
| 2026 | Type dynamics theory-driven personality detection with LLM-enhanced social profiling
Ruwen Zhang, Bo Liu 0004, Xiaorong Hao, Xinhui Huang, Jiuxin Cao |
Appl. Intell. | 2 |
| 2026 | Multi-hop commonsense knowledge injection framework for zero-shot commonsense question answering
Jiuxin Cao, Biwei Cao, Qingqing Gao, Bo Liu 0004 |
Expert Syst. Appl. | 5 |
| 2026 | PACEP: Steering large language models toward coherent user profile tracking and strategy selection in emotional support conversations
Hanyu Luo, Jiuxin Cao, Chang Liu 0113, Xinglin Li, Bo Liu 0004, Biwei Cao |
Expert Syst. Appl. | 5 |
| 2026 | MoE-MT: Unified multi-task offensive speech recognition via mixture-of-experts semantic fusion
Chuanlong Ma, Jiuxin Cao, Bo Liu 0004 |
Expert Syst. Appl. | 5 |
| 2026 | The intelligent social event observer: Multi-source continuous event integration, discovery, and induction with LLMs
Ruwen Zhang, Bo Liu 0004, Jiuxin Cao, Hantao Zhao |
Expert Syst. Appl. | 2 |
| 2026 | AutoIT: Automated Image Tagging with Random Perturbation
Xuelin Zhu, Jianshu Li, Dongqi Tang, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao |
Int. J. Comput. Vis. | 7 |
| 2026 | ARena of Privacy: Exploring Augmented Reality in Enhancing Smart Home Privacy Awareness and ControlabstractThe rapid emergence of smart homes offers convenience but also raises significant privacy concerns, such as challenges in comprehending privacy-related statuses, information, and settings. This study aims to mitigate these concerns by leveraging augmented reality technology to enhance privacy protection user experiences. Through interviews with experienced users in the Chinese culture context, we systematically identify and categorize three primary aspects of concerns encountered during smart home usage: sensor range and status, data transmission directions, and transparency of mode settings. Drawing on these insights, we design and implement a streamlined AR system for interacting with smart home devices. Our user studies demonstrate that our AR system offers significant advantages over traditional 2D smart home applications in terms of information representation, workload reduction (i.e., number of clicks), and enhanced privacy awareness. This research presents a notable step forward in privacy protection while suggesting pathways for AR-enhanced smart home systems. Hantao Zhao, Xin Yi 0001, Liru Chen, Wenze Ren, Xiaomeng Shi, Bo Liu 0004, Jiuxin Cao |
Int. J. Hum. Comput. Interact. | 7 |
| 2026 | Towards strategic persuasion: Unveiling users' susceptibility to persuasive strategies in dialogues
Bo Liu 0004, Mingrui Hu, Mingjie Dai, Jiuxin Cao, Hantao Zhao |
Knowl. Based Syst. | 2 |
| 2026 | Behavior-Driven Detection of Social Bot Groups Through Coordination Patterns and Stance Consistency
Xiaoyu Xue, Ruwen Zhang, Bo Liu 0004, Jiuxin Cao |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2026 | Multi-Label Image Classification via Contrastive Co-Occurrence LearningabstractMulti-label image classification is an essential task in computer vision that aims to identify multiple objects in images. Recently, there has been growing research interest in modeling the relationships between labels to enhance label representation learning. An intuitive approach is to train a network to estimate label co-occurrence probabilities in a supervised manner, which are then leveraged to guide the interactions between label representations. However, the extreme sparsity of label co-occurrence signals poses substantial challenges. To address this issue, we commence by examining the potential interaction behaviors between label representations under the guidance of ground-truth label co-occurrence signals. Inspired by our findings, a novel contrastive learning mechanism is crafted to mimic and enhance such behaviors, facilitating effective label representation interactions without relying on explicit supervision from label co-occurrence signals. Based on this, we develop a pioneering contrastive co-occurrence learning framework, which operates on the instance-level label co-occurrence graph for multi-label image classification. This framework involves sequential processes of label representation learning followed by co-occurrence perception learning. Cross-entropy loss for label classification learning and contrastive loss for co-occurrence perception learning are used to jointly optimize the entire framework end-to-end. In this way, label representations can interact effectively, fully perceiving their co-occurrence relationships at the instance level, thereby significantly improving the performance in label recognition. Extensive experiments on public benchmarks demonstrate the superiority of the proposed framework in multi-label image classification. Codes are available on https://github.com/jasonseu/CoCo. Xuelin Zhu, Dongqi Tang, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao |
IEEE Trans. Image Process. | 6 |
| 2025 | External Reliable Information-enhanced Multimodal Contrastive Learning for Fake News DetectionabstractWith the rapid development of the Internet, the information dissemination paradigm has changed and the efficiency has been improved greatly. While this also brings the quick spread of fake news and leads to negative impacts on cyberspace. Currently, the information presentation formats have evolved gradually, with the news formats shifting from texts to multimodal contents. As a result, detecting multimodal fake news has become one of the research hotspots. However, multimodal fake news detection research field still faces two main challenges: the inability to fully and effectively utilize multimodal information for detection, and the low credibility or static nature of the introduced external information, which limits dynamic updates. To bridge the gaps, we propose ERIC-FND, an external reliable information-enhanced multimodal contrastive learning framework for fake news detection. ERIC-FND strengthens the representation of news contents by entity-enriched external information enhancement method. It also enriches the multimodal news information via multimodal semantic interaction method where the multimodal constrative learning is employed to make different modality representations learn from each other. Moreover, an adaptive fusion method is taken to integrate the news representations from different dimensions for the eventual classification. Experiments are done on two commonly used datasets in different languages, X (Twitter) and Weibo. Experiment results demonstrate that our proposed model ERIC-FND outperforms existing state-of-the-art fake news detection methods under the same settings. Biwei Cao, Qihang Wu, Jiuxin Cao, Bo Liu 0004, Jie Gui |
AAAI | 4 |
| 2025 | PsyAdvisor: A Plug-and-Play Strategy Advice Planner with Proactive Questioning in Psychological ConversationsabstractProactive questioning is essential in psychological conversations as it helps uncover deeper issues and unspoken concerns.Current psychological LLMs are constrained by passive response mechanisms, limiting their capacity to deploy proactive strategies for psychological counseling.To bridge this gap, we first develop the ProPsyC (Proactive Psychological Conversation) dataset, a multi-turn conversation dataset with interpretive labels including strategy decision logic and reaction attribution.Based on ProPsyC, we propose PsyAdvisor by supervised fine-tuning, a plug-and-play proactive questioning strategy planner that empowers psychological LLMs to initiate welltimed questioning through strategic prompting.Experimental results demonstrate that psychological LLMs integrated with PsyAdvisor substantially improve proactive questioning capacity, conversation depth, and response quality.Furthermore, PsyAdvisor shows promising potential in assisting novice counselors by providing strategy recommendations.This study provides new optimization directions for psychological conversation systems and offers valuable insights for future research on proactive questioning mechanisms in psychological LLMs.Our code are available at https: //github.com/EthanHu777/PsyAdvisor. Bo Liu 0004, Jiuxin Cao |
ACL (1) | 3 |
| 2025 | MAGRET: Machine-generated Text Detection with Rewritten TextsabstractWith the quick advancement in text generation ability of Large Language Mode(LLM), concerns about the misuse of machine-generated content have grown, raising potential violations of legal and ethical standards. Some existing studies concentrate on detecting machine-generated text in open-source models using in-model features, but their performance on closed-source large models is limited. This limitation occurs because, in the closed-source model detection, the only reference that can be obtained is the texts, which may differ significantly due to random sampling. In this paper, we demonstrate that texts generated by the same model can align both semantically and statistically under similar prompts, facilitating effective detection and traceability. Specifically, we fine-tune a BERT encoder through contrastive learning to achieve semantic alignment in randomly generated texts from the same model. Then, we propose a method called Machine-Generated Text Detection with Rewritten Texts, which designed several prompt refactoring methods and used them to request rewritten text from LLMs. Semantic and statistical relationships between rewritten and original texts provide a basis for detection and traceability. Finally, we expanded the text dataset with multi-parameter random sampling and verified the performance of MAGRET on three text-generated datasets. Experimental results show that previous methods struggle with closed-source model detection, while our approach significantly outperforms baseline methods in this regard. It also shows MagRet’s stable performance in detection and tracing tasks across various randomly sampled texts. Jiuxin Cao, Hanyu Luo, Bo Liu 0004 |
COLING | 5 |
| 2025 | Positive Text Reframing under Multi-strategy OptimizationabstractDiffering from sentiment transfer, positive reframing seeks to substitute negative perspectives with positive expressions while preserving the original meaning. With the emergence of pre-trained language models (PLMs), it is possible to achieve acceptable results by fine-tuning PLMs. Nevertheless, generating fluent, diverse and task-constrained reframing text remains a significant challenge. To tackle this issue, a multi-strategy optimization framework (MSOF) is proposed in this paper. Starting from the objective of positive reframing, we first design positive sentiment reward and content preservation reward to encourage the model to transform the negative expressions of the original text while ensuring the integrity and consistency of the semantics. Then, different decoding optimization approaches are introduced to improve the quality of text generation. Finally, based on the modeling formula of positive reframing, we propose a multi-dimensional re-ranking method that further selects candidate sentences from three dimensions: strategy consistency, text similarity and fluency. Extensive experiments on two Seq2Seq PLMs, BART and T5, demonstrate our framework achieves significant improvements on unconstrained and controlled positive reframing tasks. Shutong Jia, Biwei Cao, Qingqing Gao, Jiuxin Cao, Bo Liu 0004 |
COLING | 5 |
| 2025 | SKL-CLIP: Learning Skeleton-Based Action Representations via Language SupervisionabstractCLIP, widely used in multimodal learning, excels due to its large-scale image-text pretraining. However, applying CLIP-like architectures to skeleton-based action representation learning presents challenges due to the incompatible non-visual data structure and the limited scale of skeleton datasets, which hinder robust generalization. To address these issues, we propose SKL-CLIP, a framework that incorporates Supervised Self-Contrastive Learning to mitigate overfitting and enhance transferable representation learning, Knowledge Distillation from the textual encoder of pretrained CLIP models to preserve generalization while adapting to skeleton-text scenarios, and Multi-Domain Parallel Training to leverage diverse support datasets, improving cross-dataset and zero-shot recognition. Extensive experiments on NTU and PKU datasets demonstrate that SKL-CLIP significantly advances skeleton-based action representation, achieving state-of-the-art performance across fully-supervised, cross-data unsupervised domain adaptation, and zero-shot tasks. Kun Wang 0057, Jiuxin Cao, Jiawei Ge 0002, Chang Liu 0113, Bo Liu 0004 |
ICME | 5 |
| 2025 | Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment RetrievalabstractCurrent text-driven Video Moment Retrieval (VMR) methods encode all video clips, including irrelevant ones, disrupting multimodal alignment and hindering optimization. To this end, we propose a denoise-then-retrieve paradigm that explicitly filters text-irrelevant clips from videos and then retrieves the target moment using purified multimodal representations. Following this paradigm, we introduce the Denoise-then-Retrieve Network (DRNet), comprising Text-Conditioned Denoising (TCD) and Text-Reconstruction Feedback (TRF) modules. TCD integrates cross-attention and structured state space blocks to dynamically identify noisy clips and produce a noise mask to purify multimodal video representations. TRF further distills a single query embedding from purified video representations and aligns it with the text embedding, serving as auxiliary supervision for denoising during training. Finally, we perform conditional retrieval using text embeddings on purified video representations for accurate VMR. Experiments on Charades-STA and QVHighlights demonstrate that our approach surpasses state-of-the-art methods on all metrics. Furthermore, our denoise-then-retrieve paradigm is adaptable and can be seamlessly integrated into advanced VMR models to boost performance. Jiuxin Cao, Bo Miao, Zhiheng Fu, Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Mehwish Nasim, Ajmal Mian |
IJCAI | 7 |
| 2025 | A Rationale-Guided Multi-Task Learning Framework for Hate Speech DetectionabstractHate Speech Detection (HSD) plays a crucial role in ensuring respectful online communication and preventing the spread of harmful content. However, existing studies on HSD often focus on the overall intent of the target message while overlooking the fine-grained details in the message. We propose a RatIonale-guided multi-taSk lEarning framework (RISE), which frames HSD as the main task and Human Rationales Tagging (HRT) as the auxiliary task. This enables simultaneous message-level understanding and the identification of expressions that trigger human judgments of hate speech, mimicking human cognitive processing to capture fine-grained information and enhance HSD performance. Furthermore, we extend rationale annotations from binary labels to BIO tagging, capturing the positional roles of tokens within rationales. Additionally, we integrate an emoji semantics interpretation module that interprets emoji meanings, enriching contextual information. Extensive experiments demonstrate that RISE outperforms state-of-the-art models, with each component contributing significantly to improved performance. Qingqing Gao, Jiuxin Cao, Fengshan Song, Biwei Cao, Bo Liu 0004 |
IJCNN | 6 |
| 2025 | Gen4Track: A Tuning-free Data Augmentation Framework via Self-correcting Diffusion Model for Vision-Language TrackingabstractThe performance of current Vision-Language Tracking (VLT) models is constrained by the limited diversity and quantity of labeled data. Compared to constructing large-scale datasets, data augmentation offers a more cost-saving strategy for VLT by synthesizing new samples from existing data, rather than generating them from scratch. However, conventional techniques like rotation and flipping may disrupt scene composition, causing conflicts between visual layouts and textual annotations. Recent advances in generative models have inspired the use of synthetic videos for data augmentation. Yet, existing approaches fail to address the core concerns of data augmentation in VLT (shown in Fig. 1)-target location accuracy, text-video consistency, and video content coherency. To bridge the gap, we propose Gen4Track, a tuning-free data augmentation framework that leverages the self-correcting mechanism to dynamically generate high-quality video data with annotations. Our approach involves (1) optimizing the attention calculations in a frozen text-to-image diffusion model to synthesize coherent videos that satisfy specific conditions (e.g., spatial location, category, color, and style), and (2) implementing a self-correcting mechanism based on a Large Language Model (LLM) to improve text-video consistency. During video augmentation, we propose content-coherent self-attention and location-enhanced cross-attention mechanisms, ensuring that image-level editings are accurately and coherently propagated throughout the video. Then, with the goal of maximizing text-video consistency, we iteratively refine the augmentation instruction with our designed self-correcting mechanism for a more aligned video. Extensive experiments validate that Gen4Track significantly boosts the performance of SOTA VLT models (achieving improvements of up to 3.2% in SUC and 3.5% in PRE), opening a new chapter of training Vision-Language trackers with synthetic videos rather than manually annotated data. Jiawei Ge 0002, Xin-Yu Zhang 0027, Jiuxin Cao, Xuelin Zhu, Qingqing Gao, Biwei Cao, Kun Wang 0057, Chang Liu 0113, Bo Liu 0004, Chen Feng 0028, Ioannis Patras |
ACM Multimedia | 10 |
| 2025 | MIAN: Multi-head Incongruity Aware Attention Network with transfer learning for sarcasm detection
Jiuxin Cao, Biwei Cao, Bo Liu 0004 |
Expert Syst. Appl. | 5 |
| 2025 | FSD-GAN: Generative Adversarial Training for Face Swap Detection via the Latent Noise Fingerprint
Jiawei Ge 0002, Jiuxin Cao, Zhixiang Zhao, Bo Liu 0004 |
J. Comput. Sci. Technol. | 4 |
| 2025 | Causality-inspired representation learning for weakly supervised skeleton-based action recognition
Kun Wang 0057, Jiuxin Cao, Jiawei Ge 0002, Chang Liu 0113, Bo Liu 0004 |
Knowl. Based Syst. | 5 |
| 2025 | Text-to-face synthesis based on facial landmarks prediction
Kun Wang 0057, Biwei Cao, Bo Liu 0004, Jiuxin Cao |
Mach. Vis. Appl. | 4 |
| 2025 | No Place to Hide: Dual Deep Interaction Channel Network for Fake News Detection With Data AugmentationabstractOnline social network has emerged as a prominent place for the propagation of fake news due to its low cost of information dissemination. Although the existing methods have made many attempts in news content and propagation structure, the detection of fake news is still facing two challenges: one is how to mine the unique key features and evolution patterns, and the other is how to tackle the problem of small samples to build the high-performance model. Different from popular methods, which take full advantage of the propagation topology structure, in this article, we propose a novel framework for fake news detection from perspectives of semantics, emotion and data enhancement. The semantic and emotional features of news and comments, the inconsistent emotion between news and news participants as well as the emotion evolution features in comments are fused by the designed dual deep interaction channel network to obtain a more comprehensive and fine-grained news representation. Meanwhile, with the construction of large language model (LLM) prompt, a LLM-based data enhancement module is used to obtain more diverse labeled data of high quality filtered by confidence, further improving the performance of the classification model. Experiments show that the proposed approach outperforms the state-of-the-art methods. Biwei Cao, Jiuxin Cao, Lulu Hua, Bo Liu 0004, Jie Gui, James T. Kwok |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | Adversarial Reinforcement Learning for Enhanced Decision-Making of Evacuation Guidance Robots in Intelligent Fire ScenariosabstractIn the context of rapid urbanization, traditional manual guidance and static evacuation signs are increasingly inadequate for addressing complex and dynamic emergencies. This study proposes an innovative emergency evacuation framework that optimizes the crowd evacuation by integrating multiagent reinforcement learning (MARL) with adversarial reinforcement learning (ARL). The developed simulation environment models realistic human behavior in complex buildings and incorporates robotic navigation and intelligent path planning. A novel simulated human behavior model was integrated, capable of complex human–robot interaction, independent escape route searching, and exhibiting herd mentality and memory mechanisms. We also proposed a multiagent framework that combines MARL and ARL to enhance overall evacuation efficiency and robustness. Additionally, we developed a new ARL evaluation framework that provides a novel method for quantifying agents’ performance. Various experiments of differing difficulty levels were conducted, and the results demonstrate that the proposed framework exhibits advantages in emergency evacuation scenarios. Specifically, our ARLR approach increased survival rates by 1.8% points in low-difficulty evacuation tasks compared to the RLR approach using only MARL algorithms. In high-difficulty evacuation tasks, the ARLR approach raised survival rates from 46.7% without robots to 64.4%, exceeding the RLR approach by 1.7% points. This study aims to enhance the efficiency and safety of human–robot collaborative fire evacuations and provides theoretical support for evaluating and improving the performance and robustness of ARL agents. Hantao Zhao, Tianxing Ma, Xiaomeng Shi, Mubbasir Kapadia, Tyler Thrash, Christoph Hölscher, Jinyuan Jia 0002, Bo Liu 0004, Jiuxin Cao |
IEEE Trans. Comput. Soc. Syst. | 9 |
| 2025 | Using Graph With Interconnection Intervals Embedding to Discover Potential Threats in the NetworkabstractGradually infiltrating through latent malicious connections to form cyberattacks has become a significant threat in cyberspace. Advanced persistent threats (APTs) are such cyberattacks, where intruders maintain a presence on networks to steal secrets. Detecting APTs by comparing user behavior with normal activity is challenging because attackers often mimic legitimate behavior sequences to evade detection. This article detects APTs by modeling the graph with interconnection intervals embedding to weaken attackers’ mimicry, namedCyberProber. We first construct an attack intent graph (AIG) to scale down the original network, while preserving key suspicious paths. Then, we embed the AIG by learning connection representations with interconnection interval features, capturing both structural and temporal dependencies. In this way, we can detect APTs in the small AIG, and further uncover more anomalies in the original network based on the identified compromised nodes in the AIG. Our experiments demonstrate that CyberProber outperforms baselines in APTs detection. Xiaorong Hao, Bo Liu 0004, Xiangguo Sun, Jiuxin Cao, Xinwen Fu |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | tGARD: Text-Guided Adversarial Reconstruction for Industrial Anomaly DetectionabstractIndustrial anomaly detection aims to identify and localize defective regions in images. Among various architectures, reconstruction-based methods have demonstrated exceptional performance. These methods reconstruct anomalous samples into normal ones and identify anomalies through a comparison between them. However, reconstruction process within these methods often focuses on PL similarity, overlooking the high-frequency consistency between the input and output, which constrains the model’s accuracy. This article proposes tGARD, a novel text-guided adversarial reconstruction method for anomaly detection. Specifically, we introduce feature aggregation module, using nonlocal block and dilated convolution to handle complex anomaly patterns. Subsequently, text-guided reconstruction module is meticulously designed to harness CLIP’s multimodal alignment capabilities, allowing for a semantically controllable reconstruction process. This control is achieved by incorporating dynamic text embeddings derived from the CLIP encoder within discriminator. Meanwhile, during reconstruction process, high-frequency details are preserved through a convolutional adversarial discriminator. Finally, category-aware loss weighting strategy is conceived to balance similarity and adversarial loss. Experiments demonstrate that our model achieves significant improvements in anomaly localization, surpassing all reconstruction-based models on MVTec-AD. It also establishes a new state-of-the-art on VisA dataset, outperforming all existing architectures. Yuchen Qiang, Jiuxin Cao, Junyang Yang, Lijia Yu, Bo Liu 0004 |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | Context-Enhanced Video Moment Retrieval With Large Language ModelsabstractCurrent methods for Video Moment Retrieval (VMR) struggle to align complex situations involving specific environmental details, character descriptions, and action narratives. To tackle this issue, we propose a Large Language Model-guided Moment Retrieval (LMR) approach that employs the extensive knowledge of Large Language Models (LLMs) to improve video context representation as well as cross-modal alignment, facilitating accurate localization of target moments. Specifically, LMR introduces a context enhancement technique with LLMs to generate crucial target-related context semantics. These semantics are integrated with visual features for producing discriminative video representations. Finally, a language-conditioned transformer is designed to decode free-form language queries, on the fly, using aligned video representations for moment retrieval. Extensive experiments demonstrate that LMR achieves state-of-the-art results, outperforming the nearest competitor by up to 3.28% and 4.06% on the challenging QVHighlights and Charades-STA benchmarks, respectively. More importantly, the performance gains are significantly higher for localization of complex queries. Bo Miao, Jiuxin Cao, Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Mehwish Nasim, Ajmal Mian |
IEEE Trans. Multim. | 6 |
| 2025 | Beyond Visual Cues: Synchronously Exploring Target-Centric Semantics for Vision-Language TrackingabstractSingle object tracking aims to locate one specific target in video sequences, given its initial state. Classical trackers rely solely on visual cues, restricting their ability to handle challenges such as appearance variations, ambiguity, and distractions. Hence, Vision-Language Tracking (VLT) has emerged as a promising approach, incorporating language descriptions to directly provide high-level semantics and enhance tracking performance. However, current Vision-Language (VL) trackers have not fully exploited the power of multi-modal learning, as they suffer from limitations such as heavily relying on off-the-shelf backbones for feature extraction, ineffective asynchronous fusion designs, and the absence of VL-related loss functions for optimizing multi-modal representation. Consequently, we present a novel tracker that progressively explores target-centric semantics for VLT. Specifically, we propose the first Synchronous Learning Backbone (SLB) for VLT, which consists of two novel modules: the Target Enhance Module (TEM) and the Semantic-Aware Module (SAM). These modules together ensure the multi-modal feature extraction and interaction at the same pace, facilitating the tracker to synchronously perceive target-related semantics from both visual and textual modalities. Moreover, we devise the dense matching loss to further strengthen multi-modal representation learning. Extensive experiments on VLT datasets demonstrate the superiority and effectiveness of our methods. Jiawei Ge 0002, Jiuxin Cao, Xiangmei Chen, Xuelin Zhu, Chang Liu 0113, Kun Wang 0057, Bo Liu 0004 |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2025 | Dual-Domain Triple Contrast for Cross-Dataset Skeleton-Based Action RecognitionabstractSkeleton-based Action Recognition (SAR) is widely recognized for its robustness and efficiency in human action analysis, but its performance in cross-dataset tasks has been limited due to domain shifts between different datasets. To address this challenge, current methods typically approach cross-dataset SAR as an Unsupervised Domain Adaptation (UDA) task, which is tackled using domain adaptation or self-supervised learning strategies. In this article, we propose a Dual-Domain Triple Contrast (D2TC) framework for cross-dataset SAR under the UDA setting. Unlike existing UDA methods that either focus on a single strategy or superficially combine strategies, our D2TC leverages contrastive learning to integrate both strategies into a unified framework. It performs three types of contrastive learning: Self-Supervised Contrastive Learning, Supervised Contrastive Learning, and UDA with Contrastive Learning, across both source and target domains. The triple contrasts go beyond mere summation, effectively bridging the domain gap and enhancing the model’s representational capacity. Additionally, we introduce multi-modal ensemble contrast and extreme skeleton augmentation methods to further enhance the skeleton-based representation learning. Extensive experiments on six cross-dataset settings validate the superiority of our D2TC framework over state-of-the-art methods, demonstrating its effectiveness in reducing domain discrepancies and improving cross-dataset SAR performance. The codes are available on https://github.com/KennCoder7/DualDomainTripleContrast . Kun Wang 0057, Jiuxin Cao, Jiawei Ge 0002, Chang Liu 0113, Bo Liu 0004 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Query-Based Knowledge Sharing for Open-Vocabulary Multi-Label ClassificationabstractIdentifying labels that are unseen during training, known as multi-label zero-shot learning, is a non-trivial task in computer vision. Recent studies have increasingly focused on utilizing vision-language pre-training (VLP) models to recognize unseen labels in an open-vocabulary manner. However, these approaches like knowledge distillation have offered only modest performance gains. The challenge of fully harnessing the potential of VLP models for effective multi-label zero-shot learning remains open. In this work, an advanced query-based knowledge sharing framework is proposed to explore the multi-modal knowledge from VLP models for open-vocabulary multi-label classification. Specifically, we introduce a set of label-agnostic query tokens that are designed to capture essential and informative visual knowledge from input images. These tokens are subsequently shared across all labels, allowing them to select pertinent one as visual clues for accurate recognition. Then, by integrating the pre-trained knowledge of VLP models, these query tokens, trained on seen labels, can be efficiently generalized to the recognition of unseen labels. Additionally, we reformulate ranking learning into a form of classification to enable the magnitude of feature vectors for prediction, which significantly benefits label recognition. Experiment results show that our framework outperforms state-of-the-art methods in multi-label zero-shot learning task by a significant margin, reaching 4.2% and 2.4% in mAP on the NUS-WIDE and Open Images datasets, respectively. Code and models are available at https://github.com/jasonseu/QKS . Xuelin Zhu, Dongqi Tang, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | CEPT: A Contrast-Enhanced Prompt-Tuning Framework for Emotion Recognition in ConversationabstractEmotion Recognition in Conversation (ERC) has attracted increasing attention due to its wide applications in public opinion analysis, empathetic conversation generation, and so on. However, ERC research suffers from the problems of data imbalance and the presence of similar linguistic expressions for different emotions. These issues can result in limited learning for minority emotions, biased predictions for common emotions, and the misclassification of different emotions with similar linguistic expressions. To alleviate these problems, we propose a Contrast-Enhanced Prompt-Tuning (CEPT) framework for ERC. We transform the ERC task into a Masked Language Modeling (MLM) generation task and generate the emotion for each utterance in the conversation based on the prompt-tuning of the Pre-trained Language Model (PLM), where a novel mixed prompt template and a label mapping strategy are introduced for better context and emotion feature modeling. Moreover, Supervised Contrastive Learning (SCL) is employed to help the PLM mine more information from the labels and learn a more discriminative representation space for utterances with different emotions. We conduct extensive experiments and the results demonstrate that CEPT outperforms the state-of-the-art methods on all three benchmark datasets and excels in recognizing minority emotions. Qingqing Gao, Jiuxin Cao, Biwei Cao, Bo Liu 0004 |
LREC/COLING | 5 |
| 2024 | Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning for Visual Story SynthesisabstractThe excellent text-to-image synthesis capability of diffusion models has driven progress in synthesizing coherent visual stories. The current state-of-the-art method combines the features of historical captions, historical frames, and the current captions as conditions for generating the current frame. However, this method treats each historical frame and caption as the same contribution. It connects them in order with equal weights, ignoring that not all historical conditions are associated with the generation of the current frame. To address this issue, we propose Causal-Story. This model incorporates a local causal attention mechanism that considers the causal relationship between previous captions, frames, and current captions. By assigning weights based on this relationship, Causal-Story generates the current frame, thereby improving the global consistency of story generation. We evaluated our model on the PororoSV and FlintstonesSV datasets and obtained state-of-the-art FID scores, and the generated frames also demonstrate better storytelling in visuals. Tianyi Song, Jiuxin Cao, Kun Wang 0057, Bo Liu 0004 |
ICASSP | 4 |
| 2024 | All in One: Multi-task Prompting for Graph Neural Networks (Extended Abstract)
Xiangguo Sun, Hong Cheng 0001, Jia Li 0009, Bo Liu 0004, Jihong Guan |
IJCAI | 4 |
| 2024 | Consistencies are All You Need for Semi-supervised Vision-Language TrackingabstractVision-Language Tracking (VLT) requires locating a specific target in video sequences, given a natural language prompt and an initial object box. Despite recent advancements, existing approaches heavily rely on expensive and time-consuming human annotations. To mitigate this limitation, directly generating pseudo labels from raw videos seems to be a straightforward solution; however, it inevitably introduces undesirable noise during the training process. Moreover, we insist that an efficient tracker should excel in tracking the target, regardless of the temporal direction. Building upon these insights, we propose the pioneering semi-supervised learning scheme for VLT task, representing a crucial step towards reducing the dependency on high-quality yet costly labeled data. Specifically, drawing inspiration from the natural attributes of a video (i.e., space, time, and semantics), our approach progressively leverages inherent consistencies from these aspects: (1) Spatially, each frame and any object cropped from it naturally form an image-bbox (bounding box) pair for self-training; (2) Temporally, bidirectional tracking trajectories should exhibit minimal differences; (3) Semantically, the correlation between visual and textual features is expected to remain consistent. Furthermore, the framework is validated with a simple yet effective tracker we devised, named ATTracker (Asymmetrical Transformer Tracker). It modifies the self-attention operation in an asymmetrical way, striving to enhance target-related features while suppressing noise. Extensive experiments confirm that our ATTracker serves as a robust baseline, outperforming fully supervised base trackers. By unveiling the potential of learning with limited annotations, this study aims to attract attention and pave the way for Semi-supervised Vision-Language Tracking (SS-VLT). Jiawei Ge 0002, Jiuxin Cao, Xuelin Zhu, Xin-Yu Zhang 0027, Chang Liu 0113, Kun Wang 0057, Bo Liu 0004 |
ACM Multimedia | 7 |
| 2024 | METER: Multimodal Hallucination Detection with Mixture of Experts via Tools Ensembling and Reasoning
Ruwen Zhang, Jinglu Chen, Mingjie Dai, Bo Liu 0004, Jiuxin Cao |
NLPCC (5) | 6 |
| 2024 | EnsCLR: Unsupervised skeleton-based action recognition via ensemble contrastive learning of representation
Kun Wang 0057, Jiuxin Cao, Biwei Cao, Bo Liu 0004 |
Comput. Vis. Image Underst. | 4 |
| 2024 | Modeling group-level public sentiment in social networks through topic and role enhancement
Ruwen Zhang, Bo Liu 0004, Jiuxin Cao, Hantao Zhao, Xuheng Sun, Xiangguo Sun |
Knowl. Based Syst. | 2 |
| 2024 | GAL: combining global and local contexts for interpersonal relation extraction toward document-level Chinese text
Jiawei Ge 0002, Jiuxin Cao, Yingxing Bao, Biwei Cao, Bo Liu 0004 |
Neural Comput. Appl. | 5 |
| 2024 | Correction: GAL: combining global and local contexts for interpersonal relation extraction toward document-level Chinese text
Jiawei Ge 0002, Jiuxin Cao, Yingxing Bao, Biwei Cao, Bo Liu 0004 |
Neural Comput. Appl. | 5 |
| 2024 | Response Generation in Social Network With Topic and Emotion ConstraintsabstractResponse generation is the task of automatically generating human-like content based on the provided context. One of its prominent applications is to simulate realistic response content for social network posts. In the digital age, social network platforms play a vital role in information exchange and social interaction. This study focuses on response generation techniques for the platform of public opinion evolution simulation that simulate realistic response content, enabling a deeper understanding of the emotional expressions of network users. Recent advancements in deep learning techniques, particularly the sequence-to-sequence (Seq2Seq) model, have shown promise in the response generation field. However, we still face two challenges: content variety, topic and emotion relevancy. To this end, we propose the EmoTG-ETRS model which comprises three parts. The first is a response generation module based on Transformer architecture. Then, an auxiliary emotion improvement module is incorporated to enhance the emotional expressiveness of the response candidates. Finally, a reverse selection module, which combines maximum mutual information (MMI) evaluation, emotional expression evaluation, and topic consistency evaluation, is devised to select the highest-scoring response. Extensive experiments have been conducted to evaluate the effectiveness of the proposed model and the results demonstrate that the EmoTG-ETRS model improves the quality of produced replies in terms of topic consistency and emotional accuracy rate when compared with the SOTA research works. Biwei Cao, Jiuxin Cao, Bo Liu 0004, Jie Gui, Jun Zhou 0027, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | A Multi-Agent System for Fine-Grained Opinion Dynamics Analysis in Online Social NetworksabstractThe influence and dissemination of users’ opinions are the essence of the opinion dynamics in online social networks (OSNs). Understanding the process of users’ opinion formation and propagation can provide better service for public opinion monitoring and product advertising. However, most opinion dynamics research typically do not distinguish between user opinion formation and propagation, instead focusing on the process of mass opinion polarization. This article presents a multiagent system (MAS) to analyze fine-grained opinion dynamics (FOD) by agent-based modeling (ABM). At the macro level, the MAS-FOD system we designed can produce a statistical analysis of public opinion evolution based on different parameter constellations and analyze the changes in personalized agent opinions. At the micro level, the agent-based model is presented to describe the single user in OSNs as an individualized agent in MAS-FOD. Specifically, we propose two mechanisms, agent-based opinion formation (AOF) mechanism and agent-based opinion propagation (AOP) mechanism, for agents to form and disseminate opinions, respectively. Meanwhile, multidimensional social influence features mining from self, local neighbors, and global topic community are defined and used in two mechanisms to train personalized agents to self-adapt to the MAS-FOD system. We demonstrate the rationality and effectiveness of the MAS-FOD system from two types of experiments: empirical analysis and simulation analysis. The former is driven by empirical social network data to analyze the predictive performance of AOF; the latter simulates the public opinion propagation through the AOP mechanism in the MAS-FOD system. The results demonstrate that: 1) the AOF outperforms SOTA methods in the accuracy of agent opinion prediction; 2) the dissemination effect based on the AOP is more closer to the actual evolution trend of public opinion than the baseline; and 3) the MAS-FOD system can perform fine-grained analysis of opinion dynamics by adjusting different parameters. Huiyu Min, Jiuxin Cao, Jiawei Ge 0002, Bo Liu 0004 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Enhancing Micro-Video Venue Recognition via Multi-Modal and Multi-Granularity Object RelationsabstractMicro-video venue recognition aims to predict the venue category where a micro-video was filmed. Different from traditional long videos which contain rich temporal context, venue prediction for micro-videos is difficult due to its limited duration (generally within 6s). The existing works usually extract features of each modality from a global perspective for prediction, neglecting the semantics carried by local objects. To this end, we propose Multi-Modal and Multi-Granularity Object Relations (M2ORE) to address the above issues, which learns multi-granularity interactive semantics between venues and multimodal semantic objects to help understand venues. Specifically, M2ORE comprises of two modules: it first extract semantic objects of different modalities, i.e. visual objects in keyframes and keywords in texts, and models the affiliation relationship between semantic objects and venues and the co-occurrence relationship among semantic objects, forming a heterogeneous venue-object relation graph. Then, to achieve the interactive semantics between venues and objects from the relation graph, a novel Parallel-Graph Inference Model (Parallel-GIM) is proposed, which updates the representation of nodes through graph propagation and fuse multi-level features (local-global-multimodal) through the devised hierarchical attention mechanism. Finally, the probability distribution of venues can be obtained through a multi-layer perceptron with the comprehensive features of the venue nodes. Extensive experiments on real-world micro-video dataset demonstrate the superiority of the proposed M2ORE. Jiuxin Cao, Xuelin Zhu, Bo Liu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Semantic-Guided Representation Enhancement for Multi-Label Image ClassificationabstractMulti-label image classification is an essential yet challenging task that requires to recognize multiple objects of images. To this end, recent studies have sought to acquire visual representations for each label by attention models, and then train binary classifiers for prediction. However, these methods have two major drawbacks: 1) They rely heavily on the precise alignments between two modalities, which is still challenging for current attention models; 2) They ignore patch-level representations rich in local object features, which are also of great importance for label recognition. In this paper, we propose a semantic-guided representation enhancement framework, which augments patch-level representations with object-level representations for robust label recognition. Concretely, the proposed framework consists of two significant components: 1) an inter-modal attention module that accounts for coarsely locating object regions and producing object-level representations for each label; 2) an intra-modal attention module that aggregates object representations to enhance patch representations based on their correlations. In this way, both local clues and global glances of objects are fully exploited simultaneously, rather than relying solely on object-level representations obtained by the inter-modal attention, thus improving the performance of label recognition. Experimental results show that our framework outperforms the state-of-the-art methods by 0.5%, 0.6%, 0.7% and 0.8% in mAP on Pascal VOC 2007, Microsoft COCO, NUS-WIDE and Visual Genome datasets, respectively. Codes and models are available on https://github.com/jasonseu/SGRE. Xuelin Zhu, Jianshu Li, Jiuxin Cao, Dongqi Tang, Bo Liu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Nowhere to Hide: Online Rumor Detection Based on Retweeting Graph Neural NetworksabstractOnline rumor detection is crucial for a healthier online environment. Traditional methods mainly rely on content understanding. However, these contents can be easily adjusted to avoid such supervision and are insufficient to improve the detection result. Compared with the content, information propagation patterns are more informative to support further performance promotion. Unfortunately, learning the propagation patterns is difficult, since the retweeting tree is more topologically complicated than linear sequences or binary trees. In light of this, we propose a novel rumor detection framework based on structure-aware retweeting graph neural networks. To capture the propagation patterns, we first design a novel conversion method to transform the complex retweeting tree as more tractable binary tree without losing the reconstruction information. Then, we serialize the retweeting tree as a corpus of meta-tree paths, where each meta-tree can preserve a basic substructure. A deep neural network is then designed to integrate all meta-trees and to generate the global structural embeddings. Furthermore, we propose to integrate content, users, and propagation patterns to enhance more reliable performance. To this end, we propose a novel self-attention-based retweeting neural network to learn individual features from both content and users. We then fuse the node-level features with our global structural embeddings via a mutual attention unit. In this way, we can generate more comprehensive representations for rumor detection. Extensive evaluations on two real-world datasets show remarkable superiorities of our model compared with existing methods. Bo Liu 0004, Xiangguo Sun, Qing Meng, Xinyan Yang, Yang Lee, Jiuxin Cao, Junzhou Luo, Roy Ka-Wei Lee |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Multi-stage dynamic disinformation detection with graph entropy guidance
Xiaorong Hao, Bo Liu 0004, Xinyan Yang, Xiangguo Sun, Qing Meng, Jiuxin Cao |
World Wide Web (WWW) | 2 |
| 2023 | Scene-Aware Label Graph Learning for Multi-Label Image ClassificationabstractMulti-label image classification refers to assigning a set of labels for an image. One of the main challenges of this task is how to effectively capture the correlation among labels. Existing studies on this issue mostly rely on the statistical label co-occurrence or semantic similarity of labels. However, an important fact is ignored that the co-occurrence of labels is closely related with image scenes (indoor, outdoor, etc.), which is a vital characteristic in multi-label image classification. In this paper, a novel scene-aware label graph learning framework is proposed, which is capable of learning visual representations for labels while fully perceiving their co-occurrence relationships under variable scenes. Specifically, our framework is able to detect scene categories of images without relying on manual annotations, and keeps track of the co-occurring labels by maintaining a global co-occurrence matrix for each scene category throughout the whole training phase. These scene-independent co-occurrence matrices are further employed to guide the interactions among label representations in a graph propagation manner towards accurate label prediction. Extensive experiments on public benchmarks demonstrate the superiority of our framework. Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao |
ICCV | 5 |
| 2023 | All in One: Multi-Task Prompting for Graph Neural NetworksabstractRecently, "pre-training and fine-tuning'' has been adopted as a standard workflow for many graph tasks since it can take general graph knowledge to relieve the lack of graph annotations from each application. However, graph tasks with node level, edge level, and graph level are far diversified, making the pre-training pretext often incompatible with these multiple tasks. This gap may even cause a "negative transfer'' to the specific application, leading to poor results. Inspired by the prompt learning in natural language processing (NLP), which has presented significant effectiveness in leveraging prior knowledge for various NLP tasks, we study the prompting topic for graphs with the motivation of filling the gap between pre-trained models and various graph tasks. In this paper, we propose a novel multi-task prompting method for graph models. Specifically, we first unify the format of graph prompts and language prompts with the prompt token, token structure, and inserting pattern. In this way, the prompting idea from NLP can be seamlessly introduced to the graph area. Then, to further narrow the gap between various graph tasks and state-of-the-art pre-training strategies, we further study the task space of various graph applications and reformulate downstream problems to the graph-level task. Afterward, we introduce meta-learning to efficiently learn a better initialization for the multi-task prompt of graphs so that our prompting framework can be more reliable and general for different tasks. We conduct extensive experiments, results from which demonstrate the superiority of our method. Xiangguo Sun, Hong Cheng 0001, Jia Li 0009, Bo Liu 0004, Jihong Guan |
KDD | 4 |
| 2023 | Real-time anomaly detection on surveillance video with two-stream spatio-temporal generative model
Jiuxin Cao, Bo Liu 0004, Xuelin Zhu |
Multim. Syst. | 4 |
| 2023 | LPP2KL: Online Location Privacy Protection Against Knowing-and-Learning Attacks for LBSsabstractWhen user-centric location privacy protections are used by the public (e.g., published as a mobile application to Mac app store), they are found to be vulnerable to a novel kind of potential inference attacks. These attacks can utilize the open-use protection’s input–output pairs to closely study the protection’s structure and mechanism. We name this kind of attacks as knowing-and-learning (KL) attacks. Targeted at such potential but threatening deep learning-based KL attacks in local search services, we propose a novel location privacy protection framework LPP2KL. LPP2KL aims to generate an obfuscated location, which is robust to the aforementioned attacks, so that the user’s current location (real) can be replaced by the generated location (pseudo) when submitted to the untrusted search server. To achieve this, we first depict a feasible net structure to implement the deep KL attack and verify its power with extensive experiments. Second, to simulate the attack-and-defense process, we develop a novel network structure with an adversary net and a protection net. These nets play a mini–max game toward privacy inference in an adversarial manner until an equilibrium is reached. Meanwhile, the protection net also achieves the balance between preserved privacy and quality of service through utility-constrained privacy optimization. We introduce a novel penalty function to handle the utility-constrained privacy optimization problem. Extensive experiments are conducted on Yelp and FoursSwarm datasets, which validate the generality of our proposed LPP2KL in protection handling and illustrate its effectiveness in utility control under the acceptable quality of service. Zhuo Ma 0002, Bo Liu 0004, Jiuxin Cao |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | Attention-Fused Deep Relevancy Matching Network for Clickbait DetectionabstractClickbait is a sort of news that contains low-quality content but intriguing or overstated titles. Most of the existing studies usually utilize titles to detect clickbait and suffer from poor performance. However, from the perspective of user cognition, the essence of clickbait news is due to the fact that the body contents fall short of the titles’ expectations. So, in addition to the explicit features of titles, many extra informative signals could help to distinguish clickbait. In this article, we proposed a novel Multiview learning approach-based ClickBait Detection (MCBD) model that takes into account useful signals from both titles and body contents. From the view of titles, a multilayer gated convolutional network is employed to learn the local- and long-distance dependence relationships between words. From the view of body contents, we first extract the key topic sentences from the body content and then we developed a novel attention-fused deep relevance matching network (ARMN) to thoroughly mine the similarity between titles and body contents. Combined with the information from the views of titles and body contents, our model can distinguish well-hidden clickbait news. Finally, extensive experiments were conducted on two real-world datasets to demonstrate that the MCBD model outperforms the state-of-the-art models, improving the final performance by 3% in terms of the F1-score. The ablation studies demonstrate that the relevancy between titles and body contents can improve the model’s performance by 1.7% on the F1-score. Qing Meng, Bo Liu 0004, Xiangguo Sun, Chengyu Liang, Jiuxin Cao, Roy Ka-Wei Lee, Xing Bao |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2023 | In Your Eyes: Modality Disentangling for Personality Analysis in Short VideoabstractWith the dramatic growth of various short video platforms, users are more likely to share their social stream online and make their social connections stronger. To better understand their preferences, personality analysis has attracted more attention. Unlike single modal data such as text or images, which is hard to comprehensively uncover one’s personal traits, personality analysis on short video is verified to be much more accurate but also more challenging because of the huge gap between incompatible data modalities. We have noticed that the key problem is how to disentangle the complexity from multimodal data to find their consistency and uniqueness. In this article, we propose a novel video analysis framework for personality detection with visual, acoustic, and textual neural networks. Specifically, to enhance our model’s sensitivity to personality detection, we first propose three deep learning channels to learn modal features. The framework can not only extract each modal feature but also learn time-varying pattern via a temporal alignment network. To identify the consistency and uniqueness across multiple modalities, we creatively propose to maximize the similarity of common information learned by a shared neural network across multiple modalities and extend the distance of exclusive information learned by private networks of different modalities. Extensive experiments on the real-world dataset demonstrate that our model can outperform existing baselines. Xiangguo Sun, Bo Liu 0004, Liya Ai, Qing Meng, Jiuxin Cao |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2023 | Self-Supervised Hypergraph Representation Learning for Sociological AnalysisabstractModern sociology has profoundly uncovered many convincing social criteria for behavioral analysis. Unfortunately, many of them are too subjective to be measured and very challenging to be presented in online social networks (OSNs) for the large data volume and complicated environments to be explored. On the other hand, data mining techniques can better find data patterns but many of them leave behind unnatural understanding to humans. Although there are some works trying to integrate social observations for specific tasks, they are still hard to be applied to more general cases. In this paper, we propose a fundamental methodology to support the further fusion of data mining techniques and sociological behavioral criteria. Our highlights are three-fold: First, we propose an effective hypergraph awareness and a fast line graph construction framework. The hypergraph can more profoundly indicate the interactions between individuals and their environments because each edge in the hypergraph (a.k.a hyperedge) contains more than two nodes, which is perfect to describe social. A line graph treats each social environment as a super node with the underlying influence between different environments. In this way, we go beyond traditional pair-wise relations and explore richer patterns under various sociological criteria; Second, we propose a novel hypergraph-based neural network to learn social influence flowing from users to users, users to environments, environment to users, and environments to environments. The neural network can be learned via a task-free method, making our model very flexible to support various data mining tasks and sociological analysis; Third, we propose both qualitative and quantitive solutions to effectively evaluate the most common sociological criteria like social conformity, social equivalence, environmental evolving and social polarization. Our extensive experiments show that our framework can better support both data mining tasks for online user behaviors and sociological analysis. Xiangguo Sun, Hong Cheng 0001, Bo Liu 0004, Jia Li 0009, Hongyang Chen 0001, Guandong Xu, Hongzhi Yin |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Structure Learning Via Meta-Hyperedge for Dynamic Rumor DetectionabstractOnline social networks have greatly facilitated our lives but have also propagated the spreading of rumours. Traditional works mostly find rumors from content, but content can be strategically manipulated to evade such detection, making these methods brittle. To improve the accuracy and robustness of rumor detection, we propose to integrate and exploit the content, propagation structure, and temporal relations because information in the networks always spreads dynamically with significant structures. In this paper, we propose a novel rumor detection framework in online temporal networks via structure learning. Specifically, to exploit the propagation structure, we propose a novel hyperedge walking strategy on a meta-hyperedge graph to learn the representations of sub-structures in the networks. Then a hyperedge expansion method is proposed to generate more global structural features. The expanded hyperedges are more hierarchical, making the learned structural embeddings more expressive. To make full use of content, we design a hypergraph learning model using hyperedge expansion to fuse node content with structural features and generate comprehensive representations for the entire graph. To exploit temporal relations, we design a masked temporal attention unit for learning the evolving patterns of the network. Extensive evaluations with six state-of-the-art baselines on two real-world datasets demonstrate the superiority of our solution. Xiangguo Sun, Hongzhi Yin, Bo Liu 0004, Qing Meng, Jiuxin Cao, Alexander Zhou 0001, Hongxu Chen 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | AlignVE: Visual Entailment Recognition Based on Alignment RelationsabstractVisual entailment (VE) is to recognize whether the semantics of a hypothesis text can be inferred from the given premise image, which is one special task among recent emerged vision and language understanding tasks. Currently, most of the existing VE approaches are derived from the methods of visual question answering. They recognize visual entailment by quantifying the similarity between the hypothesis and premise in the content semantic features from multi modalities. Such approaches, however, ignore the VE's unique nature of relation inference between the premise and hypothesis. Therefore, in this paper, a new architecture called AlignVE is proposed to solve the visual entailment problem with a relation interaction method. It models the relation between the premise and hypothesis as an alignment matrix. Then it introduces a pooling operation to get feature vectors with a fixed size. Finally, it goes through the fully-connected layer and normalization layer to complete the classification. Experiments show that our alignment-based architecture reaches 72.45% accuracy on SNLI-VE dataset, outperforming previous content-based models under the same settings. Biwei Cao, Jiuxin Cao, Jie Gui, Jiayun Shen, Bo Liu 0004, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Multim. | 5 |
| 2023 | Recognize News Transition from Collective Behavior for News RecommendationabstractIn the news recommendation, users are overwhelmed by thousands of news daily, which makes the users’ behavior data have high sparsity. Therefore, only considering a single user’s personalized preferences cannot support the news recommendation. How to improve the relatedness of news and users and reduce data sparsity has become a hot issue. Recent studies have attempted to use graph models to enrich the relationship between users and news, but they are still limited to modeling the historical behaviors of a single user. To fill the gap, we integrate user-news relationships and the overall user historical clicked news sequences to construct a global heterogeneous transition graph. And a refinement approach is proposed to recognize the news transition patterns in the graph. Based on the global heterogeneous transition graph, we propose a heterogeneous transition graph attention network to capture the common behavior patterns of most users to enhance the representation of user interest. Fusing the users’ personalized and common interest, we propose the GAINRec model to recommend news effectively. Extensive experiments are conducted on two public news recommendation datasets, and the results show the superiority of the proposed GAINRec model compared with the state-of-the-art news recommendation models. The implementation of our model is available at https://github.com/newsrec/GAINRec . Qing Meng, Bo Liu 0004, Xiangguo Sun, Mingrui Hu, Jiuxin Cao |
ACM Trans. Inf. Syst. | 3 |
| 2022 | CORN: Co-Reasoning Network for Commonsense Question AnsweringabstractCommonsense question answering (QA) requires machines to utilize the QA content and external commonsense knowledge graph (KG) for reasoning when answering questions. Existing work uses two independent modules to model the QA contextual text representation and relationships between QA entities in KG, which prevents information sharing between modules for co-reasoning. In this paper, we propose a novel model, Co-Reasoning Network (CORN), which adopts a bidirectional multi-level connection structure based on Co-Attention Transformer. The structure builds bridges to connect each layer of the text encoder and graph encoder, which can introduce the QA entity relationship from KG to the text encoder and bring contextual text information to the graph encoder, so that these features can be deeply interactively fused to form comprehensive text and graph node representations. Meanwhile, we propose a QA-aware node based KG subgraph construction method. The QA-aware nodes aggregate the question entity nodes and the answer entity nodes, and further guide the expansion and construction process of the subgraph to enhance the connectivity and reduce the introduction of noise. We evaluate our model on QA benchmarks in the CommonsenseQA and OpenBookQA datasets, and CORN achieves state-of-the-art performance. Biwei Cao, Qingqing Gao, Zheng Yin, Bo Liu 0004, Jiuxin Cao |
COLING | 5 |
| 2022 | Two-Stream Transformer for Multi-Label Image ClassificationabstractMulti-label image classification is a fundamental yet challenging task in computer vision that aims to identify multiple objects from a given image. Recent studies on this task mainly focus on learning cross-modal interactions between label semantics and high-level visual representations via an attention operation. However, these one-shot attention based approaches generally perform poorly in establishing accurate and robust alignments between vision and text due to the acknowledged semantic gap. In this paper, we propose a two-stream transformer (TSFormer) learning framework, in which the spatial stream focuses on extracting patch features with a global perception, while the semantic stream aims to learn vision-aware label semantics as well as their correlations via a multi-shot attention mechanism. Specifically, in each layer of TSFormer, a cross-modal attention module is developed to aggregate visual features from spatial stream into semantic stream and update label semantics via a residual connection. In this way, the semantic gap between two streams gradually narrows as the procedure progresses layer by layer, allowing the semantic stream to produce sophisticated visual representations for each label towards accurate label recognition. Extensive experiments on three visual benchmarks, including Pascal VOC 2007, Microsoft COCO and NUS-WIDE, consistently demonstrate that our proposed TSFormer achieves state-of-the-art performance on the multi-label image classification task. Xuelin Zhu, Jiuxin Cao, Jiawei Ge 0002, Bo Liu 0004 |
ACM Multimedia | 5 |
| 2022 | Category-universal witness discovery with attention mechanism in social network
Jiuxin Cao, Yuntao Yang, Bo Liu 0004, Qingqing Gao |
Inf. Process. Manag. | 4 |
| 2022 | Emotion recognition in conversations with emotion shift detection based on multi-task learning
Qingqing Gao, Biwei Cao, Tianyun Gu, Xing Bao, Junyan Wu, Bo Liu 0004, Jiuxin Cao |
Knowl. Based Syst. | 7 |
| 2022 | Temporal-aware and multifaceted social contexts modeling for social recommendation
Qing Meng, Bo Liu 0004, Xuheng Sun, Jiuxin Cao, Roy Ka-Wei Lee |
Knowl. Based Syst. | 2 |
| 2021 | Understanding User Topic Preferences across Multiple Social NetworksabstractIn recent years, social networks have shown diversity in function and applications. People begin to use multiple online social networks simultaneously for different demands. The ability to uncover a user’s latent topic and social network preference is critical for community detection, recommendation, and personalized service across social networks. Unfortunately, most current works focus on the single network, necessitating new technology and models to address this issue. This paper proposes a user preference discovery model on multiple social networks. Firstly, the global and local topic concepts are defined, then a latent semantic topic discovery method is used to obtain global and local topic word distributions, along with user topic and social network preferences. After that, the topic distribution characteristics of different social networks are examined, as well as the reasons why users choose one network over another to create a post. Next, a Gibbs sampling algorithm is adopted to obtain the model parameters. In the experiment, we collect data from Twitter, Instagram, and Tumblr websites to build a dataset of multiple social networks. Finally, we compare our research to previous works, and both qualitative and quantitative evaluation results have demonstrated the effectiveness. Jiuxin Cao, Huiyu Min, Bo Liu 0004 |
IEEE BigData | 5 |
| 2021 | Heterogeneous Hypergraph Embedding for Graph ClassificationabstractRecently, graph neural networks have been widely used for network embedding because of their prominent performance in pairwise relationship learning. In the real world, a more natural and common situation is the coexistence of pairwise relationships and complex non-pairwise relationships, which is, however, rarely studied. In light of this, we propose a graph neural network-based representation learning framework for heterogeneous hypergraphs, an extension of conventional graphs, which can well characterize multiple non-pairwise relations. Our framework first projects the heterogeneous hypergraph into a series of snapshots and then we take the Wavelet basis to perform localized hypergraph convolution. Since the Wavelet basis is usually much sparser than the Fourier basis, we develop an efficient polynomial approximation to the basis to replace the time-consuming Laplacian decomposition. Extensive evaluations have been conducted and the experimental results show the superiority of our method. In addition to the standard tasks of network embedding evaluation such as node classification, we also apply our method to the task of spammers detection and the superior performance of our framework shows that relationships beyond pairwise are also advantageous in the spammer detection. To make our experiment repeatable, source codes and related datasets are available at https://xiangguosun.mystrikingly.com Xiangguo Sun, Hongzhi Yin, Bo Liu 0004, Hongxu Chen 0002, Jiuxin Cao, Yingxia Shao, Nguyen Quoc Viet Hung |
WSDM | 3 |
| 2021 | Multi-level Hyperedge Distillation for Social Linking Prediction on Sparsely Observed NetworksabstractSocial linking prediction is one of the most fundamental problems in online social networks and has attracted researchers’ persistent attention. Most of the existing works predict unobserved links using graph neural networks (GNNs) to learn node embeddings upon pair-wise relations. Despite promising results given enough observed links, these models are still challenging to achieve heart-stirring performance when observed links are extremely limited. The main reason is that they only focus on the smoothness of node representations on pair-wise relations. Unfortunately, this assumption may fall when the networks do not have enough observed links to support it. To this end, we go beyond pair-wise relations and propose a new and novel framework using hypergraph neural networks with multi-level hyperedge distillation strategies. To break through the limitations of sparsely observed links, we introduce the hypergraph to uncover higher-level relations, which is exceptionally crucial to deduce unobserved links. A hypergraph allows one edge to connect multiple nodes, making it easier to learn better higher-level relations for link prediction. To overcome the restrictions of manually designed hypergraphs, which is constant in most hypergraph researches, we propose a new method to learn high-quality hyperedges using three novel hyperedges distillation strategies automatically. The generated hyperedges are hierarchical and follow the power-law distribution, which can significantly improve the link prediction performance. To predict unobserved links, we present a novel hypergraph neural networks named HNN. HNN takes the multi-level hypergraphs as input and makes the node embeddings smooth on hyperedges instead of pair-wise links only. Extensive evaluations on four real-world datasets demonstrate our model’s superior performance over state-of-the-art baselines, especially when the observed links are extremely reduced. Xiangguo Sun, Hongzhi Yin, Bo Liu 0004, Hongxu Chen 0002, Qing Meng, Wang Han, Jiuxin Cao |
WWW | 3 |
| 2021 | Asymptotically Optimal Online Scheduling With Arbitrary Hard Deadlines in Multi-Hop Communication NetworksabstractThis paper firstly proposes a greedy online packet scheduling algorithm for the problem raised by Mao, Koksal and Shroff that allows arbitrary hard deadlines in multi-hop networks aiming at maximizing the total revenue. With the same assumption of$\rho _{M} / \rho _{m}={O}(1)$where$\rho _{M}$and$\rho _{\textit {}m}$are the maximum and minimum revenue a packet may carry, our algorithm is${O}$(${P} _{M}$)-competitive improving on MKS algorithm by a factor of${O}$(log${P} _{M}$), where${P} _{M}$is the length of the longest path a packet may travel in the network. We prove that it is asymptotically optimal by presenting a lower bound of${P} _{M}$on the competitiveness for this problem. Secondly, this paper studies the extension of this problem that includes routing as a part of the solution. We prove that using the fastest path algorithm for the routing part, the greedy online algorithm also achieves asymptotically optimal competitiveness for the extended problem. Furthermore, we present a non-greedy online algorithm that not only is asymptotically optimal, but also can adaptively achieve a better competitiveness when the network has a larger${C} _{\textit {min}}$, where${C} _{\textit {min}}$is the minimum link capacity in the network. Finally, simulation results are reported, showing that not only do the greedy online algorithms achieve asymptotically optimal bounds, but also practically achieve better performance than the previously proposed algorithms. Bo Liu 0004, Xiaojun Shen 0002 |
IEEE/ACM Trans. Netw. | 2 |
| 2020 | Towards tenant demand-aware bandwidth allocation strategy in cloud datacenter
Jiuxin Cao, Zhuo Ma 0002, Jue Xie, Xiangying Zhu, Fang Dong 0001, Bo Liu 0004 |
Future Gener. Comput. Syst. | 6 |
| 2020 | Co-Detection of crowdturfing microblogs and spammers in online social networks
Bo Liu 0004, Xiangguo Sun, Zeyang Ni, Jiuxin Cao, Junzhou Luo, Benyuan Liu, Xinwen Fu |
World Wide Web | 1 |
| 2020 | Topic based time-sensitive influence maximization in online social networks
Huiyu Min, Jiuxin Cao, Tangfei Yuan, Bo Liu 0004 |
World Wide Web | 4 |
| 2020 | Group-level personality detection based on text generated networks
Xiangguo Sun, Bo Liu 0004, Qing Meng, Jiuxin Cao, Junzhou Luo, Hongzhi Yin |
World Wide Web | 2 |
| 2020 | Survey on user location prediction based on geo-social networking data
Xiaoming Fu 0001, Jiuxin Cao, Bo Liu 0004 |
World Wide Web | 4 |
| 2019 | Quantifying Group Influence on Individuals in Online Social NetworksabstractAccording to social psychology studies, social networks are strongly subject to group influence; that is, users' behaviors or sentiment are primarily influenced by their group environments. But only a few studies to date have examined group influence in online social networks (OSNs). In this paper, a Factor Graph-based Group Influence model (F2GI) is proposed to depict the influence of perceived groups on a user. First, we analyze individuals' group environments and propose definitions and detecting methods of two types of groups: following groups and interacting groups. Secondly, based on these groups, we quantity the groups' features and propose the F2GI model able to describe group influence. Finally, using the retweet behaviors as the primarily manifestations of group influence, we apply our model to predict individuals' retweet behaviors and to validate how much influence the group puts on individuals. The results show that our group influence model can depict group influence more effectively and predict the actions of users who are affected by perceived groups accurately with social psychology ideas than classical machine learning methods. We believe that, moreover, our model can also be applied to marketing strategies, advertisement strategies and the prediction of public opinion. Qing Meng, Junzhou Luo, Bo Liu 0004, Xiangguo Sun, Jiuxin Cao |
ISCC | 3 |
| 2019 | Joint Visual-Textual Sentiment Analysis Based on Cross-Modality Attention Mechanism
Xuelin Zhu, Biwei Cao, Bo Liu 0004, Jiuxin Cao |
MMM (1) | 4 |
| 2019 | Analysis of and defense against crowd-retweeting based spam in social networks
Bo Liu 0004, Zeyang Ni, Junzhou Luo, Jiuxin Cao, Xudong Ni, Benyuan Liu, Xinwen Fu |
World Wide Web | 1 |
| 2018 | An Application-Driven-based Network Resource Control MethodabstractWith the expansion of the scale of Internet users and the increasing of network application types, the motive force of the network development has gradually changed from technologies to applications, and the network resources need be customized for different applications in order to guarantee different QoS. However, network resources are always scarce for application requirements, and the existing works only attach importance to guarantee QoS of network applications and ignore the promotion of the network resources utilization. So, an Application-Driven-based Network Resource Control Method (ADNRC) is presented in this paper. The thought of the separation of the transmission and the control based on SDN is introduced in this method, and the network slices for different applications are constructed to customize the network resources and guarantee QoS of network applications. At the same time, the multi-path routes between the sources and the destinations are generated during the procedure of constructing the network slices, and the network flows are scheduled in real time based on the multi-path routes and the network status to enhance the utilization of network resources. The experimental results show that the proposed method in this paper is better than existing methods in the aspect of both the QoS guarantee of network applications and the promotion of the network resources utilization. Wei Li 0017, Runhuan Zhang, Wenjiang Ding, Bo Liu 0004, Junzhou Luo |
CSCWD | 4 |
| 2018 | Who Am I? Personality Detection Based on Deep Learning for TextsabstractRecently, personality detection based on texts from online social networks has attracted more and more attentions. However, most related models are based on letter, word or phrase, which is not sufficient to get good results. In this paper, we present our preliminary but interesting and useful research results to show that the structure of texts can be also an important feature in the study of personality detection from texts. We propose a model named 2CLSTM, which is a bidirectional LSTMs (Long Short Term Memory networks) concatenated with CNN (Convolutional Neural Network), to detect user's personality using structures of texts. Besides, a concept, Latent Sentence Group (LSG), is put forward to express the abstract feature combination based on closely connected sentences and we use our model to capture it. To the best of our knowledge, most related works only conducted their experiments on one data set, which may not well explain the versatility of their models. We implement our evaluations on two different kinds of datasets, containing long texts and short texts. Evaluations on both datasets have achieved better results, which demonstrate that our model can efficiently learn valid text structure features to accomplish the task. Xiangguo Sun, Bo Liu 0004, Jiuxin Cao, Junzhou Luo, Xiaojun Shen 0002 |
ICC | 2 |
| 2018 | Community Discovery Based on Social Relations and Temporal-Spatial Topics in LBSNs
Jiuxin Cao, Xuelin Zhu, Bo Liu 0004 |
PAKDD (3) | 5 |
| 2018 | Effective fine-grained location prediction based on user check-in pattern in LBSNs
Jiuxin Cao, Xuelin Zhu, Renjun Lv, Bo Liu 0004 |
J. Netw. Comput. Appl. | 5 |
| 2018 | GLPP: A Game-Based Location Privacy-Preserving Framework in Account Linked Mixed Location-Based ServicesabstractIn Location-Based Services (LBSs) platforms, such as Foursquare and Swarm, the submitted position for a share or search leads to the exposure of users’ activities. Additionally, the cross-platform account linkage could aggravate this exposure, as the fusion of users’ information can enhance inference attacks on users’ next submitted location. Hence, in this paper, we propose GLPP, a personalized and continuous location privacy-preserving framework in account linked platforms with different LBSs (i.e., search-based LBSs and share-based LBSs). The key point of GLPP is to obfuscate every location submitted in search-based LBSs so as to defend dynamic inference attacks. Specifically, first, possible inference attacks are listed through user behavioral analysis. Second, for each specific attack, an obfuscation model is proposed to minimize location privacy leakage under a given location distortion, which ensures submitted locations’ utility for search-based LBSs. Third, for dynamic attacks, a framework based on zero-sum game is adopted to joint specific obfuscation above and minimize the location privacy leakage to a balanced point. Experiments on real dataset prove the effectiveness of our proposed attacks in Accuracy, Certainty, and Correctness and, meanwhile, also show the performance of our preserving solution in defense of attacks and guarantee of location utility. Zhuo Ma 0002, Jiuxin Cao, Xiusheng Chen, Bo Liu 0004, Yuntao Yang |
Secur. Commun. Networks | 5 |
| 2016 | On crowd-retweeting spamming campaign in social networksabstractCrowdsourcing is often used to solicit contributions from an online community for ideas, evaluation and opinions. However, spamming can pollute such a system and manipulate the results of crowdsourcing. For detection of those spammers, the training data used in previous studies is often derived by experts labeling collected data and manually identifying spammers. The reliability of such training data is questionable. In this paper, we utilize two web based service providers Zhubajie (ZBJ) and Sandaha (SDH) and obtain reliable data about the spammers. We use such data to investigate the crowd-retweeting spam in Sina Weibo. We analyze profile features, social relationship and retweeting behavior of such spammers. We find that although these spammers are likely to connect more closely than legitimate users, the underlying social tie is different from the social relationship in other spam campaigns because of the unique retweeting features with the information cascade effect. Based on these findings, we propose retweeting-aware link based ranking algorithms to detect suspect spam accounts using seeds of identified spammers. Our evaluation shows that our algorithm is more effective than other link-based methods. Bo Liu 0004, Junzhou Luo, Jiuxin Cao, Xudong Ni, Benyuan Liu, Xinwen Fu |
ICC | 1 |
| 2016 | Attribute-Based Influence Maximization in Social Networks
Jiuxin Cao, Dan Dong, Zhuo Ma 0002, Bo Liu 0004 |
WISE (1) | 7 |
| 2016 | A mobile phone-based physical-social location proof system for mobile social network serviceabstractAbstract Location‐related mobile social network services are popular nowadays, and their methods to obtain end users’ location information are based on people's self‐report location claims, using mobile devices to check positions and send them back to the service providers. However, this mechanism has a serious vulnerability that makes malicious users be able to access restricted resource by transmitting fake locations. Both academic and industrial researchers are recently aware of this problem's importance since the commercialized trend of location‐related mobile social network services. To address this issue, we propose mobile phone‐based physical‐social location, a mobile phone‐based location proof system to verify users’ location claims and defend various fake location information. Our core idea is that a user's location claim can be proved by a set of selective physical encountered people serving as “witnesses” who are co‐located with him/her in that area. The system is composed of proof generation and verification. In the proof generation phase, we leverage a certain number of co‐located people to generate certificates as location proofs during their encounters via bluetooth interface. In the verification phase, we propose an efficient verification scheme to make our system accurate and adaptive. We have implemented the MPSL system using real world Nokia N82 (Nokia, Espoo, Finland) phones. Our experimental results show that our mobile phone‐based system can achieve high verification accuracy and good performance. Copyright © 2014 John Wiley & Sons, Ltd. Xudong Ni, Junzhou Luo, Boying Zhang, Jin Teng, Xiaole Bai, Bo Liu 0004, Dong Xuan |
Secur. Commun. Networks | 6 |
| 2015 | Location-Based Influence Maximization in Social NetworksabstractIn this paper, we aim at the product promotion in O2O model and carry out the research of location-based influence maximization on the platform of LBSN. As offline consuming behavior exists under the O2O environment, the traditional online influence diffusion model could not describe the product acceptance accurately. Moreover, the existing researches of influence maximization tend to only concern on the online network of relationships but rarely take the offline part into consideration. This paper introduces the location property into the influence maximization to accord with the characteristic of O2O model. Firstly, we propose an improved influence diffusion model called TP Model which could accurately describe the process of accepting products under the O2O environment. Meanwhile, the definition of location-based influence maximization is presented. Then the user mobility pattern is analyzed and the calculation method of offline probability is designed. Considering the influence ability, a location-based influence maximization algorithm named TPH is proposed. Experiments prove TPH algorithm has general advantage. Finally, focusing on the performance of TPH algorithm under special circumstances, MR algorithm is designed as complement and experiments also verify its high effectiveness. Jiuxin Cao, Bo Liu 0004, Junzhou Luo |
CIKM | 3 |
| 2015 | On Computing Multi-Agent Itinerary Planning in Distributed Wireless Sensor Networks
Bo Liu 0004, Jiuxin Cao, Wei Yu 0002, Benyuan Liu, Xinwen Fu |
WASA | 1 |
| 2014 | Heterogeneous Metric Learning with Content-Based Regularization for Software Artifact RetrievalabstractThe problem of software artifact retrieval has the goal to effectively locate software artifacts, such as a piece of source code, in a large code repository. This problem has been traditionally addressed through the textual query. In other words, information retrieval techniques will be exploited based on the textual similarity between queries and textual representation of software artifacts, which is generated by collecting words from comments, identifiers, and descriptions of programs. However, in addition to these semantic information, there are rich information embedded in source codes themselves. These source codes, if analyzed properly, can be a rich source for enhancing the efforts of software artifact retrieval. To this end, in this paper, we develop a feature extraction method on source codes. Specifically, this method can capture both the inherent information in the source codes and the semantic information hidden in the comments, descriptions, and identifiers of the source codes. Moreover, we design a heterogeneous metric learning approach, which allows to integrate code features and text features into the same latent semantic space. This, in turn, can help to measure the artifact similarity by exploiting the joint power of both code and text features. Finally, extensive experiments on real-world data show that the proposed method can help to improve the performances of software artifact retrieval with a significant margin. Liang Wu 0011, Liang Du 0003, Bo Liu 0004, Guandong Xu, Yong Ge 0001, Yanjie Fu, Yuanchun Zhou, Hui Xiong 0001 |
ICDM | 3 |
| 2013 | Max-Cut based overlapping channel assignment for 802.11 multi-radio wireless mesh networksabstractDue to the limited number of orthogonal channels in a multi-radio multi-channel wireless mesh network (MR-WMN), overlapping channel assignment (CA) is one of the main factors that greatly affect the network capacity. In this paper, we first propose a model for measuring achieved network capacity in MR-WMNs. Then we prove that finding an optimal overlapping CA in a given MR-WMN with odd number of channels, is equivalent to finding an optimal assignment by only using its orthogonal channels. This theory allows us to use fewer channels to solve complicated CA problems. Third, we prove that in 802.11b/g-based MR-WMN, the simplified optimization problem is a Max-3-Cut problem. Although this problem is NP-hard, it has an efficient approximation algorithm that achieves approximation ratio of 1.19616 probabilistically by using the algorithm for Max-Cut whose approximation ratio is 1.1383 probabilistically. Based on the algorithm for Max-Cut, this paper proposes Max-Cut-based channel assignment (MCCA) which uses a heuristic method to adjust the result produced by the Max-Cut algorithm to achieve an even better result. Finally, we perform extensive simulations to compare the MCCA with a state-of-the-art Tabu-Search based algorithm. The results show that the Max-Cut-based overlapping CA algorithm effectively improves on the network capacity. Wei Wang 0089, Bo Liu 0004, Ming Yang 0001, Junzhou Luo, Xiaojun Shen 0002 |
CSCWD | 2 |
| 2013 | Execution Recovery in Transactional Composite ServiceabstractIn the composite service which runs for a long time under the heterogeneous and loose-coupled circumstance, the failure of service tends to occur. The transaction and recovery mechanism is urgently needed in order to guarantee the end-to-end QoS of the workflow and satisfy the user requirement. In this paper we address the composite service recovery issue in the way of substitution with the consideration of QoS constraint, based on our previous research work of the transactional construction and processing rules and the global optimization service selection algorithm, TSSA. Firstly, the service execution graph (SEG) is introduced and a service execution solution selection algorithm is proposed to choose one from the solution set of TSSA which has the highest success rate of recovery. Then, the concepts of execution backup path and switch cost are introduced and a search algorithm is presented to search for the optimal backup path when current service failure occurs. Meanwhile, a local induction algorithm based on positive feedback is described which could rapidly construct an transactional execution path when no backup execution path can be found in SEG and also guarantees the transactional constraint of composite service. Finally, experimental results show the recovery algorithm proposed in this paper is efficient and has high reliability. Jiuxin Cao, Gongrui Zhu, Bo Liu 0004, Junzhou Luo |
ICWS | 4 |
| 2012 | TASS: Transaction Assurance in Service SelectionabstractAs there are various risks of failure when Web Services are deployed in unreliable environment, the execution of a composite service requires the assurance of the transaction mechanism. However, existing QoS-aware composition approaches do not consider the transactional constraints during service selection. We address this issue by considering the combination of transactional and QoS requirements. Firstly, the novel construction and processing rules are proposed to guarantee the atomic consistency of the composite service and the correctness of these rules is proved subsequently. Then, on basis of these rules, an Ant Colony System based service selection algorithm is presented to guarantee the end-to-end QoS constraints on the premise of ensuring the atomic consistency during service selection. Therein, an optimization strategy is suggested to shrink the searching space of the algorithm tremendously. Finally, experimental results show the efficiency and effectiveness of the algorithm and demonstrate further the correctness of the construction and processing rules through simulations. Jiuxin Cao, Gongrui Zhu, Bo Liu 0004, Fang Dong 0001 |
ICWS | 4 |
| 2011 | QoS Preference-Aware Replica Selection Strategy Using MapReduce-Based PGA in Data GridsabstractData replication is an important technique to reduce access latency and bandwidth consumption in Grid environment. As one of the major functions of data replication, replica selection determines the best replica according to some specific criteria in Data Grid environment, where the data resources are limited and Grid users compete for these resources. In this paper, we focus mainly on a novel QoS preference-aware replica selection strategy which will meet individual QoS sensitivity (IQS) constraints for different users/applications. We first present a framework that characterize QoS properties of replica services and establish its mathematical model by introducing quantification methods. In order to deal with the IQS constraints and to perceive Grid users' QoS preferences accurately, we propose a QoS preference acquisition algorithm based on Analytic Hierarchy Process (AHP). We then design and implement a novel effective and efficient parallel genetic algorithm (PGA) based on Map Reduce paradigm for optimizing the objective function which corresponds to the optimal replica. Simulation results show that our strategy has a better performance in validity as well as scalability, and the optimal replica can always be obtained for Grid users with different IQS constraints under Data Grid environments that vary in system loads, scheduling strategies and user types. Runqun Xiong, Junzhou Luo, Aibo Song, Bo Liu 0004, Fang Dong 0001 |
ICPP | 4 |
| 2010 | Efficient Multi-objective Services Selection Algorithm Based on Particle Swarm OptimizationabstractWith the development of Web Service, it has become a key issue to select appropriate services from a large number of candidates for creating complex composite services according to users' different QoS levels requirements. However, the existing service selection algorithms have many defects such as high time complexity, non-global optimal solutions, and poor quality solutions. To solve these defects, an efficient multi-objective services selection algorithm, EMOSS, is proposed in this paper based on particle swarm optimization. The essence of EMOSS is to model the service selection problem as a constrained multi-objective optimization problem. First the services in each sub-service set are sorted by their concept of domination, then the new sub-service set nSi, whose size is far less than the original one, is constructed and finally output pare to optimal set. The theoretical analysis and experimental results show that EMOSS can effectively obtain high quality solutions. Jiuxin Cao, Xuesheng Sun, Bo Liu 0004 |
APSCC | 4 |
| 2010 | Semantic-based self-organizing mechanism for service registry and discoveryabstractService-oriented computing is the new paradigm for distributed applications. A key issue of utilizing web services is to design a scalable service discovery mechanism. Current discovery method based on centralized registries can easily suffer from performance bottleneck in the service network of large scale. To address such problem, this paper presents a dynamic and scalable mechanism for discovery and registry of semantic web services. Registry proxies are introduced to achieve the data distribution of web service. Behavior of registry proxies are carefully designed to maintain the overall structure in a self-organization way. A distributed balance-aware algorithm is designed to improve the system performance. The result of simulation experiment shows that the proposed mechanism can serve as a scalable solution for semantic web service publication and discovery. Jiuxin Cao, Bo Liu 0004 |
CSCWD | 4 |
| 2010 | Multi-agent based QoS-aware Service CompositionabstractService composition is the current research focus in the field of Service-Oriented Computing. However, the service discovery and selection mechanisms are static and not flexible in existing approaches on service composition, and the end-to-end QoS of a composite service can not also be ensured. In this paper, Multi-agent based QoS-aware Service Composition solution (MQSC) is presented, in which the concept of user satisfaction degree is introduced to depict the QoS of the service and the composite service execution mechanism based on Directed Acyclic Graph (DAG) is presented in detail. Through the collaboration between agents, MQSC not only can provide a mechanism for the dynamic service composition but also can ensure the end-to-end QoS of the composite service. Wei Li 0017, Junzhou Luo, Bo Liu 0004, Jiuxin Cao |
SMC | 3 |
| 2010 | Agent-based task representation and processing in pervasive computing environmentabstractAbstract It is an effective approach to adopt software agent in pervasive computing (Per‐Com) environment. Meanwhile, task execution is the key problem which is affected directly by task representation method in Per‐Com environment. The task of agent in Per‐Com environment is dynamic and changeful, so the task representation should be flexible and agent‐oriented. However, most of the existing task representation methods cannot present task relationship clearly and lack unified format, which leads to a drawback that the corresponding task processing efficiency is not satisfying. In this paper, firstly, an agent‐based Per‐Com architecture is presented in which the agent is endowed with certain role. Secondly, through analyzing task relationship, an agent role based task relationship tree (TRT) model and an XML‐based agent task representation method are proposed. Thirdly, based on TRT, an agent task decomposing and online multi‐agent scheduling algorithm are presented to solve time‐optimal problems. Meanwhile, a load balancing strategy based on probability is proposed to increase agent utilization. Finally, a prototype of an application on traffic monitoring is presented. The results of evaluation indicate that the architecture is feasible and the simulation tests of time performance indicate that the proposed TRT‐based task processing methods have a better performance than others. Copyright © 2009 John Wiley & Sons, Ltd. Bo Liu 0004, Junzhou Luo, Jiuxin Cao |
Wirel. Commun. Mob. Comput. | 1 |
| 2008 | SQUARE: A New TCP Variant for Future High Speed and Long Delay EnvironmentsabstractThe increasing diversity of Internet application requirements has spurred recent interest in transport protocol for high speed delay product connectivity. Addressing the deficiency of previous protocols, this paper presents a spectrum of time based odd function congestion control protocols for deployment in high speed and long distance networks. Through extensive analysis we focus on SQUARE, one of the protocols in odd function congestion control. Without the need to revise the current end-to-end architecture, SQUARE is shown to be efficient when there is bandwidth available, to be fair when many flows compete with each other for the same bottleneck, to be friendly when deployed with the conventional TCP, to be robust when there are oscillations in the network. Yansheng Qu, Junzhou Luo, Wei Li 0017, Bo Liu 0004, Laurence T. Yang |
AINA | 4 |
| 2006 | Semi-online Task Allocation Algorithm among Cooperative AgentsabstractTask allocation algorithm has great influence on the efficiency multi-agent task system. The performance of the existing allocation algorithms will decline with the increasing of task complexity. So a task allocation framework is presented and a semi-online multi-agent task allocation algorithm(SOAL) based on dependences of sub-tasks is proposed, the relationship of dependences among sub-tasks is the partial knowledge for SOAL. In contrast to the existing approaches, the performance of SOAL is more close to the optimal offline algorithm. The competitive analysis results and the tests of time performance demonstrate the advantage of SOAL Bo Liu 0004, Junzhou Luo, Wei Li 0017 |
CSCWD | 1 |
| 2005 | Multi-Agent Based Network Management Task Decomposition and SchedulingabstractThe rapid development of Internet makes network management on large-scale network a critical issue. But with the management task of large-scale network becoming more complicated, neither centralized network management nor agent based network management can satisfy the increasing demands. This paper presents a network management framework to support dynamic scheduling decisions. In this framework, some algorithms are proposed to decompose the whole network management task into several groups of sub-tasks. During the course of decomposition, different priorities are assigned to sub-tasks. Then based on the priorities of these sub-tasks, the strategies of agent scheduling are established. Priority-ranked sub-tasks are grouped according to their inter-dependences. Sub-tasks with the same priority are put into the same group and they can be performed in parallel manner, while different groups of sub-tasks with different priorities must be implemented according to the order of their priorities. An experiment has been done with the algorithms, the results of which demonstrate the advantage of the algorithms. Bo Liu 0004, Junzhou Luo, Wei Li 0017 |
AINA | 1 |
| 2005 | Distributed network self-management model based on CSCWabstractComputer network system has become a large-scale distributed system. But various kinds of existing network management solutions cannot manage it efficiently and network management is facing the new challenges. Applying the principles of CSCW to network management by multi-agent system is a novel train of thought of constructing the new generation network management system. Distributed network self-management model (DNSM), which is based on multi-agent system, adopts the management policies based on management domain and provides users with Web-based management way, was put forward in this paper. This model not only provides network management with more intelligence but also avoids the usage of amounts of network bandwidth. The experimental results show that DNSM is better than the existing network management solutions on performance for the large-scale computer network. Junzhou Luo, Wei Li 0017, Bo Liu 0004 |
CSCWD (1) | 3 |