Biwei Cao

dblp:204/0889 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Multi-hop commonsense knowledge injection framework for zero-shot commonsense question answering
Jiuxin Cao, Biwei Cao, Qingqing Gao, Bo Liu 0004
Expert Syst. Appl.3
2026 PACEP: Steering large language models toward coherent user profile tracking and strategy selection in emotional support conversations
Hanyu Luo, Jiuxin Cao, Chang Liu 0113, Xinglin Li, Bo Liu 0004, Biwei Cao
Expert Syst. Appl.6
2025 External Reliable Information-enhanced Multimodal Contrastive Learning for Fake News Detection
abstract
With the rapid development of the Internet, the information dissemination paradigm has changed and the efficiency has been improved greatly. While this also brings the quick spread of fake news and leads to negative impacts on cyberspace. Currently, the information presentation formats have evolved gradually, with the news formats shifting from texts to multimodal contents. As a result, detecting multimodal fake news has become one of the research hotspots. However, multimodal fake news detection research field still faces two main challenges: the inability to fully and effectively utilize multimodal information for detection, and the low credibility or static nature of the introduced external information, which limits dynamic updates. To bridge the gaps, we propose ERIC-FND, an external reliable information-enhanced multimodal contrastive learning framework for fake news detection. ERIC-FND strengthens the representation of news contents by entity-enriched external information enhancement method. It also enriches the multimodal news information via multimodal semantic interaction method where the multimodal constrative learning is employed to make different modality representations learn from each other. Moreover, an adaptive fusion method is taken to integrate the news representations from different dimensions for the eventual classification. Experiments are done on two commonly used datasets in different languages, X (Twitter) and Weibo. Experiment results demonstrate that our proposed model ERIC-FND outperforms existing state-of-the-art fake news detection methods under the same settings.
Biwei Cao, Qihang Wu, Jiuxin Cao, Bo Liu 0004, Jie Gui
AAAI1
2025 Positive Text Reframing under Multi-strategy Optimization
abstract
Differing from sentiment transfer, positive reframing seeks to substitute negative perspectives with positive expressions while preserving the original meaning. With the emergence of pre-trained language models (PLMs), it is possible to achieve acceptable results by fine-tuning PLMs. Nevertheless, generating fluent, diverse and task-constrained reframing text remains a significant challenge. To tackle this issue, a multi-strategy optimization framework (MSOF) is proposed in this paper. Starting from the objective of positive reframing, we first design positive sentiment reward and content preservation reward to encourage the model to transform the negative expressions of the original text while ensuring the integrity and consistency of the semantics. Then, different decoding optimization approaches are introduced to improve the quality of text generation. Finally, based on the modeling formula of positive reframing, we propose a multi-dimensional re-ranking method that further selects candidate sentences from three dimensions: strategy consistency, text similarity and fluency. Extensive experiments on two Seq2Seq PLMs, BART and T5, demonstrate our framework achieves significant improvements on unconstrained and controlled positive reframing tasks.
Shutong Jia, Biwei Cao, Qingqing Gao, Jiuxin Cao, Bo Liu 0004
COLING2
2025 A Rationale-Guided Multi-Task Learning Framework for Hate Speech Detection
abstract
Hate Speech Detection (HSD) plays a crucial role in ensuring respectful online communication and preventing the spread of harmful content. However, existing studies on HSD often focus on the overall intent of the target message while overlooking the fine-grained details in the message. We propose a RatIonale-guided multi-taSk lEarning framework (RISE), which frames HSD as the main task and Human Rationales Tagging (HRT) as the auxiliary task. This enables simultaneous message-level understanding and the identification of expressions that trigger human judgments of hate speech, mimicking human cognitive processing to capture fine-grained information and enhance HSD performance. Furthermore, we extend rationale annotations from binary labels to BIO tagging, capturing the positional roles of tokens within rationales. Additionally, we integrate an emoji semantics interpretation module that interprets emoji meanings, enriching contextual information. Extensive experiments demonstrate that RISE outperforms state-of-the-art models, with each component contributing significantly to improved performance.
Qingqing Gao, Jiuxin Cao, Fengshan Song, Biwei Cao, Bo Liu 0004
IJCNN4
2025 Gen4Track: A Tuning-free Data Augmentation Framework via Self-correcting Diffusion Model for Vision-Language Tracking
abstract
The performance of current Vision-Language Tracking (VLT) models is constrained by the limited diversity and quantity of labeled data. Compared to constructing large-scale datasets, data augmentation offers a more cost-saving strategy for VLT by synthesizing new samples from existing data, rather than generating them from scratch. However, conventional techniques like rotation and flipping may disrupt scene composition, causing conflicts between visual layouts and textual annotations. Recent advances in generative models have inspired the use of synthetic videos for data augmentation. Yet, existing approaches fail to address the core concerns of data augmentation in VLT (shown in Fig. 1)-target location accuracy, text-video consistency, and video content coherency. To bridge the gap, we propose Gen4Track, a tuning-free data augmentation framework that leverages the self-correcting mechanism to dynamically generate high-quality video data with annotations. Our approach involves (1) optimizing the attention calculations in a frozen text-to-image diffusion model to synthesize coherent videos that satisfy specific conditions (e.g., spatial location, category, color, and style), and (2) implementing a self-correcting mechanism based on a Large Language Model (LLM) to improve text-video consistency. During video augmentation, we propose content-coherent self-attention and location-enhanced cross-attention mechanisms, ensuring that image-level editings are accurately and coherently propagated throughout the video. Then, with the goal of maximizing text-video consistency, we iteratively refine the augmentation instruction with our designed self-correcting mechanism for a more aligned video. Extensive experiments validate that Gen4Track significantly boosts the performance of SOTA VLT models (achieving improvements of up to 3.2% in SUC and 3.5% in PRE), opening a new chapter of training Vision-Language trackers with synthetic videos rather than manually annotated data.
Jiawei Ge 0002, Xin-Yu Zhang 0027, Jiuxin Cao, Xuelin Zhu, Qingqing Gao, Biwei Cao, Kun Wang 0057, Chang Liu 0113, Bo Liu 0004, Chen Feng 0028, Ioannis Patras
ACM Multimedia7
2025 MIAN: Multi-head Incongruity Aware Attention Network with transfer learning for sarcasm detection
Jiuxin Cao, Biwei Cao, Bo Liu 0004
Expert Syst. Appl.4
2025 Text-to-face synthesis based on facial landmarks prediction
Kun Wang 0057, Biwei Cao, Bo Liu 0004, Jiuxin Cao
Mach. Vis. Appl.3
2025 No Place to Hide: Dual Deep Interaction Channel Network for Fake News Detection With Data Augmentation
abstract
Online social network has emerged as a prominent place for the propagation of fake news due to its low cost of information dissemination. Although the existing methods have made many attempts in news content and propagation structure, the detection of fake news is still facing two challenges: one is how to mine the unique key features and evolution patterns, and the other is how to tackle the problem of small samples to build the high-performance model. Different from popular methods, which take full advantage of the propagation topology structure, in this article, we propose a novel framework for fake news detection from perspectives of semantics, emotion and data enhancement. The semantic and emotional features of news and comments, the inconsistent emotion between news and news participants as well as the emotion evolution features in comments are fused by the designed dual deep interaction channel network to obtain a more comprehensive and fine-grained news representation. Meanwhile, with the construction of large language model (LLM) prompt, a LLM-based data enhancement module is used to obtain more diverse labeled data of high quality filtered by confidence, further improving the performance of the classification model. Experiments show that the proposed approach outperforms the state-of-the-art methods.
Biwei Cao, Jiuxin Cao, Lulu Hua, Bo Liu 0004, Jie Gui, James T. Kwok
IEEE Trans. Comput. Soc. Syst.1
2024 CEPT: A Contrast-Enhanced Prompt-Tuning Framework for Emotion Recognition in Conversation
abstract
Emotion Recognition in Conversation (ERC) has attracted increasing attention due to its wide applications in public opinion analysis, empathetic conversation generation, and so on. However, ERC research suffers from the problems of data imbalance and the presence of similar linguistic expressions for different emotions. These issues can result in limited learning for minority emotions, biased predictions for common emotions, and the misclassification of different emotions with similar linguistic expressions. To alleviate these problems, we propose a Contrast-Enhanced Prompt-Tuning (CEPT) framework for ERC. We transform the ERC task into a Masked Language Modeling (MLM) generation task and generate the emotion for each utterance in the conversation based on the prompt-tuning of the Pre-trained Language Model (PLM), where a novel mixed prompt template and a label mapping strategy are introduced for better context and emotion feature modeling. Moreover, Supervised Contrastive Learning (SCL) is employed to help the PLM mine more information from the labels and learn a more discriminative representation space for utterances with different emotions. We conduct extensive experiments and the results demonstrate that CEPT outperforms the state-of-the-art methods on all three benchmark datasets and excels in recognizing minority emotions.
Qingqing Gao, Jiuxin Cao, Biwei Cao, Bo Liu 0004
LREC/COLING3
2024 EnsCLR: Unsupervised skeleton-based action recognition via ensemble contrastive learning of representation
Kun Wang 0057, Jiuxin Cao, Biwei Cao, Bo Liu 0004
Comput. Vis. Image Underst.3
2024 A data-driven approach to choosing privacy parameters for clinical trial data sharing under differential privacy
abstract
OBJECTIVES: Clinical trial data sharing is crucial for promoting transparency and collaborative efforts in medical research. Differential privacy (DP) is a formal statistical technique for anonymizing shared data that balances privacy of individual records and accuracy of replicated results through a "privacy budget" parameter, ε. DP is considered the state of the art in privacy-protected data publication and is underutilized in clinical trial data sharing. This study is focused on identifying ε values for the sharing of clinical trial data. MATERIALS AND METHODS: We analyzed 2 clinical trial datasets with privacy budget ε ranging from 0.01 to 10. Smaller values of ε entail adding greater amounts of random noise, with better privacy as a result. Comparison of rates, odds ratios, means, and mean differences between the original clinical trial datasets and the empirical distribution of the DP estimator was performed. RESULTS: The DP rate closely approximated the original rate of 6.5% when ε > 1. The DP odds ratio closely aligned with the original odds ratio of 0.689 when ε ≥ 3. The DP mean closely approximated the original mean of 164.64 when ε ≥ 1. As ε increased to 5, both the minimum and maximum DP means converged toward the original mean. DISCUSSION: There is no consensus on how to choose the privacy budget ε. The definition of DP does not specify the required level of privacy, and there is no established formula for determining ε. CONCLUSION: Our findings suggest that the application of DP holds promise in the context of sharing clinical trial data.
Henian Chen, Jinyong Pang, Yayi Zhao, Spencer Giddens, Joseph Ficek, Matthew J. Valente, Biwei Cao, Ellen Daley
J. Am. Medical Informatics Assoc.7
2024 GAL: combining global and local contexts for interpersonal relation extraction toward document-level Chinese text
Jiawei Ge 0002, Jiuxin Cao, Yingxing Bao, Biwei Cao, Bo Liu 0004
Neural Comput. Appl.4
2024 Correction: GAL: combining global and local contexts for interpersonal relation extraction toward document-level Chinese text
Jiawei Ge 0002, Jiuxin Cao, Yingxing Bao, Biwei Cao, Bo Liu 0004
Neural Comput. Appl.4
2024 Response Generation in Social Network With Topic and Emotion Constraints
abstract
Response generation is the task of automatically generating human-like content based on the provided context. One of its prominent applications is to simulate realistic response content for social network posts. In the digital age, social network platforms play a vital role in information exchange and social interaction. This study focuses on response generation techniques for the platform of public opinion evolution simulation that simulate realistic response content, enabling a deeper understanding of the emotional expressions of network users. Recent advancements in deep learning techniques, particularly the sequence-to-sequence (Seq2Seq) model, have shown promise in the response generation field. However, we still face two challenges: content variety, topic and emotion relevancy. To this end, we propose the EmoTG-ETRS model which comprises three parts. The first is a response generation module based on Transformer architecture. Then, an auxiliary emotion improvement module is incorporated to enhance the emotional expressiveness of the response candidates. Finally, a reverse selection module, which combines maximum mutual information (MMI) evaluation, emotional expression evaluation, and topic consistency evaluation, is devised to select the highest-scoring response. Extensive experiments have been conducted to evaluate the effectiveness of the proposed model and the results demonstrate that the EmoTG-ETRS model improves the quality of produced replies in terms of topic consistency and emotional accuracy rate when compared with the SOTA research works.
Biwei Cao, Jiuxin Cao, Bo Liu 0004, Jie Gui, Jun Zhou 0027, Yuan Yan Tang, James T. Kwok
IEEE Trans. Comput. Soc. Syst.1
2023 AlignVE: Visual Entailment Recognition Based on Alignment Relations
abstract
Visual entailment (VE) is to recognize whether the semantics of a hypothesis text can be inferred from the given premise image, which is one special task among recent emerged vision and language understanding tasks. Currently, most of the existing VE approaches are derived from the methods of visual question answering. They recognize visual entailment by quantifying the similarity between the hypothesis and premise in the content semantic features from multi modalities. Such approaches, however, ignore the VE's unique nature of relation inference between the premise and hypothesis. Therefore, in this paper, a new architecture called AlignVE is proposed to solve the visual entailment problem with a relation interaction method. It models the relation between the premise and hypothesis as an alignment matrix. Then it introduces a pooling operation to get feature vectors with a fixed size. Finally, it goes through the fully-connected layer and normalization layer to complete the classification. Experiments show that our alignment-based architecture reaches 72.45% accuracy on SNLI-VE dataset, outperforming previous content-based models under the same settings.
Biwei Cao, Jiuxin Cao, Jie Gui, Jiayun Shen, Bo Liu 0004, Yuan Yan Tang, James T. Kwok
IEEE Trans. Multim.1
2022 CORN: Co-Reasoning Network for Commonsense Question Answering
abstract
Commonsense question answering (QA) requires machines to utilize the QA content and external commonsense knowledge graph (KG) for reasoning when answering questions. Existing work uses two independent modules to model the QA contextual text representation and relationships between QA entities in KG, which prevents information sharing between modules for co-reasoning. In this paper, we propose a novel model, Co-Reasoning Network (CORN), which adopts a bidirectional multi-level connection structure based on Co-Attention Transformer. The structure builds bridges to connect each layer of the text encoder and graph encoder, which can introduce the QA entity relationship from KG to the text encoder and bring contextual text information to the graph encoder, so that these features can be deeply interactively fused to form comprehensive text and graph node representations. Meanwhile, we propose a QA-aware node based KG subgraph construction method. The QA-aware nodes aggregate the question entity nodes and the answer entity nodes, and further guide the expansion and construction process of the subgraph to enhance the connectivity and reduce the introduction of noise. We evaluate our model on QA benchmarks in the CommonsenseQA and OpenBookQA datasets, and CORN achieves state-of-the-art performance.
Biwei Cao, Qingqing Gao, Zheng Yin, Bo Liu 0004, Jiuxin Cao
COLING2
2022 Efficient gradient boosting for prognostic biomarker discovery
abstract
MOTIVATION: A gradient boosting decision tree (GBDT) is a powerful ensemble machine-learning method that has the potential to accelerate biomarker discovery from high-dimensional molecular data. Recent algorithmic advances, such as extreme gradient boosting (XGB) and light gradient boosting (LGB), have rendered the GBDT training more efficient, scalable and accurate. However, these modern techniques have not yet been widely adopted in discovering biomarkers for censored survival outcomes, which are key clinical outcomes or endpoints in cancer studies. RESULTS: In this paper, we present a new R package 'Xsurv' as an integrated solution that applies two modern GBDT training frameworks namely, XGB and LGB, for the modeling of right-censored survival outcomes. Based on our simulations, we benchmark the new approaches against traditional methods including the stepwise Cox regression model and the original gradient boosting function implemented in the package 'gbm'. We also demonstrate the application of Xsurv in analyzing a melanoma methylation dataset. Together, these results suggest that Xsurv is a useful and computationally viable tool for screening a large number of prognostic candidate biomarkers, which may facilitate future translational and clinical research. AVAILABILITY AND IMPLEMENTATION: 'Xsurv' is freely available as an R package at: https://github.com/topycyao/Xsurv. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kaiqiao Li, Sijie Yao, Biwei Cao, Denise Kalos, Pei Fen Kuan, Ruoqing Zhu
Bioinform.4
2022 Emotion recognition in conversations with emotion shift detection based on multi-task learning
Qingqing Gao, Biwei Cao, Tianyun Gu, Xing Bao, Junyan Wu, Bo Liu 0004, Jiuxin Cao
Knowl. Based Syst.2
2020 PDXGEM: patient-derived tumor xenograft-based gene expression model for predicting clinical response to anticancer therapy in cancer patients
abstract
BACKGROUND: Cancer is a highly heterogeneous disease with varying responses to anti-cancer drugs. Although several attempts have been made to predict the anti-cancer therapeutic responses, there remains a great need to develop highly accurate prediction models of response to the anti-cancer drugs for clinical applications toward a personalized medicine. Patient derived xenografts (PDXs) are preclinical cancer models in which the tissue or cells from a patient's tumor are implanted into an immunodeficient or humanized mouse. In the present study, we develop a bioinformatics analysis pipeline to build a predictive gene expression model (GEM) for cancer patients' drug responses based on gene expression and drug activity data from PDX models. RESULTS: Drug sensitivity biomarkers were identified by performing an association analysis between gene expression levels and post-treatment tumor volume changes in PDX models. We built a drug response prediction model (called PDXGEM) in a random-forest algorithm by using a subset of the drug sensitvity biomarkers with concordant co-expression patterns between the PDXs and pretreatment cancer patient tumors. We applied the PDXGEM to several cytotoxic chemotherapies as well as targeted therapy agents that are used to treat breast cancer, pancreatic cancer, colorectal cancer, or non-small cell lung cancer. Significantly accurate predictions of PDXGEM for pathological response or survival outcomes were observed in extensive independent validations on multiple cancer patient datasets obtained from retrospective observational studies and prospective clinical trials. CONCLUSION: Our results demonstrated the strong potential of using molecular profiles and drug activity data of PDX tumors in developing a clinically translatable predictive cancer biomarkers for cancer patients. The PDXGEM web application is publicly available at http://pdxgem.moffitt.org .
Youngchul Kim, Biwei Cao, Rodrigo Carvajal
BMC Bioinform.3
2019 Joint Visual-Textual Sentiment Analysis Based on Cross-Modality Attention Mechanism
Xuelin Zhu, Biwei Cao, Bo Liu 0004, Jiuxin Cao
MMM (1)2