Boyang Liu 0002

dblp:165/8466-2 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0002-0814-002XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 ConspEmoLLM: Conspiracy Theory Detection Using an Emotion-Based Large Language Model
abstract
The internet has brought both benefits and harms to society. A prime example of the latter is misinformation, including conspiracy theories, which flood the web. Recent advances in natural language processing, particularly the emergence of large language models (LLMs), have improved the prospects of accurate misinformation detection. However, most LLM-based approaches to conspiracy theory detection focus only on binary classification and fail to account for the important relationship between misinformation and affective features (i.e., sentiment and emotions). Driven by a comprehensive analysis of conspiracy text that reveals its distinctive affective features, we propose ConspEmoLLM, the first open-source LLM that integrates affective information and is able to perform diverse tasks relating to conspiracy theories. These tasks include not only conspiracy theory detection, but also classification of theory type and detection of related discussion (e.g., opinions towards theories). ConspEmoLLM is fine-tuned based on an emotion-oriented LLM using our novel ConDID dataset, which includes five tasks to support LLM instruction tuning and evaluation. We demonstrate that when applied to these tasks, ConspEmoLLM largely outperforms several open-source general domain LLMs and ChatGPT, as well as an LLM that has been fine-tuned using ConDID, but which does not use affective features. ConspEmoLLM can be easily applied to identify and classify conspiracy-related text in the real world. The work has been released at https://github.com/lzw108/ConspEmoLLM/.
Zhiwei Liu 0003, Boyang Liu 0002, Paul Thompson 0002, Kailai Yang, Sophia Ananiadou
ECAI2
2024 SuicidEmoji: Derived Emoji Dataset and Tasks for Suicide-Related Social Content
abstract
Early suicidal ideation detection using social media is crucial for mental health surveillance. Simultaneously, emojis from the posts can help us better understand users' emotions and predict mental health conditions. However, research in emoji-based suicide analysis remains underexplored, with few resources available, which can restrict the development of studying emoji usage patterns among users with suicidal ideation. In this work, we build a derived suicide-related emoji dataset named SuicidEmoji, which contains 25k emoji posts (2,329 suicide-related posts and 22,722 posts for the control group users) filtered from about 1.3 million crawled Reddit data. To the best of our knowledge, SuicidEmoji is the first suicide-related emoji dataset. Based on SuicidEmoji, we propose two novel tasks: emoji-aware suicidal ideation detection and emoji prediction, for which we build two benchmark subdatasets from SuicidEmoji to evaluate the performance of advanced methods including pre-trained language models (PLMs) and large language models (LLMs). We analyze the experimental results of two PLMs and the highly capable LLMs, which reveal the significance and challenges of emoji-based suicide-related NLP tasks. The dataset is avaliable at https://github.com/TianlinZhang668/SuicidEmoji.
Kailai Yang, Shaoxiong Ji, Boyang Liu 0002, Qianqian Xie, Sophia Ananiadou
SIGIR4
2023 Global information-aware argument mining based on a top-down multi-turn QA model
abstract
Argument mining (AM) aims to automatically generate a graph that represents the argument structure of a document. Most previous AM models only pay attention to a single argument component (AC) to classify the type of the AC or a pair of ACs to identify and classify the argumentative relation (AR) between the two ACs. These models ignore the impact of global argument structure of the documents, which is important, especially in some highly structured genres such as scientific papers, where the process of argumentation is relatively fixed. Inspired by this, we propose a novel two-stage model which leverages global structure information to support AM. The first stage uses a multi-turn question-answering model to incrementally generate an initial argumentative graph that identifies relations among ACs. At each turn, all ACs related to the query AC are generated simultaneously, such that the sibling global information between the answer ACs is considered. In addition, the partially constructed graph is used as global structure information to support the extension of the graph with additional ACs. After the whole initial graph structure has been determined, the second stage assigns semantic types to both the ACs and ARs among them, leveraging information from this initial graph as global structure information. We test the proposed methods on two scientific datasets (one is the AbstRCT dataset including 659 abstracts about cancer research and the other is the SciARG dataset that consists of 225 computer linguistic abstracts and 285 biomedical abstracts) and a student essay dataset PE with 402 essays. Our experiments show that our model improves the state-of-the-art performance on two scientific datasets for different AM subtasks, with average improvements of 1%, 2.41%, 1.1% for the ACC, ARI and ARC task respectively on the AbstRCT dataset, and 2.36%, 1.84%, 8.87% for the ACC, ARI and ARC task on the SciARG dataset. Our model also achieves comparative results on the PE datasets: 87.7% of F1 scores for the ACC task, 81.4% for the ARI task and 78.8% for the ARC task.
Boyang Liu 0002, Viktor Schlegel, Paul Thompson 0002, Riza Theresa Batista-Navarro, Sophia Ananiadou
Inf. Process. Manag.1
2023 PHQ-aware depressive symptoms identification with similarity contrastive learning on social media
abstract
Depressive symptoms identification on social media aims to identify posts from social media expressing symptoms of depression. This can be beneficial for developing mental health support systems and for understanding the symptoms of depression. The Patient Health Questionnaire-9 (PHQ-9) is an instrument that healthcare professionals widely use to assess and monitor symptoms of depression. However, most existing models only consider capturing semantic information from posts, without considering PHQ-9 descriptive information related to symptoms. In addition, they are not devised to capture features that are specific to each symptom, especially in the case of multi-label symptoms identification. To tackle these challenges, we present a Span-based PHQ-aware and similarity contrastive network (SpanPHQ). We first adopt a novel span-based framework casting depressive symptoms identification task as a span-prediction problem. Then, we introduce context-aware and PHQ-aware self-guided cross-attention modules to enhance the model’s ability to consider both semantic contextual information and PHQ-9 descriptive information. Besides, a similarity contrastive learning is designed to effectively utilise the label information in identifying class-specific features. Our model is evaluated on two depressive symptoms identification datasets, i.e., the D2S dataset with 1,850 Twitter posts and the PRIMATE dataset with 2,000 Reddit posts. Moreover, our model achieves competitive performance compared to existing models on both datasets, with macro-F1 of 63.22%, 68.84%, micro-F1 of 73.34%, 75.92%, weighted-F1 of 72.86%, 76.65%, JacS of 69.94%, 63.82% and HamL of 0.0665, 0.1832 on these two datasets, respectively. The ablation study further provides evidence of the effectiveness of each module we proposed. Furthermore, we include visualisations and case studies verifying the ability of our model to learn PHQ information and its superior performance over existing baselines. Our work is expected to help the future identification and analysis of depressive symptoms on social media.
Kailai Yang, Hassan Alhuzali, Boyang Liu 0002, Sophia Ananiadou
Inf. Process. Manag.4
2022 Incorporating Zoning Information into Argument Mining from Biomedical Literature
abstract
The goal of text zoning is to segment a text into zones (i.e., Background, Conclusion) that serve distinct functions. Argumentative zoning, a specific text zoning scheme for the scientific domain, is considered as the antecedent for argument mining by many researchers. Surprisingly, however, little work is concerned with exploiting zoning information to improve the performance of argument mining models, despite the relatedness of the two tasks. In this paper, we propose two transformer-based models to incorporate zoning information into argumentative component identification and classification tasks. One model is for the sentence-level argument mining task and the other is for the token-level task. In particular, we add the zoning labels predicted by an off-the-shelf model to the beginning of each sentence, inspired by the convention commonly used biomedical abstracts. Moreover, we employ multi-head attention to transfer the sentence-level zoning information to each token in a sentence. Based on experiment results, we find a significant improvement in F1-scores for both sentence- and token-level tasks. It is worth mentioning that these zoning labels can be obtained with high accuracy by utilising readily available automated methods. Thus, existing argument mining models can be improved by incorporating zoning information without any additional annotation cost.
Boyang Liu 0002, Viktor Schlegel, Riza Theresa Batista-Navarro, Sophia Ananiadou
LREC1