Haorui He

dblp:319/5437 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System
abstract
State-of-the-art (SOTA) fact-checking systems combat misinformation by employing autonomous LLM-based agents to decompose complex claims into smaller sub-claims, verify each sub-claim individually, and aggregate the partial results to produce verdicts with justifications (explanations for the verdicts). The security of these systems is crucial, as compromised fact-checkers can amplify misinformation, but remains largely underexplored. To bridge this gap, this work introduces a novel threat model against such fact-checking systems and presents Fact2Fiction, the first poisoning attack framework targeting SOTA agentic fact-checking systems. Fact2Fiction employs LLMs to mimic the decomposition strategy and exploit system-generated justifications to craft tailored malicious evidences that compromise sub-claim verification. Extensive experiments demonstrate that Fact2Fiction achieves 8.9%-21.2% higher attack success rates than SOTA attacks across various poisoning budgets and exposes security weaknesses in existing fact-checking systems, highlighting the need for defensive countermeasures.
Haorui He, Yupeng Li 0001, Bin B. Zhu, Dacheng Wen, Reynold Cheng, Francis C. M. Lau 0001
AAAI1
2026 Debating Truth: Debate-driven Claim Verification with Multiple Large Language Model Agents
abstract
State-of-the-art single-agent claim verification methods struggle with complex claims that require nuanced analysis of multifaceted evidence. Inspired by real-world professional fact-checkers, we propose DebateCV, the first debate-driven claim verification framework powered by multiple LLM agents. In DebateCV, two Debaters argue opposing stances to surface subtle errors in single-agent assessments. A decisive Moderator is then required to weigh the evidential strength of conflicting arguments to deliver an accurate verdict. Yet, zero-shot Moderators are biased toward neutral judgments, and no datasets exist for training them. To bridge this gap, we propose Debate-SFT, a post-training framework that leverages synthetic data to enhance agents' ability to effectively adjudicate debates for claim verification. Results show that our methods surpass state-of-the-art non-debate approaches in both accuracy (across various evidence conditions) and justification quality.
Haorui He, Yupeng Li 0001, Dacheng Wen, Yang Chen 0001, Reynold Cheng, Donald Donglong Chen, Francis C. M. Lau 0001
WWW1
2025 Deep learning based coronary vessels segmentation in X-ray angiography using temporal information
abstract
Invasive coronary angiography (ICA) is the gold standard imaging modality during cardiac interventions. Accurate segmentation of coronary vessels in ICA is required for aiding diagnosis and creating treatment plans. Current automated algorithms for vessel segmentation face task-specific challenges, including motion artifacts and unevenly distributed contrast, as well as the general challenge inherent to X-ray imaging, which is the presence of shadows from overlapping organs in the background. To address these issues, we present Temporal Vessel Segmentation Network (TVS-Net) model that fuses sequential ICA information into a novel densely connected 3D encoder-2D decoder structure with a loss function based on elastic interaction. We develop our model using an ICA dataset comprising 323 samples, split into 173 for training, 82 for validation, and 68 for testing, with a relatively relaxed annotation protocol that produced coarse-grained samples, and achieve 83.4% Dice and 84.3% recall on the test dataset. We additionally perform an external evaluation over 60 images from a local hospital, achieving 78.5% Dice and 82.4% recall and outperforming the state-of-the-art approaches. We also conduct a detailed manual re-segmentation for evaluation only on a subset of the first dataset under strict annotation protocol, achieving a Dice score of 86.2% and recall of 86.3% and surpassing even the coarse-grained gold standard used in training. The results indicate our TVS-Net is effective for multi-frame ICA segmentation, highlights the network's generalizability and robustness across diverse settings, and showcases the feasibility of weak supervision in ICA segmentation.
Haorui He, Abhirup Banerjee, Robin Choudhury, Vicente Grau
Medical Image Anal.1
2024 Emilia: An Extensive, Multilingual, and Diverse Speech Dataset For Large-Scale Speech Generation
abstract
Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and spontaneous speech datasets. In response, we introduce Emilia, the first large-scale, multilingual, and diverse speech generation dataset. Emilia starts with over 101k hours of speech across six languages, covering a wide range of speaking styles to enable more natural and spontaneous speech generation. To facilitate the scale-up of Emilia, we also present Emilia-Pipe, the first open-source preprocessing pipeline designed to efficiently transform raw, in-the-wild speech data into high-quality training data with speech annotations. Experimental results demonstrate the effectiveness of both Emilia and Emilia-Pipe. Demos are available at: https://emilia-dataset.github.io/Emilia-Demo-Page/.
Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu 0008, Jiaqi Li 0030, Peiyang Shi, Yuancheng Wang, Kai Chen 0026, Pengyuan Zhang, Zhizheng Wu 0001
SLT1
2024 SPMIS: An Investigation of Synthetic Spoken Misinformation Detection
abstract
In recent years, speech generation technology has advanced rapidly, fueled by generative models and large-scale training techniques. While these developments have enabled the production of high-quality synthetic speech, they have also raised concerns about the misuse of this technology, particularly for generating synthetic misinformation. Current research primarily focuses on distinguishing machine-generated speech from human-produced speech, but the more urgent challenge is detecting misinformation within spoken content. This task requires a thorough analysis of factors such as speaker identity, topic, and synthesis. To address this need, we conduct an initial investigation into synthetic spoken misinformation detection by introducing an open-source dataset, SpMis. SpMis includes speech synthesized from over 1,000 speakers across five common topics, utilizing state-of-the-art text-to-speech systems. Although our results show promising detection capabilities, they also reveal substantial challenges for practical implementation, underscoring the importance of ongoing research in this critical area.
Peizhuo Liu, Renqiang He, Haorui He, Huadi Zheng, Jie Shi 0005, Tong Xiao 0001, Zhizheng Wu 0001
SLT4
2024 Amphion: an Open-Source Audio, Music, and Speech Generation Toolkit
abstract
Amphion is an open-source toolkit for Audio, Music, and Speech Generation, targeting to ease the way for junior researchers and engineers into these fields. It presents a unified framework that includes diverse generation tasks and models, with the added bonus of being easily extendable for new incorporation. The toolkit is designed with beginner-friendly workflows and pre-trained models, allowing both beginners and seasoned researchers to kick-start their projects with relative ease. The initial release of Amphion v0.1 supports a range of tasks including Text to Speech (TTS), Text to Audio (TTA), and Singing Voice Conversion (SVC), supplemented by essential components like data preprocessing, state-of-the-art vocoders, and evaluation metrics. This paper presents a high-level overview of Amphion. Amphion is open-sourced at https://github.com/open-mmlab/Amphion.
Xueyao Zhang, Liumeng Xue, Yicheng Gu, Yuancheng Wang, Jiaqi Li 0030, Haorui He, Chaoren Wang, Songting Liu, Junan Zhang, Zihao Fang, Haopeng Chen, Tze Ying Tang, Lexiao Zou, Mingxuan Wang, Kai Chen 0026, Haizhou Li 0001, Zhizheng Wu 0001
SLT6
2024 MCFEND: A Multi-source Benchmark Dataset for Chinese Fake News Detection
abstract
The prevalence of fake news across various online sources has had a significant influence on the public. Existing Chinese fake news detection datasets are limited to news sourced solely from Weibo. However, fake news originating from multiple sources exhibits diversity in various aspects, including its content and social context. Methods trained on purely one single news source can hardly be applicable to real-world scenarios. Our pilot experiment demonstrates that the F1 score of the state-of-the-art method that learns from a large Chinese fake news detection dataset, Weibo-21, drops significantly from 0.943 to 0.470 when the test data is changed to multi-source news data, failing to identify more than one-third of the multi-source fake news. To address this limitation, we constructed the first multi-source benchmark dataset for Chinese fake news detection, termed MCFEND, which is composed of news we collected from diverse sources such as social platforms, messaging apps, and traditional online news outlets. Notably, such news has been fact-checked by 14 authoritative fact-checking agencies worldwide. In addition, various existing Chinese fake news detection methods are thoroughly evaluated on our proposed dataset in cross-source, multi-source, and unseen source ways. MCFEND, as a benchmark dataset, aims to advance Chinese fake news detection approaches in real-world scenarios.
Yupeng Li 0001, Haorui He, Dacheng Wen
WWW2
2023 Contextual Target-Specific Stance Detection on Twitter: Dataset and Method
abstract
To understand different aspects of online human behaviors, e.g., the public stances toward various social and political issues, contextual target-specific stance detection has become one of the most important studies on social media. Considering the lack of appropriate data for the studies of contextual target-specific stance detection on Twitter, which is one of the most popular online social platforms worldwide, we introduce CTSDT, a new dataset that consists of a large number of annotated target-specific conversations collected from Twitter. Furthermore, we propose a new contextual target-specific stance detection model called ConMulAttn, which is the first method that can learn both the contents of the posts and the concrete relationships between the posts in a conversation. We conduct extensive evaluation using CTSDT as well as another two popular datasets, CreateDebate and ConvinceMe, for contextual target-specific stance detection. The evaluation results validate the necessity of introducing our dataset CTSDT. Besides, according to the evaluation results, our proposed model ConMulAttn can outperform the state-of-the-art contextual target-specific stance detection method by up to 25% in F1score, indicating the effectiveness and superiority of our solution. Our study has the potential to assist policymakers in utilizing conversation data from online social platforms to efficiently gain real-time insights into public stances on target topics, such as vaccination.
Yupeng Li 0001, Dacheng Wen, Haorui He, Jianxiong Guo, Xuan Ning, Francis C. M. Lau 0001
ICDM3
2023 Improved Target-Specific Stance Detection on Social Media Platforms by Delving Into Conversation Threads
abstract
Target-specific stance detection on social media, which aims at classifying a textual data instance such as a post or a comment into a stance class of a target issue, is an emerging opinion mining paradigm of importance. An example application would be to overcome vaccine hesitancy in combating the coronavirus pandemic. Existing stance detection strategies rely merely on the individual instances which cannot always capture the expressed stance of a given target. We address a new task called conversational stance detection (CSD) which is to infer the stance toward a given target (e.g., COVID-19 vaccination) when given a data instance and its corresponding conversation thread. To carry out the task, we first propose a benchmarking CSD dataset with annotations of stances and the structures of conversation threads among the instances, which is based on six major social media platforms in Hong Kong. To infer the desired stances from both data instances and conversation threads, we propose a model called Branch-bidirectional encoder representations from transformers (BERT) that incorporates contextual information in conversation threads. Extensive experiments on our CSD dataset show that our proposed model outperforms all the baseline models that do not make use of contextual information. Specifically, it improves the F1 score by 10.3% compared with the state-of-the-art method in the SemEval-2016 Task 6 competition. This shows the potential of incorporating rich contextual information on detecting target-specific stances on social media platforms and suggests a more practical way to construct future stance detection tasks.
Yupeng Li 0001, Haorui He, Shaonan Wang, Francis C. M. Lau 0001, Yunya Song
IEEE Trans. Comput. Soc. Syst.2
2022 Adaptive Knowledge Distillation for Efficient Relation Classification
Haorui He, Yuanzhe Ren
ICANN (2)1