EDBT 2026 Demo / reviewers in the wild / expert
Xingyi Song
dblp:185/5566
· DBLP profile ↗
18ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0002-4188-6974ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination MitigationabstractAs corporate responsibility increasingly incorporates environmental, social, and governance (ESG) criteria, ESG reporting is becoming a legal requirement in many regions and a key channel for documenting sustainability practices and assessing firms’ long-term and ethical performance. However, the length and complexity of ESG disclosures make them difficult to interpret and automate the analysis reliably. To support scalable and trustworthy analysis, this paper introduces ESG-Bench, a benchmark dataset for ESG report understanding and hallucination mitigation in large language models (LLMs). ESG-Bench contains human-annotated question–answer (QA) pairs grounded in real-world ESG report contexts, with fine-grained labels indicating whether model outputs are factually supported or hallucinated. Framing ESG report analysis as a QA task with verifiability constraints enables systematic evaluation of LLMs’ ability to extract and reason over ESG content and provides a new use case: mitigating hallucinations in socially sensitive, compliance-critical settings. We design task-specific Chain-of-Thought (CoT) prompting strategies and fine-tune multiple state-of-the-art LLMs on ESG-Bench using CoT-annotated rationales. Our experiments show that these CoT-based methods substantially outperform standard prompting and direct fine-tuning in reducing hallucinations, and that the gains transfer to existing QA benchmarks beyond the ESG domain. Ben Wu 0001, Mali Jin, Peizhen Bai, Hanpei Zhang, Xingyi Song |
AAAI | 6 |
| 2025 | A Dataset for Analysing News Framing in Chinese MediaabstractFraming is an essential device in news reporting, allowing writers to influence public perceptions of current affairs. While automatic news framing detection datasets exist in various languages, none focus on news framing in the Chinese language, which presents unique challenges with complex character meanings and unique linguistic features. This study introduces the first Chinese News Framing dataset, to be used as either a stand-alone dataset or a supplementary resource to the SemEval-2023 task 3 dataset. We detail its creation and conduct baseline experiments to demonstrate the need for such a dataset and create benchmarks for future research, providing results obtained through fine-tuning XLM-RoBERTa-Base and using GPT-4o in the zero-shot setting. We find that GPT-4o performs significantly worse than fine-tuned XLM-RoBERTa across all languages. For the Chinese language, we obtain an F1-micro (the performance metric for SemEval task 3, subtask 2) score of 0.719 using only samples from our Chinese News Framing dataset and a score of 0.753 when we augment the SemEval dataset with Chinese news framing samples. With positive news frame detection results, this dataset is a valuable resource for detecting news frames in the Chinese language and is a useful supplement to the SemEval-2023 task 3 dataset. Owen Cook, Yida Mu, Xingyi Song, Kalina Bontcheva |
ICWSM | 4 |
| 2025 | Cross-modal augmentation for few-shot multimodal fake news detection
Ye Jiang 0001, Taihang Wang, Xiaoman Xu, Yimin Wang 0002, Xingyi Song, Diana Maynard |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Identifying and Aligning Medical Claims Made on Social Media with Medical EvidenceabstractEvidence-based medicine is the practise of making medical decisions that adhere to the latest, and best known evidence at that time. Currently, the best evidence is often found in the form of documents, such as randomized control trials, meta-analyses and systematic reviews. This research focuses on aligning medical claims made on social media platforms with this medical evidence. By doing so, individuals without medical expertise can more effectively assess the veracity of such medical claims. We study three core tasks: identifying medical claims, extracting medical vocabulary from these claims, and retrieving evidence relevant to those identified medical claims. We propose a novel system that can generate synthetic medical claims to aid each of these core tasks. We additionally introduce a novel dataset produced by our synthetic generator that, when applied to these tasks, demonstrates not only a more flexible and holistic approach, but also an improvement in all comparable metrics. We make our dataset, the Expansive Medical Claim Corpus (EMCC), available at https://zenodo.org/records/8321460. Anthony James Hughes, Xingyi Song |
LREC/COLING | 2 |
| 2024 | Large Language Models Offer an Alternative to the Traditional Approach of Topic ModellingabstractTopic modelling, as a well-established unsupervised technique, has found extensive use in automatically detecting significant topics within a corpus of documents. However, classic topic modelling approaches (e.g., LDA) have certain drawbacks, such as the lack of semantic understanding and the presence of overlapping topics. In this work, we investigate the untapped potential of large language models (LLMs) as an alternative for uncovering the underlying topics within extensive text corpora. To this end, we introduce a framework that prompts LLMs to generate topics from a given set of documents and establish evaluation protocols to assess the clustering efficacy of LLMs. Our findings indicate that LLMs with appropriate prompts can stand out as a viable alternative, capable of generating relevant topic titles and adhering to human guidelines to refine and merge topics. Through in-depth experiments and evaluation, we summarise the advantages and constraints of employing LLMs in topic extraction. Yida Mu, Chun Dong, Kalina Bontcheva, Xingyi Song |
LREC/COLING | 4 |
| 2024 | Examining Temporalities on Stance Detection towards COVID-19 VaccinationabstractPrevious studies have highlighted the importance of vaccination as an effective strategy to control the transmission of the COVID-19 virus. It is crucial for policymakers to have a comprehensive understanding of the public’s stance towards vaccination on a large scale. However, attitudes towards COVID-19 vaccination, such as pro-vaccine or vaccine hesitancy, have evolved over time on social media. Thus, it is necessary to account for possible temporal shifts when analysing these stances. This study aims to examine the impact of temporal concept drift on stance detection towards COVID-19 vaccination on Twitter. To this end, we evaluate a range of transformer-based models using chronological (splitting the training, validation, and test sets in order of time) and random splits (randomly splitting these three sets) of social media data. Our findings reveal significant discrepancies in model performance between random and chronological splits in several existing COVID-19-related datasets; specifically, chronological splits significantly reduce the accuracy of stance classification. Therefore, real-world stance detection approaches need to be further refined to incorporate temporal factors as a key consideration. Yida Mu, Mali Jin, Kalina Bontcheva, Xingyi Song |
LREC/COLING | 4 |
| 2024 | Examining the Limitations of Computational Rumor Detection Models Trained on Static DatasetsabstractA crucial aspect of a rumor detection model is its ability to generalize, particularly its ability to detect emerging, previously unknown rumors. Past research has indicated that content-based (i.e., using solely source post as input) rumor detection models tend to perform less effectively on unseen rumors. At the same time, the potential of context-based models remains largely untapped. The main contribution of this paper is in the in-depth evaluation of the performance gap between content and context-based models specifically on detecting new, unseen rumors. Our empirical findings demonstrate that context-based models are still overly dependent on the information derived from the rumors’ source post and tend to overlook the significant role that contextual information can play. We also study the effect of data split strategies on classifier performance. Based on our experimental results, the paper also offers practical suggestions on how to minimize the effects of temporal concept drift in static datasets during the training of rumor detection methods. Yida Mu, Xingyi Song, Kalina Bontcheva, Nikolaos Aletras |
LREC/COLING | 2 |
| 2024 | Navigating Prompt Complexity for Zero-Shot Classification: A Study of Large Language Models in Computational Social ScienceabstractInstruction-tuned Large Language Models (LLMs) have exhibited impressive language understanding and the capacity to generate responses that follow specific prompts. However, due to the computational demands associated with training these models, their applications often adopt a zero-shot setting. In this paper, we evaluate the zero-shot performance of two publicly accessible LLMs, ChatGPT and OpenAssistant, in the context of six Computational Social Science classification tasks, while also investigating the effects of various prompting strategies. Our experiments investigate the impact of prompt complexity, including the effect of incorporating label definitions into the prompt; use of synonyms for label names; and the influence of integrating past memories during foundation model training. The findings indicate that in a zero-shot setting, current LLMs are unable to match the performance of smaller, fine-tuned baseline transformer models (such as BERT-large). Additionally, we find that different prompting strategies can significantly affect classification accuracy, with variations in accuracy and F1 scores exceeding 10%. Yida Mu, Ben Wu 0001, William Thorne, Ambrose Robinson, Nikolaos Aletras, Carolina Scarton, Kalina Bontcheva, Xingyi Song |
LREC/COLING | 8 |
| 2024 | The CLEF-2024 CheckThat! Lab: Check-Worthiness, Subjectivity, Persuasion, Roles, Authorities, and Adversarial Robustness
Alberto Barrón-Cedeño, Firoj Alam, Tanmoy Chakraborty 0002, Tamer Elsayed, Preslav Nakov, Piotr Przybyla, Julia Maria Struß, Fatima Haouari, Maram Hasanain, Federico Ruggeri, Xingyi Song, Reem Suwaileh |
ECIR (5) | 11 |
| 2024 | Enhancing Data Quality through Simple De-duplication: Navigating Responsible Computational Social Science ResearchabstractResearch in natural language processing (NLP) for Computational Social Science (CSS) heavily relies on data from social media platforms.This data plays a crucial role in the development of models for analysing socio-linguistic phenomena within online communities.In this work, we conduct an in-depth examination of 20 datasets extensively used in NLP for CSS to comprehensively examine data quality.Our analysis reveals that social media datasets exhibit varying levels of data duplication.Consequently, this gives rise to challenges like label inconsistencies and data leakage, compromising the reliability of models.Our findings also suggest that data duplication has an impact on the current claims of state-of-the-art performance, potentially leading to an overestimation of model effectiveness in real-world scenarios.Finally, we propose new protocols and best practices for improving dataset development from social media data and its usage. Yida Mu, Mali Jin, Xingyi Song, Nikolaos Aletras |
EMNLP | 3 |
| 2024 | Confidence Regulation Neurons in Language ModelsabstractDespite their widespread use, the mechanisms by which large language models (LLMs) represent and regulate uncertainty in next-token predictions remain largely unexplored. This study investigates two critical components believed to influence this uncertainty: the recently discovered entropy neurons and a new set of components that we term token frequency neurons. Entropy neurons are characterized by an unusually high weight norm and influence the final layer normalization (LayerNorm) scale to effectively scale down the logits. Our work shows that entropy neurons operate by writing onto an \textit{unembedding null space}, allowing them to impact the residual stream norm with minimal direct effect on the logits themselves. We observe the presence of entropy neurons across a range of models, up to 7 billion parameters. On the other hand, token frequency neurons, which we discover and describe here for the first time, boost or suppress each token’s logit proportionally to its log frequency, thereby shifting the output distribution towards or away from the unigram distribution. Finally, we present a detailed case study where entropy neurons actively manage confidence: the setting of induction, i.e. detecting and continuing repeated subsequences. Alessandro Stolfo, Ben Wu 0001, Wes Gurnee, Yonatan Belinkov, Xingyi Song, Mrinmaya Sachan, Neel Nanda |
NeurIPS | 5 |
| 2023 | VaxxHesitancy: A Dataset for Studying Hesitancy towards COVID-19 Vaccination on TwitterabstractVaccine hesitancy has been a common concern, probably since vaccines were created and, with the popularisation of social media, people started to express their concerns about vaccines online alongside those posting pro- and anti-vaccine content. Predictably, since the first mentions of a COVID-19 vaccine, social media users posted about their fears and concerns or about their support and belief into the effectiveness of these rapidly developing vaccines. Identifying and understanding the reasons behind public hesitancy towards COVID-19 vaccines is important for policy markers that need to develop actions to better inform the population with the aim of increasing vaccine take-up. In the case of COVID-19, where the fast development of the vaccines was mirrored closely by growth in anti-vaxx disinformation, automatic means of detecting citizen attitudes towards vaccination became necessary. This is an important computational social sciences task that requires data analysis in order to gain in-depth understanding of the phenomena at hand. Annotated data is also necessary for training data-driven models for more nuanced analysis of attitudes towards vaccination. To this end, we created a new collection of over 3,101 tweets annotated with users' attitudes towards COVID-19 vaccination (stance). Besides, we also develop a domain-specific language model (VaxxBERT) that achieves the best predictive performance (73.0 accuracy and 69.3 F1-score) as compared to a robust set of baselines. To the best of our knowledge, these are the first dataset and model that model vaccine hesitancy as a category distinct from pro- and anti-vaccine stance. Yida Mu, Mali Jin, Charlie Grimshaw, Carolina Scarton, Kalina Bontcheva, Xingyi Song |
ICWSM | 6 |
| 2023 | Similarity-Aware Multimodal Prompt Learning for fake news detection
Ye Jiang 0001, Xiaomin Yu, Yimin Wang 0002, Xiaoman Xu, Xingyi Song, Diana Maynard |
Inf. Sci. | 5 |
| 2020 | Comparing Topic-Aware Neural Networks for Bias Detection of NewsabstractThe commercial pressure on media has increasingly dominated the institutional rules of news media, and consequently, more and more sensational and dramatized frames and biases are in evidence in newspaper articles. Increased bias in the news media, which can result in misunderstanding and misuse of facts, leads to polarized opinions which can heavily influence the perspectives of the reader. This paper investigates learning models for detecting bias in the news. First, we look at incorporating into the models Latent Dirichlet Allocation (LDA) distributions which could enrich the feature space by adding word co-occurrence distribution and local topic probability in each document. In our proposed models, the LDA distributions are regarded as additive features on the sentence level and document level respectively. Second, we compare the performance of different popular neural network architectures incorporating these LDA distributions on a hyperpartisan newspaper article detection task. Preliminary experiment results show that the hierarchical models benefit more than non-hierarchical models when incorporating LDA features, and the former also outperform the latter. Ye Jiang 0001, Yimin Wang 0002, Xingyi Song, Diana Maynard |
ECAI | 3 |
| 2020 | RP-DNN: A Tweet Level Propagation Context Based Deep Neural Networks for Early Rumor Detection in Social MediaabstractEarly rumor detection (ERD) on social media platform is very challenging when limited, incomplete and noisy information is available. Most of the existing methods have largely worked on event-level detection that requires the collection of posts relevant to a specific event and relied only on user-generated content. They are not appropriate to detect rumor sources in the very early stages, before an event unfolds and becomes widespread. In this paper, we address the task of ERD at the message level. We present a novel hybrid neural network architecture, which combines a task-specific character-based bidirectional language model and stacked Long Short-Term Memory (LSTM) networks to represent textual contents and social-temporal contexts of input source tweets, for modelling propagation patterns of rumors in the early stages of their development. We apply multi-layered attention models to jointly learn attentive context embeddings over multiple context inputs. Our experiments employ a stringent leave-one-out cross-validation (LOO-CV) evaluation setup on seven publicly available real-life rumor event data sets. Our models achieve state-of-the-art(SoA) performance for detecting unseen rumors on large augmented data which covers more than 12 events and 2,967 rumors. An ablation study is conducted to understand the relative contribution of each component of our proposed model. Jie Gao 0009, Sooji Han, Xingyi Song, Fabio Ciravegna |
LREC | 3 |
| 2020 | Using Deep Neural Networks with Intra- and Inter-Sentence Context to Classify Suicidal BehaviourabstractIdentifying statements related to suicidal behaviour in psychiatric electronic health records (EHRs) is an important step when modeling that behaviour, and when assessing suicide risk. We apply a deep neural network based classification model with a lightweight context encoder, to classify sentence level suicidal behaviour in EHRs. We show that incorporating information from sentences to left and right of the target sentence significantly improves classification accuracy. Our approach achieved the best performance when classifying suicidal behaviour in Autism Spectrum Disorder patient records. The results could have implications for suicidality research and clinical surveillance. Xingyi Song, Johnny Downs, Sumithra Velupillai, Rachel Holden, Maxim Kikoler, Kalina Bontcheva, Rina Dutta, Angus Roberts |
LREC | 1 |
| 2018 | A Deep Neural Network Sentence Level Classification Method with Context InformationabstractIn the sentence classification task, context formed from sentences adjacent to the sentence being classified can provide important information for classification.This context is, however, often ignored.Where methods do make use of context, only small amounts are considered, making it difficult to scale.We present a new method for sentence classification, Context-LSTM-CNN, that makes use of potentially large contexts.The method also utilizes long-range dependencies within the sentence being classified, using an LSTM, and short-span features, using a stacked CNN.Our experiments demonstrate that this approach consistently improves over previous methods on two different datasets. Xingyi Song, Johann Petrak, Angus Roberts |
EMNLP | 1 |
| 2014 | Data selection for discriminative training in statistical machine translation
Xingyi Song, Lucia Specia, Trevor Cohn |
EAMT | 1 |