Yichuan Li 0001

dblp:216/7478-1 · DBLP profile ↗
← Back
6ranked-venue papers in the field
3as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 4 (2 first)Data Mining & Knowledge Discovery · 2 (1 first)
YearPublicationVenuePosition
2025 LLM Profiling and Fine-Tuning with Limited Neighbor Information for Node Classification on Text-Attributed Graphs
Xinyi Fang, Kyumin Lee, Yichuan Li 0001
IEEE Big Data3
2023 What Boosts Fake News Dissemination on Social Media? A Causal Inference View
Yichuan Li 0001, Kyumin Lee, Nima Kordzadeh, Ruocheng Guo
PAKDD (4)1
2021 Multi-Source Domain Adaptation with Weak Supervision for Early Fake News Detection
abstract
Recently, the massive and diverse fake news from politics to entertainment and health has amplified the social distrust problem and has become a big challenge for the society and research community. The existing fake news detection methods are mostly designed for either a specific domain or require huge labeled data from various domains. If there is not enough labeled data in a certain domain, existing models may not work well for detecting fake news from that domain. To overcome these limitations we propose a novel framework based on multisource domain adaptation and weak supervision for early fake news detection. The framework transfers sufficient labeled source domains’ knowledge into a target/new domain with limited or even no labeled data by the multi-source domain adaptation, and applies researchers’ prior knowledge about fake news to the target domain by the weak supervision. The weak supervision assigns the weak labels to the unlabeled samples in the target domain through known heuristic rules. Our experimental results show that our approach outperforms 7 state-of-the-art methods in three real-world datasets. In particular, our model achieves, on average, 5.2% higher accuracy than the best baseline. Our model with a more advanced encoder can further boost the performance by 3.7%. The code is available at this clickable link.
Yichuan Li 0001, Kyumin Lee, Nima Kordzadeh, Brenton D. Faber, Cameron Fiddes, Elaine Chen, Kai Shu
IEEE BigData1
2021 Reducing and Exploiting Data Augmentation Noise through Meta Reweighting Contrastive Learning for Text Classification
abstract
Data augmentation has shown its effectiveness in resolving the data-hungry problem and improving model's generalization ability. However, the quality of augmented data can be varied, especially compared with the raw/original data. To boost deep learning models' performance given augmented data/samples in text classification tasks, we propose a novel framework, which leverages both meta learning and contrastive learning techniques as parts of our design for reweighting the augmented samples and refining their feature representations based on their quality. As part of the framework, we propose novel weight-dependent enqueue and dequeue algorithms to utilize augmented samples' weight/quality information effectively. Through experiments, we show that our framework can reasonably cooperate with existing deep learning models (e.g., RoBERTa-base and Text-CNN) and augmentation techniques (e.g., Wordnet and Easydata) for specific supervised learning tasks. Experiment results show that our framework achieves an average of 1.6%, up to 4.3% absolute improvement on Text-CNN encoders and an average of 1.4%, up to 4.4% absolute improvement on RoBERTa-base encoders on seven GLUE benchmark datasets compared with the best baseline. We present an indepth analysis of our framework design, revealing the non-trivial contributions of our network components. Our code is publicly available for better reproducibility.1
Guanyi Mou, Yichuan Li 0001, Kyumin Lee
IEEE BigData2
2020 Toward A Multilingual and Multimodal Data Repository for COVID-19 Disinformation
abstract
The COVID-19 epidemic is considered as the global health crisis of the whole society and the greatest challenge mankind faced since World War Two. Unfortunately, the fake news about COVID-19 is spreading as fast as the virus itself. The incorrect health measurements, anxiety, and hate speeches will have bad consequences on people's physical health, as well as their mental health in the whole world. To help better combat the COVID-19 fake news, we propose a new fake news detection dataset MM-COVID1(Multilingual and Multidimensional COVID-19 Fake News Data Repository). This dataset provides the multilingual fake news and the relevant social context. We collect 3981 pieces of fake news content and 7192 trustworthy information from English, Spanish, Portuguese, Hindi, French and Italian, 6 different languages. We present a detailed and exploratory analysis of MM-COVID from different perspectives.
Yichuan Li 0001, Bohan Jiang, Kai Shu, Huan Liu 0001
IEEE BigData1
2020 Early Detection of Fake News with Multi-source Weak Social Supervision
Kai Shu, Guoqing Zheng, Yichuan Li 0001, Subhabrata Mukherjee, Ahmed Awadallah 0001, Scott W. Ruston, Huan Liu 0001
ECML/PKDD (3)3