Michael Shliselberg

dblp:326/1278 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2024
0009-0007-9220-7226ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Web and social media mining · 44% Information retrieval · 44% Data mining · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › fact-checking
claim matching
0.812024
SynDy: Synthetic Dynamic Dataset Generation Framework for Misinformation Tasks · SIGIR 2024
Web and social media mining
misinformation
0.812024
SynDy: Synthetic Dynamic Dataset Generation Framework for Misinformation Tasks · SIGIR 2024
Data mining › clustering › document clustering
topical clustering
0.212024
SynDy: Synthetic Dynamic Dataset Generation Framework for Misinformation Tasks · SIGIR 2024

Methods — techniques the papers use, named apart from their topics

synthetic data generation · 0.8large language model · 0.8distant supervision · 0.8
YearPublicationVenuePosition
2024 SynDy: Synthetic Dynamic Dataset Generation Framework for Misinformation Tasks
abstract
Diaspora communities are disproportionately impacted by off-the-radar misinformation and often neglected by mainstream fact-checking efforts, creating a critical need to scale-up efforts of nascent fact-checking initiatives. In this paper we present SynDy, a framework for Synthetic Dynamic Dataset Generation to leverage the capabilities of the largest frontier Large Language Models (LLMs) to train local, specialized language models. To the best of our knowledge, SynDy is the first paper utilizing LLMs to create fine-grained synthetic labels for tasks of direct relevance to misinformation mitigation, namely Claim Matching, Topical Clustering, and Claim Relationship Classification. SynDy utilizes LLMs and social media queries to automatically generate distantly-supervised, topically-focused datasets with synthetic labels on these three tasks, providing essential tools to scale up human-led fact-checking at a fraction of the cost of human-annotated data. Training on SynDy's generated labels shows improvement over a standard baseline and is not significantly worse compared to training on human labels (which may be infeasible to acquire). SynDy is being integrated into Meedan's chatbot tiplines that are used by over 50 organizations, serve over 230K users annually, and automatically distribute human-written fact-checks via messaging apps such as WhatsApp. SynDy will also be integrated into our deployed Co-Insights toolkit, enabling low-resource organizations to launch tiplines for their communities. Finally, we envision SynDy enabling additional fact-checking tools such as matching new misinformation claims to high-quality explainers on common misinformation topics.
Michael Shliselberg, Ashkan Kazemi, Scott A. Hale, Shiri Dori-Hacohen
SIGIR1