Ashkan Kazemi

dblp:277/9400 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0002-2475-1007ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 56% Web and social media mining · 34% Data mining · 10%
Artificial intelligence
1 paper
Information extraction and text analysis · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › fact-checking
claim matching
0.812024
SynDy: Synthetic Dynamic Dataset Generation Framework for Misinformation Tasks · SIGIR 2024
Web and social media mining
misinformation
0.812024
SynDy: Synthetic Dynamic Dataset Generation Framework for Misinformation Tasks · SIGIR 2024
Information retrieval
cross-language information retrieval
0.512021
Claim Matching Beyond English to Scale Global Fact-Checking · ACL/IJCNLP (1) 2021
Data mining › clustering › document clustering
topical clustering
0.212024
SynDy: Synthetic Dynamic Dataset Generation Framework for Misinformation Tasks · SIGIR 2024

Methods — techniques the papers use, named apart from their topics

claim matching · 1.0synthetic data generation · 0.8large language model · 0.8distant supervision · 0.8
YearPublicationVenuePosition
2024 Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language Models
abstract
Recent progress in large language models (LLMs) has enabled the deployment of many generative NLP applications. At the same time, it has also led to a misleading public discourse that “it’s all been solved.” Not surprisingly, this has, in turn, made many NLP researchers – especially those at the beginning of their careers – worry about what NLP research area they should focus on. Has it all been solved, or what remaining questions can we work on regardless of LLMs? To address this question, this paper compiles NLP research directions rich for exploration. We identify fourteen different research areas encompassing 45 research directions that require new research and are not directly solvable by LLMs. While we identify many research areas, many others exist; we do not cover areas currently addressed by LLMs, but where LLMs lag behind in performance or those focused on LLM development. We welcome suggestions for other research directions to include: https://bit.ly/nlp-era-llm.
Oana Ignat, Zhijing Jin 0001, Artem Abzaliev, Laura Biester, Santiago Castro, Naihao Deng, Xinyi Gao 0004, Aylin Gunal, Jacky He, Ashkan Kazemi, Muhammad Khalifa, Namho Koh, Andrew Lee 0001, Siyang Liu 0003, Do June Min, Shinka Mori, Joan Nwatu, Verónica Pérez-Rosas, Zekun Wang 0002, Winston Wu, Rada Mihalcea
LREC/COLING10
2024 SynDy: Synthetic Dynamic Dataset Generation Framework for Misinformation Tasks
abstract
Diaspora communities are disproportionately impacted by off-the-radar misinformation and often neglected by mainstream fact-checking efforts, creating a critical need to scale-up efforts of nascent fact-checking initiatives. In this paper we present SynDy, a framework for Synthetic Dynamic Dataset Generation to leverage the capabilities of the largest frontier Large Language Models (LLMs) to train local, specialized language models. To the best of our knowledge, SynDy is the first paper utilizing LLMs to create fine-grained synthetic labels for tasks of direct relevance to misinformation mitigation, namely Claim Matching, Topical Clustering, and Claim Relationship Classification. SynDy utilizes LLMs and social media queries to automatically generate distantly-supervised, topically-focused datasets with synthetic labels on these three tasks, providing essential tools to scale up human-led fact-checking at a fraction of the cost of human-annotated data. Training on SynDy's generated labels shows improvement over a standard baseline and is not significantly worse compared to training on human labels (which may be infeasible to acquire). SynDy is being integrated into Meedan's chatbot tiplines that are used by over 50 organizations, serve over 230K users annually, and automatically distribute human-written fact-checks via messaging apps such as WhatsApp. SynDy will also be integrated into our deployed Co-Insights toolkit, enabling low-resource organizations to launch tiplines for their communities. Finally, we envision SynDy enabling additional fact-checking tools such as matching new misinformation claims to high-quality explainers on common misinformation topics.
Michael Shliselberg, Ashkan Kazemi, Scott A. Hale, Shiri Dori-Hacohen
SIGIR2
2023 Query Rewriting for Effective Misinformation Discovery
abstract
Ashkan Kazemi, Artem Abzaliev, Naihao Deng, Rui Hou, Scott Hale, Veronica Perez-Rosas, Rada Mihalcea. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ashkan Kazemi, Artem Abzaliev, Naihao Deng, Rui Hou 0007, Scott A. Hale, Verónica Pérez-Rosas, Rada Mihalcea
IJCNLP (1)1
2021 Claim Matching Beyond English to Scale Global Fact-Checking
abstract
Ashkan Kazemi, Kiran Garimella, Devin Gaffney, Scott Hale. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ashkan Kazemi, Venkata Rama Kiran Garimella, Devin Gaffney, Scott A. Hale
ACL/IJCNLP (1)1
2020 Biased TextRank: Unsupervised Graph-Based Content Extraction
abstract
We introduce Biased TextRank, a graph-based content extraction method inspired by the popular TextRank algorithm that ranks text spans according to their importance for language processing tasks and according to their relevance to an input "focus."Biased TextRank enables focused content extraction for text by modifying the random restarts in the execution of TextRank.The random restart probabilities are assigned based on the relevance of the graph nodes to the focus of the task.We present two applications of Biased TextRank: focused summarization and explanation extraction, and show that our algorithm leads to improved performance on two different datasets by significant ROUGE-N score margins.Much like its predecessor, Biased TextRank is unsupervised, easy to implement and orders of magnitude faster and lighter than current state-ofthe-art Natural Language Processing methods for similar tasks.
Ashkan Kazemi, Verónica Pérez-Rosas, Rada Mihalcea
COLING1