Vedaant Shah

dblp:348/4704 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Information extraction and text analysis · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › multilingual NLP
cross-lingual information extraction
0.912025
Translation and Fusion Improves Cross-lingual Information Extraction · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › named entity recognition
low-resource named entity recognition
0.912025
Translation and Fusion Improves Cross-lingual Information Extraction · ACL (1) 2025
Natural language and speech › Information extraction and text analysis
named entity recognition
0.912025
Translation and Fusion Improves Cross-lingual Information Extraction · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

translation · 0.9instruction tuning · 0.9annotation fusion · 0.9
YearPublicationVenuePosition
2025 Translation and Fusion Improves Cross-lingual Information Extraction
abstract
Large language models (LLMs) combined with instruction tuning have shown significant progress in information extraction (IE) tasks, exhibiting strong generalization capabilities to unseen datasets by following annotation guidelines.However, their applicability to lowresource languages remains limited due to lack of both labeled data for fine-tuning, and unlabeled text for pre-training.In this paper, we propose TransFusion, a framework in which models are fine-tuned to use English translations of low-resource language data, enabling more precise predictions through annotation fusion.Based on TransFusion, we introduce GoLLIE-TF, a cross-lingual instruction-tuned LLM for IE tasks, designed to close the performance gap between high and low-resource languages.Our experiments across twelve multilingual IE datasets spanning 50 languages demonstrate that GoLLIE-TF achieves better cross-lingual transfer over the base model.In addition, we show that TransFusion significantly improves low-resource language named entity recognition when applied to proprietary models such as GPT-4 (+5 F1) with a prompting approach, or fine-tuning different language models including decoder-only (+14 F1) and encoder-only (+13 F1) architectures.
Yang Chen 0065, Vedaant Shah, Alan Ritter
ACL (1)2