VLDB 2026 Research / reviewers in the wild / expert
Toyin Aguda
dblp:348/6888 · also Toyin D. Aguda
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0005-1582-754XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Machine translation · 65% Information extraction and text analysis · 35% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational finance and economics · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
domain-specific machine translation |
0.9 | 1 | 2025 | Translating Domain-Specific Terminology in Typologically-Diverse Languages: A Study in Tax and Financial Education · EMNLP 2025 |
Natural language and speech › Machine translation
terminology translation |
0.9 | 1 | 2025 | Translating Domain-Specific Terminology in Typologically-Diverse Languages: A Study in Tax and Financial Education · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.7 | 1 | 2023 | REFinD: Relation Extraction Financial Dataset · SIGIR 2023 |
Computational finance and economics › financial data analysis
financial document analysis |
0.2 | 1 | 2023 | REFinD: Relation Extraction Financial Dataset · SIGIR 2023 |
Methods — techniques the papers use, named apart from their topics
deep learning · 1.3terminology-aided translation · 0.9LLM evaluation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Translating Domain-Specific Terminology in Typologically-Diverse Languages: A Study in Tax and Financial EducationabstractDomain-specific multilingual terminology is essential for accurate machine translation (MT) and cross-lingual NLP applications.We present a gold-standard terminology resource for the tax and financial education domains, built from curated governmental publications and covering seven typologically diverse languages: English, Spanish, Russian, Vietnamese, Korean, Chinese (traditional and simplified) and Haitian Creole.Using this resource, we assess various MT systems and LLMs on translation quality and term accuracy.We annotate over 3,000 terms for domain-specificity, facilitating a comparison between domain-specific and general term translations, and observe models' challenges with specialized tax terms.We also analyze the case of terminology-aided translation, and the LLMs' performance in extracting the translated term given the context.Our results highlight model limitations and the value of high-quality terminologies for advancing MT research in specialized contexts.1 * Contribution done while working at JPMorgan. 1 Please contact the author(s) if you want to have access to the terminologies and parallel data. Arturo Oncevay, Elena Kochkina, Keshav Ramani, Toyin Aguda, Simerjot Kaur, Charese Smiley |
EMNLP | 4 |
| 2024 | Large Language Models as Financial Data Annotators: A Study on Effectiveness and EfficiencyabstractCollecting labeled datasets in finance is challenging due to scarcity of domain experts and higher cost of employing them. While Large Language Models (LLMs) have demonstrated remarkable performance in data annotation tasks on general domain datasets, their effectiveness on domain specific datasets remains under-explored. To address this gap, we investigate the potential of LLMs as efficient data annotators for extracting relations in financial documents. We compare the annotations produced by three LLMs (GPT-4, PaLM 2, and MPT Instruct) against expert annotators and crowdworkers. We demonstrate that the current state-of-the-art LLMs can be sufficient alternatives to non-expert crowdworkers. We analyze models using various prompts and parameter settings and find that customizing the prompts for each relation group by providing specific examples belonging to those groups is paramount. Furthermore, we introduce a reliability index (LLM-RelIndex) used to identify outputs that may require expert attention. Finally, we perform an extensive time, cost and error analysis and provide recommendations for the collection and usage of automated annotations in domain-specific settings. Toyin Aguda, Suchetha Siddagangappa, Elena Kochkina, Simerjot Kaur, Dongsheng Wang 0005, Charese Smiley |
LREC/COLING | 1 |
| 2023 | REFinD: Relation Extraction Financial DatasetabstractA number of datasets for Relation Extraction (RE) have been created to aide downstream tasks such as information retrieval, semantic search, question answering and textual entailment. However, these datasets fail to capture financial-domain specific challenges since most of these datasets are compiled using general knowledge sources such as Wikipedia, web-based text and news articles, hindering real-life progress and adoption within the financial world. To address this limitation, we propose REFinD, the first large-scale annotated dataset of relations, with ~29K instances and 22 relations amongst 8 types of entity pairs, generated entirely over financial documents. We also provide an empirical evaluation with various state-of-the-art models as benchmarks for the RE task and highlight the challenges posed by our dataset. We observed that various state-of-the-art deep learning models struggle with numeric inference, relational and directional ambiguity. To encourage further research in this direction, REFinD is available at https://www.jpmorgan.com/technology/artificial-intelligence/initiatives/refind-dataset/problem-motivation-outcome. Simerjot Kaur, Charese Smiley, Akshat Gupta, Joy Prakash Sain, Dongsheng Wang 0005, Suchetha Siddagangappa, Toyin Aguda, Sameena Shah |
SIGIR | 7 |