VLDB 2026 Research / reviewers in the wild / expert
Simerjot Kaur
dblp:305/3802
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-5863-4749ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Information extraction and text analysis · 34% Machine translation · 31% Trustworthy machine learning · 16% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational finance and economics · 100% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › calibration
confidence calibration |
0.9 | 1 | 2025 | Calibrating LLM Confidence by Probing Perturbed Representation Stability · EMNLP 2025 |
Natural language and speech › Machine translation
domain-specific machine translation |
0.9 | 1 | 2025 | Translating Domain-Specific Terminology in Typologically-Diverse Languages: A Study in Tax and Financial Education · EMNLP 2025 |
Natural language and speech › Machine translation
terminology translation |
0.9 | 1 | 2025 | Translating Domain-Specific Terminology in Typologically-Diverse Languages: A Study in Tax and Financial Education · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis › document understanding
layout-aware document understanding |
0.8 | 1 | 2024 | DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding · ACL (1) 2024 |
Computer vision › Vision and language › multimodal understanding
multimodal document understanding |
0.8 | 1 | 2024 | DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding · ACL (1) 2024 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.7 | 1 | 2023 | REFinD: Relation Extraction Financial Dataset · SIGIR 2023 |
Natural language and speech › Information extraction and text analysis
document understanding |
0.2 | 1 | 2024 | DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding · ACL (1) 2024 |
Computational finance and economics › financial data analysis
financial document analysis |
0.2 | 1 | 2023 | REFinD: Relation Extraction Financial Dataset · SIGIR 2023 |
Methods — techniques the papers use, named apart from their topics
deep learning · 1.3terminology-aided translation · 0.9representation perturbation · 0.9probing · 0.9LLM evaluation · 0.9layout-aware modeling · 0.8generative language modeling · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Calibrating LLM Confidence by Probing Perturbed Representation StabilityabstractReza Khanmohammadi, Erfan Miahi, Mehrsa Mardikoraem, Simerjot Kaur, Ivan Brugere, Charese Smiley, Kundan S Thind, Mohammad M. Ghassemi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Reza Khanmohammadi, Erfan Miahi, Mehrsa Mardikoraem, Simerjot Kaur, Ivan Brugere, Charese Smiley, Kundan Thind, Mohammad M. Ghassemi |
EMNLP | 4 |
| 2025 | Translating Domain-Specific Terminology in Typologically-Diverse Languages: A Study in Tax and Financial EducationabstractDomain-specific multilingual terminology is essential for accurate machine translation (MT) and cross-lingual NLP applications.We present a gold-standard terminology resource for the tax and financial education domains, built from curated governmental publications and covering seven typologically diverse languages: English, Spanish, Russian, Vietnamese, Korean, Chinese (traditional and simplified) and Haitian Creole.Using this resource, we assess various MT systems and LLMs on translation quality and term accuracy.We annotate over 3,000 terms for domain-specificity, facilitating a comparison between domain-specific and general term translations, and observe models' challenges with specialized tax terms.We also analyze the case of terminology-aided translation, and the LLMs' performance in extracting the translated term given the context.Our results highlight model limitations and the value of high-quality terminologies for advancing MT research in specialized contexts.1 * Contribution done while working at JPMorgan. 1 Please contact the author(s) if you want to have access to the terminologies and parallel data. Arturo Oncevay, Elena Kochkina, Keshav Ramani, Toyin Aguda, Simerjot Kaur, Charese Smiley |
EMNLP | 5 |
| 2025 | Investigating the Temporal Association of Biomedical Research on Small Business Funding: A Bibliometric and Data Analytic ApproachabstractThe relationship between scientific innovation in biomedical sciences and its impact on industrial activities is a complex and dynamic process. This article investigates the relationship between science and industrial innovation, focusing on how the historical impact and content of scientific paper abstracts are associated with future funding and innovation grant application content for small businesses. The research incorporates bibliometric analyses along with small business innovation research (SBIR) data to yield a holistic view of the science-industry interface. We quantify the temporal effects and impact latency of scientific advancements on industrial activity across 10873 topics and take into account their taxonomic relationships, spanning from 2010 to 2021. We find that the impact of scientific advances on industrial projects across different thematic depths consistently exhibitedp-values less than 0.05, underscoring the significant predictive power of contemporary scientific activities on future industrial projects. Further, we demonstrate that the semantic contents of scientific paper abstracts within a topic are associated with future industrial project description text embeddings. The frequency analysis reveals that various scientific activities significantly inform future industrial project funding across varying depths of MeSH topic categorization, highlighting the significant role of science in steering industrial innovation. This study demonstrates that the impact of scientific research on industrial innovation extends beyond the mere volume of scientific output, but is greatly influenced by its impact, the broader themes it advances, and the meaningful narratives it presents. Reza Khanmohammadi, Simerjot Kaur, Charese Smiley, Tuka Al Hanai, Ivan Brugere, Armineh Nourbakhsh, Mohammad M. Ghassemi |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | DocLLM: A Layout-Aware Generative Language Model for Multimodal Document UnderstandingabstractDongsheng Wang, Natraj Raman, Mathieu Sibue, Zhiqiang Ma, Petr Babkin, Simerjot Kaur, Yulong Pei, Armineh Nourbakhsh, Xiaomo Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dongsheng Wang 0005, Natraj Raman, Mathieu Sibue, Petr Babkin, Simerjot Kaur, Yulong Pei, Armineh Nourbakhsh, Xiaomo Liu |
ACL (1) | 6 |
| 2024 | Large Language Models as Financial Data Annotators: A Study on Effectiveness and EfficiencyabstractCollecting labeled datasets in finance is challenging due to scarcity of domain experts and higher cost of employing them. While Large Language Models (LLMs) have demonstrated remarkable performance in data annotation tasks on general domain datasets, their effectiveness on domain specific datasets remains under-explored. To address this gap, we investigate the potential of LLMs as efficient data annotators for extracting relations in financial documents. We compare the annotations produced by three LLMs (GPT-4, PaLM 2, and MPT Instruct) against expert annotators and crowdworkers. We demonstrate that the current state-of-the-art LLMs can be sufficient alternatives to non-expert crowdworkers. We analyze models using various prompts and parameter settings and find that customizing the prompts for each relation group by providing specific examples belonging to those groups is paramount. Furthermore, we introduce a reliability index (LLM-RelIndex) used to identify outputs that may require expert attention. Finally, we perform an extensive time, cost and error analysis and provide recommendations for the collection and usage of automated annotations in domain-specific settings. Toyin Aguda, Suchetha Siddagangappa, Elena Kochkina, Simerjot Kaur, Dongsheng Wang 0005, Charese Smiley |
LREC/COLING | 4 |
| 2023 | REFinD: Relation Extraction Financial DatasetabstractA number of datasets for Relation Extraction (RE) have been created to aide downstream tasks such as information retrieval, semantic search, question answering and textual entailment. However, these datasets fail to capture financial-domain specific challenges since most of these datasets are compiled using general knowledge sources such as Wikipedia, web-based text and news articles, hindering real-life progress and adoption within the financial world. To address this limitation, we propose REFinD, the first large-scale annotated dataset of relations, with ~29K instances and 22 relations amongst 8 types of entity pairs, generated entirely over financial documents. We also provide an empirical evaluation with various state-of-the-art models as benchmarks for the RE task and highlight the challenges posed by our dataset. We observed that various state-of-the-art deep learning models struggle with numeric inference, relational and directional ambiguity. To encourage further research in this direction, REFinD is available at https://www.jpmorgan.com/technology/artificial-intelligence/initiatives/refind-dataset/problem-motivation-outcome. Simerjot Kaur, Charese Smiley, Akshat Gupta, Joy Prakash Sain, Dongsheng Wang 0005, Suchetha Siddagangappa, Toyin Aguda, Sameena Shah |
SIGIR | 1 |