Simerjot Kaur

dblp:305/3802 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-5863-4749ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Information extraction and text analysis · 34% Machine translation · 31% Trustworthy machine learning · 16%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational finance and economics · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › calibration
confidence calibration
0.912025
Calibrating LLM Confidence by Probing Perturbed Representation Stability · EMNLP 2025
Natural language and speech › Machine translation
domain-specific machine translation
0.912025
Translating Domain-Specific Terminology in Typologically-Diverse Languages: A Study in Tax and Financial Education · EMNLP 2025
Natural language and speech › Machine translation
terminology translation
0.912025
Translating Domain-Specific Terminology in Typologically-Diverse Languages: A Study in Tax and Financial Education · EMNLP 2025
Natural language and speech › Information extraction and text analysis › document understanding
layout-aware document understanding
0.812024
DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding · ACL (1) 2024
Computer vision › Vision and language › multimodal understanding
multimodal document understanding
0.812024
DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding · ACL (1) 2024
Natural language and speech › Information extraction and text analysis
relation extraction
0.712023
REFinD: Relation Extraction Financial Dataset · SIGIR 2023
Natural language and speech › Information extraction and text analysis
document understanding
0.212024
DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding · ACL (1) 2024
Computational finance and economics › financial data analysis
financial document analysis
0.212023
REFinD: Relation Extraction Financial Dataset · SIGIR 2023

Methods — techniques the papers use, named apart from their topics

deep learning · 1.3terminology-aided translation · 0.9representation perturbation · 0.9probing · 0.9LLM evaluation · 0.9layout-aware modeling · 0.8generative language modeling · 0.8
YearPublicationVenuePosition
2025 Calibrating LLM Confidence by Probing Perturbed Representation Stability
abstract
Reza Khanmohammadi, Erfan Miahi, Mehrsa Mardikoraem, Simerjot Kaur, Ivan Brugere, Charese Smiley, Kundan S Thind, Mohammad M. Ghassemi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Reza Khanmohammadi, Erfan Miahi, Mehrsa Mardikoraem, Simerjot Kaur, Ivan Brugere, Charese Smiley, Kundan Thind, Mohammad M. Ghassemi
EMNLP4
2025 Translating Domain-Specific Terminology in Typologically-Diverse Languages: A Study in Tax and Financial Education
abstract
Domain-specific multilingual terminology is essential for accurate machine translation (MT) and cross-lingual NLP applications.We present a gold-standard terminology resource for the tax and financial education domains, built from curated governmental publications and covering seven typologically diverse languages: English, Spanish, Russian, Vietnamese, Korean, Chinese (traditional and simplified) and Haitian Creole.Using this resource, we assess various MT systems and LLMs on translation quality and term accuracy.We annotate over 3,000 terms for domain-specificity, facilitating a comparison between domain-specific and general term translations, and observe models' challenges with specialized tax terms.We also analyze the case of terminology-aided translation, and the LLMs' performance in extracting the translated term given the context.Our results highlight model limitations and the value of high-quality terminologies for advancing MT research in specialized contexts.1 * Contribution done while working at JPMorgan. 1 Please contact the author(s) if you want to have access to the terminologies and parallel data.
Arturo Oncevay, Elena Kochkina, Keshav Ramani, Toyin Aguda, Simerjot Kaur, Charese Smiley
EMNLP5
2025 Investigating the Temporal Association of Biomedical Research on Small Business Funding: A Bibliometric and Data Analytic Approach
abstract
The relationship between scientific innovation in biomedical sciences and its impact on industrial activities is a complex and dynamic process. This article investigates the relationship between science and industrial innovation, focusing on how the historical impact and content of scientific paper abstracts are associated with future funding and innovation grant application content for small businesses. The research incorporates bibliometric analyses along with small business innovation research (SBIR) data to yield a holistic view of the science-industry interface. We quantify the temporal effects and impact latency of scientific advancements on industrial activity across 10873 topics and take into account their taxonomic relationships, spanning from 2010 to 2021. We find that the impact of scientific advances on industrial projects across different thematic depths consistently exhibitedp-values less than 0.05, underscoring the significant predictive power of contemporary scientific activities on future industrial projects. Further, we demonstrate that the semantic contents of scientific paper abstracts within a topic are associated with future industrial project description text embeddings. The frequency analysis reveals that various scientific activities significantly inform future industrial project funding across varying depths of MeSH topic categorization, highlighting the significant role of science in steering industrial innovation. This study demonstrates that the impact of scientific research on industrial innovation extends beyond the mere volume of scientific output, but is greatly influenced by its impact, the broader themes it advances, and the meaningful narratives it presents.
Reza Khanmohammadi, Simerjot Kaur, Charese Smiley, Tuka Al Hanai, Ivan Brugere, Armineh Nourbakhsh, Mohammad M. Ghassemi
IEEE Trans. Comput. Soc. Syst.2
2024 DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding
abstract
Dongsheng Wang, Natraj Raman, Mathieu Sibue, Zhiqiang Ma, Petr Babkin, Simerjot Kaur, Yulong Pei, Armineh Nourbakhsh, Xiaomo Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Dongsheng Wang 0005, Natraj Raman, Mathieu Sibue, Petr Babkin, Simerjot Kaur, Yulong Pei, Armineh Nourbakhsh, Xiaomo Liu
ACL (1)6
2024 Large Language Models as Financial Data Annotators: A Study on Effectiveness and Efficiency
abstract
Collecting labeled datasets in finance is challenging due to scarcity of domain experts and higher cost of employing them. While Large Language Models (LLMs) have demonstrated remarkable performance in data annotation tasks on general domain datasets, their effectiveness on domain specific datasets remains under-explored. To address this gap, we investigate the potential of LLMs as efficient data annotators for extracting relations in financial documents. We compare the annotations produced by three LLMs (GPT-4, PaLM 2, and MPT Instruct) against expert annotators and crowdworkers. We demonstrate that the current state-of-the-art LLMs can be sufficient alternatives to non-expert crowdworkers. We analyze models using various prompts and parameter settings and find that customizing the prompts for each relation group by providing specific examples belonging to those groups is paramount. Furthermore, we introduce a reliability index (LLM-RelIndex) used to identify outputs that may require expert attention. Finally, we perform an extensive time, cost and error analysis and provide recommendations for the collection and usage of automated annotations in domain-specific settings.
Toyin Aguda, Suchetha Siddagangappa, Elena Kochkina, Simerjot Kaur, Dongsheng Wang 0005, Charese Smiley
LREC/COLING4
2023 REFinD: Relation Extraction Financial Dataset
abstract
A number of datasets for Relation Extraction (RE) have been created to aide downstream tasks such as information retrieval, semantic search, question answering and textual entailment. However, these datasets fail to capture financial-domain specific challenges since most of these datasets are compiled using general knowledge sources such as Wikipedia, web-based text and news articles, hindering real-life progress and adoption within the financial world. To address this limitation, we propose REFinD, the first large-scale annotated dataset of relations, with ~29K instances and 22 relations amongst 8 types of entity pairs, generated entirely over financial documents. We also provide an empirical evaluation with various state-of-the-art models as benchmarks for the RE task and highlight the challenges posed by our dataset. We observed that various state-of-the-art deep learning models struggle with numeric inference, relational and directional ambiguity. To encourage further research in this direction, REFinD is available at https://www.jpmorgan.com/technology/artificial-intelligence/initiatives/refind-dataset/problem-motivation-outcome.
Simerjot Kaur, Charese Smiley, Akshat Gupta, Joy Prakash Sain, Dongsheng Wang 0005, Suchetha Siddagangappa, Toyin Aguda, Sameena Shah
SIGIR1