VLDB 2026 Research / reviewers in the wild / expert
Charese Smiley
dblp:165/0746 · also Charese H. Smiley
· DBLP profile ↗
19ranked-venue papers
2as first author
13since 2021 · last 2026
0009-0007-5575-2313ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Systematic Multi-Aspect Evaluation of Time Series-Based Report Generation: The Case of Financial Analysis from Stock Data
Elizabeth Fons, Elena Kochkina, Rachneet Kaur, Berowne Hlavaty, Charese Smiley, Svitlana Vyetrenko, Manuela M. Veloso |
LREC | 6 |
| 2025 | AfroCS-xs: Creating a Compact, High-Quality, Human-Validated Code-Switched Dataset for African LanguagesabstractKayode Olaleye, Arturo Oncevay, Mathieu Sibue, Nombuyiselo Zondi, Michelle Terblanche, Sibongile Mapikitla, Richard Lastrucci, Charese Smiley, Vukosi Marivate. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Kayode Olaleye, Arturo Oncevay, Mathieu Sibue, Nombuyiselo Zondi, Michelle Terblanche, Sibongile Mapikitla, Richard Lastrucci, Charese Smiley, Vukosi Marivate |
ACL (1) | 8 |
| 2025 | Calibrating LLM Confidence by Probing Perturbed Representation StabilityabstractReza Khanmohammadi, Erfan Miahi, Mehrsa Mardikoraem, Simerjot Kaur, Ivan Brugere, Charese Smiley, Kundan S Thind, Mohammad M. Ghassemi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Reza Khanmohammadi, Erfan Miahi, Mehrsa Mardikoraem, Simerjot Kaur, Ivan Brugere, Charese Smiley, Kundan Thind, Mohammad M. Ghassemi |
EMNLP | 6 |
| 2025 | Translating Domain-Specific Terminology in Typologically-Diverse Languages: A Study in Tax and Financial EducationabstractDomain-specific multilingual terminology is essential for accurate machine translation (MT) and cross-lingual NLP applications.We present a gold-standard terminology resource for the tax and financial education domains, built from curated governmental publications and covering seven typologically diverse languages: English, Spanish, Russian, Vietnamese, Korean, Chinese (traditional and simplified) and Haitian Creole.Using this resource, we assess various MT systems and LLMs on translation quality and term accuracy.We annotate over 3,000 terms for domain-specificity, facilitating a comparison between domain-specific and general term translations, and observe models' challenges with specialized tax terms.We also analyze the case of terminology-aided translation, and the LLMs' performance in extracting the translated term given the context.Our results highlight model limitations and the value of high-quality terminologies for advancing MT research in specialized contexts.1 * Contribution done while working at JPMorgan. 1 Please contact the author(s) if you want to have access to the terminologies and parallel data. Arturo Oncevay, Elena Kochkina, Keshav Ramani, Toyin Aguda, Simerjot Kaur, Charese Smiley |
EMNLP | 6 |
| 2025 | The Impact of Domain-Specific Terminology on Machine Translation for Finance in European LanguagesabstractArturo Oncevay, Charese Smiley, Xiaomo Liu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Arturo Oncevay, Charese Smiley, Xiaomo Liu |
NAACL (Long Papers) | 2 |
| 2025 | Investigating the Temporal Association of Biomedical Research on Small Business Funding: A Bibliometric and Data Analytic ApproachabstractThe relationship between scientific innovation in biomedical sciences and its impact on industrial activities is a complex and dynamic process. This article investigates the relationship between science and industrial innovation, focusing on how the historical impact and content of scientific paper abstracts are associated with future funding and innovation grant application content for small businesses. The research incorporates bibliometric analyses along with small business innovation research (SBIR) data to yield a holistic view of the science-industry interface. We quantify the temporal effects and impact latency of scientific advancements on industrial activity across 10873 topics and take into account their taxonomic relationships, spanning from 2010 to 2021. We find that the impact of scientific advances on industrial projects across different thematic depths consistently exhibitedp-values less than 0.05, underscoring the significant predictive power of contemporary scientific activities on future industrial projects. Further, we demonstrate that the semantic contents of scientific paper abstracts within a topic are associated with future industrial project description text embeddings. The frequency analysis reveals that various scientific activities significantly inform future industrial project funding across varying depths of MeSH topic categorization, highlighting the significant role of science in steering industrial innovation. This study demonstrates that the impact of scientific research on industrial innovation extends beyond the mere volume of scientific output, but is greatly influenced by its impact, the broader themes it advances, and the meaningful narratives it presents. Reza Khanmohammadi, Simerjot Kaur, Charese Smiley, Tuka Al Hanai, Ivan Brugere, Armineh Nourbakhsh, Mohammad M. Ghassemi |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Large Language Models as Financial Data Annotators: A Study on Effectiveness and EfficiencyabstractCollecting labeled datasets in finance is challenging due to scarcity of domain experts and higher cost of employing them. While Large Language Models (LLMs) have demonstrated remarkable performance in data annotation tasks on general domain datasets, their effectiveness on domain specific datasets remains under-explored. To address this gap, we investigate the potential of LLMs as efficient data annotators for extracting relations in financial documents. We compare the annotations produced by three LLMs (GPT-4, PaLM 2, and MPT Instruct) against expert annotators and crowdworkers. We demonstrate that the current state-of-the-art LLMs can be sufficient alternatives to non-expert crowdworkers. We analyze models using various prompts and parameter settings and find that customizing the prompts for each relation group by providing specific examples belonging to those groups is paramount. Furthermore, we introduce a reliability index (LLM-RelIndex) used to identify outputs that may require expert attention. Finally, we perform an extensive time, cost and error analysis and provide recommendations for the collection and usage of automated annotations in domain-specific settings. Toyin Aguda, Suchetha Siddagangappa, Elena Kochkina, Simerjot Kaur, Dongsheng Wang 0005, Charese Smiley |
LREC/COLING | 6 |
| 2023 | Robust NLP for Finance (RobustFin)abstractNatural language processing (NLP) technologies have been widely applied in business domains such as e-commerce and customer service, but their adoption in the financial sector has been constrained by industry-specific performance standards and regulatory restrictions. This challenge has created new opportunities for core research in related areas. Recent advancements in NLP, such as the advent of large language models, has encouraged adoption in the finance sector. However, compared to other domains, finance has stricter requirements for robustness, explainability, and generalizability. Given this background, we propose to organize the first Robust NLP for Finance (RobustFin) workshop at KDD '23 to encourage the study of and research on robustness and explainability technologies with regard to financial NLP. The goal of the workshop is to extend the applications of NLP in finance, while motivating further research in robust NLP. Sameena Shah, Xiaodan Zhu 0001, Gerard de Melo, Armineh Nourbakhsh, Xiaomo Liu, Charese Smiley, Zhiyu Chen 0002 |
KDD | 7 |
| 2023 | REFinD: Relation Extraction Financial DatasetabstractA number of datasets for Relation Extraction (RE) have been created to aide downstream tasks such as information retrieval, semantic search, question answering and textual entailment. However, these datasets fail to capture financial-domain specific challenges since most of these datasets are compiled using general knowledge sources such as Wikipedia, web-based text and news articles, hindering real-life progress and adoption within the financial world. To address this limitation, we propose REFinD, the first large-scale annotated dataset of relations, with ~29K instances and 22 relations amongst 8 types of entity pairs, generated entirely over financial documents. We also provide an empirical evaluation with various state-of-the-art models as benchmarks for the RE task and highlight the challenges posed by our dataset. We observed that various state-of-the-art deep learning models struggle with numeric inference, relational and directional ambiguity. To encourage further research in this direction, REFinD is available at https://www.jpmorgan.com/technology/artificial-intelligence/initiatives/refind-dataset/problem-motivation-outcome. Simerjot Kaur, Charese Smiley, Akshat Gupta, Joy Prakash Sain, Dongsheng Wang 0005, Suchetha Siddagangappa, Toyin Aguda, Sameena Shah |
SIGIR | 2 |
| 2023 | Knowledge Discovery from Unstructured Data in Financial Services (KDF) WorkshopabstractKnowledge discovery from unstructured data, including business documents, web content, and news articles, has been a key AI challenge for the financial services industry. Comprehending these corpora and discovering knowledge from them, which could be textual, tabular, or graphic, are the cornerstone of supporting business decisions in the financial services domain, where information retrieval and content analysis techniques are of fundamental importance. We propose a workshop on knowledge discovery from unstructured data in financial services at SIGIR 2023 to highlight the current and emerging opportunities, invite original research, and prompt success sharing between researchers. Sameena Shah, Xiaodan Zhu 0001, Wenhu Chen, Manling Li, Armineh Nourbakhsh, Xiaomo Liu, Charese Smiley, Yulong Pei, Akshat Gupta |
SIGIR | 8 |
| 2022 | ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question AnsweringabstractWith the recent advance in large pre-trained language models, researchers have achieved record performances in NLP tasks that mostly focus on language pattern matching.The community is experiencing the shift of the challenge from how to model language to the imitation of complex reasoning abilities like human beings.In this work, we investigate the application domain of finance that involves realworld, complex numerical reasoning.We propose a new large-scale dataset, CONVFINQA, aiming to study the chain of numerical reasoning in conversational question answering.Our dataset poses great challenge in modeling longrange, complex numerical reasoning paths in real-world conversations.We conduct comprehensive experiments and analyses with both the neural symbolic methods and the promptingbased methods, to provide insights into the reasoning mechanisms of these two divisions.We believe our new dataset should serve as a valuable resource to push forward the exploration of real-world, complex reasoning tasks as the next research focus.Our dataset and code is publicly available 1 . Zhiyu Chen 0002, Charese Smiley, Sameena Shah, William Yang Wang |
EMNLP | 3 |
| 2022 | When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial DomainabstractRaj Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah, Wendi Du, Sudheer Chava, Natraj Raman, Charese Smiley, Jiaao Chen, Diyi Yang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Raj Sanjay Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah, Wendi Du, Sudheer Chava, Natraj Raman, Charese Smiley, Jiaao Chen, Diyi Yang |
EMNLP | 8 |
| 2021 | FinQA: A Dataset of Numerical Reasoning over Financial DataabstractZhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, William Yang Wang. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zhiyu Chen 0002, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao 'Kenneth' Huang, Bryan R. Routledge, William Yang Wang |
EMNLP (1) | 3 |
| 2019 | Building and Querying an Enterprise Knowledge GraphabstractInformation providers are faced with a critical challenge to process, retrieve and present information to their users in order to satisfy their complex information needs, because data has been increasing in an unprecedented manner, coming from diverse sources, and covering a variety of domains in heterogeneous formats. In this paper, we present Thomson Reuters' effort in developing a family of services for building and querying an enterprise knowledge graph in order to address this challenge. We first acquire data from various sources via different approaches. Furthermore, we mine useful information from the data by adopting a variety of techniques, including Named Entity Recognition and Relation Extraction; such mined information is further integrated with existing structured data (e.g., via Entity Linking techniques) in order to obtain relatively comprehensive descriptions of the entities. By modeling the data as an RDF graph model, we enable easy data management and the embedding of rich semantics in our data. Finally, in order to facilitate the querying of this mined and integrated data, i.e., the knowledge graph, we propose TR Discover, a natural language interface that allows users to ask questions of our knowledge graph in their own words; such natural language questions are then translated into executable queries for answer retrieval. We evaluate our services, i.e., named entity recognition, relation extraction, entity linking and natural language interface, on real-world datasets, and demonstrate and discuss their practicability and limitations. Dezhao Song, Frank Schilder, Shai Hertz, Giuseppe Saltini, Charese Smiley, Phani Nivarthi, Oren Hazai, Dudi Landau, Mike Zaharkin, Tom Zielund, Hugo Molina-Salgado, Chris Brew, Dan Bennett |
IEEE Trans. Serv. Comput. | 5 |
| 2018 | The E2E NLG Challenge: A Tale of Two SystemsabstractThis paper presents the two systems we entered into the 2017 E2E NLG Challenge: TemplGen, a templated-based system and SeqGen, a neural network-based system.Through the automatic evaluation, SeqGen achieved competitive results compared to the template-based approach and to other participating systems as well.In addition to the automatic evaluation, in this paper we present and discuss the human evaluation results of our two systems. Charese Smiley, Elnaz Davoodi, Dezhao Song, Frank Schilder |
INLG | 1 |
| 2016 | When to Plummet and When to Soar: Corpus Based Verb Selection for Natural Language GenerationabstractFor data-to-text tasks in Natural Language Generation (NLG), researchers are often faced with choices about the right words to express phenomena seen in the data.One common phenomenon centers around the description of trends between two data points and selecting the appropriate verb to express both the direction and intensity of movement.Our research shows that rather than simply selecting the same verbs again and again, variation and naturalness can be achieved by quantifying writers' patterns of usage around verbs. Charese Smiley, Vassilis Plachouras, Frank Schilder, Hiroko Bretz, Jochen L. Leidner, Dezhao Song |
INLG | 1 |
| 2016 | Interacting with Financial Data using Natural LanguageabstractFinancial and economic data are typically available in the form of tables and comprise mostly of monetary amounts, numeric and other domain-specific fields. They can be very hard to search and they are often made available out of context, or in forms which cannot be integrated with systems where text is required, such as voice-enabled devices. This work presents a novel system that enables both experts in the finance domain and non-expert users to search financial data with both keyword and natural language queries. Our system answers the queries with an automatically generated textual description using Natural Language Generation (NLG). The answers are further enriched with derived information, not explicitly asked in the user query, to provide the context of the answer. The system is designed to be flexible in order to accommodate new use cases without significant development effort, thus allowing fast integration of new datasets. Vassilis Plachouras, Charese Smiley, Hiroko Bretz, Ola Taylor, Jochen L. Leidner, Dezhao Song, Frank Schilder |
SIGIR | 2 |
| 2015 | Natural Language Question Answering and Analytics for Diverse and Interlinked DatasetsabstractPrevious systems for natural language questions over complex linked datasets require the user to enter a complete and well-formed question, and present the answers as raw lists of entities. Using a feature-based grammar with a full formal semantics, we have developed a system that is able to support rich autosuggest, and to deliver dynamically generated analytics for each result that it returns. Dezhao Song, Frank Schilder, Charese Smiley, Chris Brew |
HLT-NAACL | 3 |
| 2015 | TR Discover: A Natural Language Interface for Querying and Analyzing Interlinked Datasets
Dezhao Song, Frank Schilder, Charese Smiley, Chris Brew, Tom Zielund, Hiroko Bretz, Robert Martin, Chris Dale, John Duprey, Johanna Harrison |
ISWC (2) | 3 |