Sameena Shah

dblp:78/3078 · DBLP profile ↗
← Back
22ranked-venue papers in the field
3as first author
6since 2021 · last 2023
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 11 (1 first)Data Mining & Knowledge Discovery · 3 (1 first)Big Data, Cloud & Distributed Data Systems · 3Other / Interdisciplinary · 3Database Systems & Data Management · 1Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
YearPublicationVenuePosition
2023 Robust NLP for Finance (RobustFin)
abstract
Natural language processing (NLP) technologies have been widely applied in business domains such as e-commerce and customer service, but their adoption in the financial sector has been constrained by industry-specific performance standards and regulatory restrictions. This challenge has created new opportunities for core research in related areas. Recent advancements in NLP, such as the advent of large language models, has encouraged adoption in the finance sector. However, compared to other domains, finance has stricter requirements for robustness, explainability, and generalizability. Given this background, we propose to organize the first Robust NLP for Finance (RobustFin) workshop at KDD '23 to encourage the study of and research on robustness and explainability technologies with regard to financial NLP. The goal of the workshop is to extend the applications of NLP in finance, while motivating further research in robust NLP.
Sameena Shah, Xiaodan Zhu 0001, Gerard de Melo, Armineh Nourbakhsh, Xiaomo Liu, Charese Smiley, Zhiyu Chen 0002
KDD1
2023 BizGraphQA: A Dataset for Image-based Inference over Graph-structured Diagrams from Business Domains
abstract
Graph-structured diagrams, such as enterprise ownership charts or management hierarchies, are a challenging medium for deep learning models as they not only require the capacity to model language and spatial relations but also the topology of links between entities and the varying semantics of what those links represent. Devising Question Answering models that automatically process and understand such diagrams have vast applications to many enterprise domains, and can move the state-of-the-art on multimodal document understanding to a new frontier. Curating real-world datasets to train these models can be difficult, due to scarcity and confidentiality of the documents where such diagrams are included. Recently released synthetic datasets are often prone to repetitive structures that can be memorized or tackled using heuristics. In this paper, we present a collection of 10,000 synthetic graphs that faithfully reflect properties of real graphs in four business domains, and are realistically rendered within a PDF document with varying styles and layouts. In addition, we have generated over 130,000 question instances that target complex graphical relationships specific to each domain. We hope this challenge will encourage the development of models capable of robust reasoning about graph structured images, which are ubiquitous in numerous sectors in business and across scientific disciplines.
Petr Babkin, William Watson, Lucas Cecchi, Natraj Raman, Armineh Nourbakhsh, Sameena Shah
SIGIR7
2023 REFinD: Relation Extraction Financial Dataset
abstract
A number of datasets for Relation Extraction (RE) have been created to aide downstream tasks such as information retrieval, semantic search, question answering and textual entailment. However, these datasets fail to capture financial-domain specific challenges since most of these datasets are compiled using general knowledge sources such as Wikipedia, web-based text and news articles, hindering real-life progress and adoption within the financial world. To address this limitation, we propose REFinD, the first large-scale annotated dataset of relations, with ~29K instances and 22 relations amongst 8 types of entity pairs, generated entirely over financial documents. We also provide an empirical evaluation with various state-of-the-art models as benchmarks for the RE task and highlight the challenges posed by our dataset. We observed that various state-of-the-art deep learning models struggle with numeric inference, relational and directional ambiguity. To encourage further research in this direction, REFinD is available at https://www.jpmorgan.com/technology/artificial-intelligence/initiatives/refind-dataset/problem-motivation-outcome.
Simerjot Kaur, Charese Smiley, Akshat Gupta, Joy Prakash Sain, Dongsheng Wang 0005, Suchetha Siddagangappa, Toyin Aguda, Sameena Shah
SIGIR8
2023 Knowledge Discovery from Unstructured Data in Financial Services (KDF) Workshop
abstract
Knowledge discovery from unstructured data, including business documents, web content, and news articles, has been a key AI challenge for the financial services industry. Comprehending these corpora and discovering knowledge from them, which could be textual, tabular, or graphic, are the cornerstone of supporting business decisions in the financial services domain, where information retrieval and content analysis techniques are of fundamental importance. We propose a workshop on knowledge discovery from unstructured data in financial services at SIGIR 2023 to highlight the current and emerging opportunities, invite original research, and prompt success sharing between researchers.
Sameena Shah, Xiaodan Zhu 0001, Wenhu Chen, Manling Li, Armineh Nourbakhsh, Xiaomo Liu, Charese Smiley, Yulong Pei, Akshat Gupta
SIGIR1
2023 DocGraphLM: Documental Graph Language Model for Information Extraction
abstract
Advances in Visually Rich Document Understanding (VrDU) have enabled information extraction and question answering over documents with complex layouts. Two tropes of architectures have emerged-transformer-based models inspired by LLMs, and Graph Neural Networks. In this paper, we introduce DocGraphLM, a novel framework that combines pre-trained language models with graph semantics. To achieve this, we propose 1) a joint encoder architecture to represent documents, and 2) a novel link prediction approach to reconstruct document graphs. DocGraphLM predicts both directions and distances between nodes using a convergent joint loss function that prioritizes neighborhood restoration and downweighs distant node detection. Our experiments on three SotA datasets show consistent improvement on IE and QA tasks with the adoption of graph features. Moreover, we report that adopting the graph features accelerates convergence in the learning process druing training, despite being solely constructed through link prediction.
Dongsheng Wang 0005, Armineh Nourbakhsh, Kang Gu, Sameena Shah
SIGIR5
2022 Structure and Semantics Preserving Document Representations
abstract
Retrieving relevant documents from a corpus is typically based on the semantic similarity between the document content and query text. The inclusion of structural relationship between documents can benefit the retrieval mechanism by addressing semantic gaps. However, incorporating these relationships requires tractable mechanisms that balance structure with semantics and take advantage of the prevalent pre-train/fine-tune paradigm. We propose here a holistic approach to learning document representations by integrating intra-document content with inter-document relations. Our deep metric learning solution analyzes the complex neighborhood structure in the relationship network to efficiently sample similar/dissimilar document pairs and defines a novel quintuplet loss function that simultaneously encourages document pairs that are semantically relevant to be closer and structurally unrelated to be far apart in the representation space. Furthermore, the separation margins between the documents are varied flexibly to encode the heterogeneity in relationship strengths. The model is fully fine-tunable and natively supports query projection during inference. We demonstrate that it outperforms competing methods on multiple datasets for document retrieval tasks.
Natraj Raman, Sameena Shah, Manuela M. Veloso
SIGIR2
2018 An Extensible Event Extraction System With Cross-Media Event Resolution
abstract
The automatic extraction of breaking news events from natural language text is a valuable capability for decision support systems. Traditional systems tend to focus on extracting events from a single media source and often ignore cross-media references. Here, we describe a large-scale automated system for extracting natural disasters and critical events from both newswire text and social media. We outline a comprehensive architecture that can identify, categorize and summarize seven different event types - namely floods, storms, fires, armed conflict, terrorism, infrastructure breakdown, and labour unavailability. The system comprises fourteen modules and is equipped with a novel coreference mechanism, capable of linking events extracted from the two complementary data sources. Additionally, the system is easily extensible to accommodate new event types. Our experimental evaluation demonstrates the effectiveness of the system.
Fabio Petroni, Natraj Raman, Timothy Nugent, Armineh Nourbakhsh, Zarko Panic, Sameena Shah, Jochen L. Leidner
KDD6
2017 Reuters tracer: Toward automated news production using large scale social media data
abstract
To deal with the sheer volume of information and gain competitive advantage, the news industry has started to explore and invest in news automation. In this paper, we present Reuters Tracer, a system that automates end-to-end news production using Twitter data. It is capable of detecting, classifying, annotating, and disseminating news in real time for Reuters journalists without manual intervention. In contrast to other similar systems, Tracer is topic and domain agnostic. It has a bottom-up approach to news detection, and does not rely on a predefined set of sources or subjects. Instead, it identifies emerging conversations from 12+ million tweets per day and selects those that are news-like. Then, it contextualizes each story by adding a summary and a topic to it, estimating its newsworthiness, veracity, novelty, and scope, and geotags it. Designing algorithms to generate news that meets the standards of Reuters journalists in accuracy and timeliness is quite challenging. But Tracer is able to achieve competitive precision, recall, timeliness, and veracity on news detection and delivery. In this paper, we reveal our key algorithm designs and evaluations that helped us achieve this goal, and lessons learned along the way.
Xiaomo Liu, Armineh Nourbakhsh, Quanzhi Li, Sameena Shah, Robert Martin, John Duprey
IEEE BigData4
2017 Real-Time Novel Event Detection from Social Media
abstract
In this paper, we present a new approach for detecting novel events from social media, specially Twitter, at real-time. An event is usually defined by who, what, where and when, and an event tweet usually contains terms corresponding to these aspects. To exploit this information, we propose a method that incorporates simple semantics by splitting the tweet term space into groups of terms that have the meaning of the same type. These groups are called semantic categories (classes) and each reflects one or more event aspects. The semantic classes include named entity, mention, location, hashtag, verb, noun and embedded link. To group tweets talking about the same event into the same cluster, similarity measuring is conducted by calculating class-wise similarity and then aggregating them together. Users of a real-time event detection system are usually only interested in novel (new) events, which are happening now or just happened a short time ago. To fulfill this requirement, a temporal identification module is used to filter out event clusters that are about old stories. The clustering module also computes a novelty score for each event cluster, which reflects how novel the event is, compared to previous events. We evaluated our event detection method using multiple quality metrics and a large-scale event corpus having millions of tweets. The experiment results show that the proposed online event detection method achieves the state-of-the-art performance. Our experiment also shows that the temporal identification module can effectively detect old events.
Quanzhi Li, Armineh Nourbakhsh, Sameena Shah, Xiaomo Liu
ICDE3
2017 Data Sets: Word Embeddings Learned from Tweets and General Data
Quanzhi Li, Sameena Shah, Xiaomo Liu, Armineh Nourbakhsh
ICWSM2
2016 Table classification using both structure and content information: A case study of financial documents
abstract
Tables are significant document components. Table extraction and classification are critical for us to explore, retrieve and mine knowledge encoded in tables. This paper presents a learning based approach for classifying tables based on their content and structural information, with focus on financial document tables. To the best of our knowledge, this is the first study on classifying tables in financial domain, and also the first study of table classification based on its semantics, a more fine-grained level than previous studies. The experimental results show that it can effectively classify financial tables. We also analyzed what features are important and how to generate them. The feature identification and generation approach can potentially apply to other domains.
Quanzhi Li, Sameena Shah
IEEE BigData2
2016 Using paraphrases to improve tweet classification: Comparing WordNet and word embedding approaches
abstract
Two of the major problems in social media message classification are the data sparseness issue and the high degree of lexical variation. Paraphrases, or synonyms, are alternative ways of expressing the same meaning using different lexical variations. In this study, we try to use paraphrases to improve tweet topic classification performance. We explored two approaches to generating paraphrases, WordNet, which is a lexical database grouping English words into sets of synonyms, and word embeddings, which are learned from millions of tweets and billions of words. Our experiment shows that using paraphrases can improve the topic classification task, and the word embedding approach outperforms the WordNet method. To our knowledge, this is the first study exploiting paraphrases for tweet classification.
Quanzhi Li, Sameena Shah, Mohammad M. Ghassemi, Armineh Nourbakhsh, Xiaomo Liu
IEEE BigData2
2016 TweetSift: Tweet Topic Classification Based on Entity Knowledge Base and Topic Enhanced Word Embedding
abstract
Classifying tweets into topic categories is necessary and important for many applications, since tweets are about a variety of topics and users are only interested in certain topical areas. Many tweet classification approaches fail to achieve high accuracy due to data sparseness issue. Tweet, as a special type of short text, in additional to its text, also has other metadata that can be used to enrich its context, such as user name, mention, hashtag and embedded link. In this demonstration, we present TweetSift, an efficient and effective real time tweet topic classifier. TweetSift exploits external tweet-specific entity knowledge to provide more topical context for a tweet, and integrates them with topic enhanced word embeddings for topic classification. The demonstration will show how TweetSift works and how it is incorporated with our social media event detection system.
Quanzhi Li, Sameena Shah, Xiaomo Liu, Armineh Nourbakhsh
CIKM2
2016 Hashtag Recommendation Based on Topic Enhanced Embedding, Tweet Entity Data and Learning to Rank
abstract
In this paper, we present a new approach of recommending hashtags for tweets. It uses Learning to Rank algorithm to incorporate features built from topic enhanced word embeddings, tweet entity data, hashtag frequency, hashtag temporal data and tweet URL domain information. The experiments using millions of tweets and hashtags show that the proposed approach outperforms the three baseline methods -- the LDA topic, the tf.idf based and the general word embedding approaches.
Quanzhi Li, Sameena Shah, Armineh Nourbakhsh, Xiaomo Liu
CIKM2
2016 Reuters Tracer: A Large Scale System of Detecting & Verifying Real-Time News Events from Twitter
abstract
News professionals are facing the challenge of discovering news from more diverse and unreliable information in the age of social media. More and more news events break on social media first and are picked up by news media subsequently. The recent Brussels attack is such an example. At Reuters, a global news agency, we have observed the necessity of providing a more effective tool that can help our journalists to quickly discover news on social media, verify them and then inform the public.
Xiaomo Liu, Quanzhi Li, Armineh Nourbakhsh, Merine Thomas, Kajsa Anderson, Russ Kociuba, Mark Vedder, Steven Pomerville, Ramdev Wudali, Robert Martin, John Duprey, Arun Vachher, William Keenan, Sameena Shah
CIKM15
2016 Perceived, Projected, and True Investment Expertise: Not All Experts Provide Expert Recommendations
abstract
Social networks enable knowledge sharing that inevitably begs the question of expertise analysis. Many online profiles claim expertise, but possessing true expertise is rare. We characterize expertise as projected expertise (claims of a person), perceived expertise (how the crowd perceives the individual) and true expertise (factual). StockTwits, an investor-focused microblogging platform, allows us to study all three aspects simultaneously. We analyze more than 18 million tweets spanning 1700 days. The large time scale allows us to also analyze expertise and its categories as they evolve over time, which is the first study of its kind on StockTwits. We propose a method to capture perceived expertise by how significantly a user's follower network grows and how often the user is brought up in conversations. We also quantify actual, market-based, true expertise based on the user's trade and investment recommendations. Finally we provide an analysis bringing out the differences between how users project themselves, how the crowd perceives them, and how they are actually performing on the market. Our results show that users who project themselves as experts are ones that talk the most and provide the least recommendation-to-tweet ratio (that is, most of their conversations are mundane). The recommendations from users who project novice expertise slightly outperform (≈5%) the overall stock market. On the other hand, the trade recommendations from self-proclaimed experts yield 80% less than those of intermediate traders. Interestingly, users who are perceived as experts by others, as measured by centrality measurements, resulted in net negative returns after a four year trading period. Our study also looks at the evolution of expertise, and begins to understand why and what makes users change the way they project their own expertise. For this topic, however, this paper introduces more questions than it answers, which will serve as the basis for future studies.
Amit Shavit, Sameena Shah
DSAA2
2016 User Behaviors in Newsworthy Rumors: A Case Study of Twitter
Quanzhi Li, Xiaomo Liu, Armineh Nourbakhsh, Sameena Shah
ICWSM5
2016 Tweet Sentiment Analysis by Incorporating Sentiment-Specific Word Embedding and Weighted Text Features
abstract
Previous studies have used many manually identified features and word embeddings for tweet sentiment classification. In this paper, we propose a new approach, which incorporates sentiment-specific word embeddings (SSWE) and a weighted text feature model (WTFM). WTFM produces features based on text negation, tf.idf weighting scheme, and a Rocchio text classification method. Compared to other tweet sentiment feature generation approaches, WTFM is easy to build, simple, yet effective. Experiments show that the proposed approach outperforms the two state-of-the-art tweet sentiment classification methods, SSWE and National Research Council Canada's (NRC) model.
Quanzhi Li, Sameena Shah, Armineh Nourbakhsh, Xiaomo Liu
WI2
2016 Tweet Topic Classification Using Distributed Language Representations
abstract
Many classification tasks on short text, such as tweet, fail to achieve high accuracy due to data sparseness. One approach to solving this problem is to enrich the context of data by using external data sources, or distributed language representations trained on huge amount of data. In this paper, we present several tweet topic classification methods by exploiting different types of data: tweet text, tweet text plus entity knowledge base, word embeddings derived from tweet text, distributed representations of tweets, and topical word embeddings. The word embedding, topical word embedding and sentence representation models are generated from billions of words from tweets without supervision. To the best of our knowledge, this is the first study of applying distributed language representations to tweet topic classification task.
Quanzhi Li, Sameena Shah, Xiaomo Liu, Armineh Nourbakhsh
WI2
2015 Real-time Rumor Debunking on Twitter
abstract
In this paper, we propose the first real time rumor debunking algorithm for Twitter. We use cues from 'wisdom of the crowds', that is, the aggregate 'common sense' and investigative journalism of Twitter users. We concentrate on identification of a rumor as an event that may comprise of one or more conflicting microblogs. We continue monitoring the rumor event and generate real time updates dynamically based on any additional information received. We show using real streaming data that it is possible, using our approach, to debunk rumors accurately and efficiently, often much faster than manual verification by professionals.
Xiaomo Liu, Armineh Nourbakhsh, Quanzhi Li, Sameena Shah
CIKM5
2013 Stock Prediction Using Event-Based Sentiment Analysis
abstract
We propose a novel approach to label social media text using significant stock market events (big losses or gains). Since stock events are easily quantifiable using returns from indices or individual stocks, they provide meaningful and automated labels. We extract significant stock movements and collect appropriate pre, post and contemporaneous text from social media sources (for example, tweets from twitter). Subsequently, we assign the respective label (positive or negative) for each tweet. We train a model on this collected set and make predictions for labels of future tweets. We aggregate the net sentiment per each day (amongst other metrics) and show that it holds significant predictive power for subsequent stock market movement. We create successful trading strategies based on this system and find significant returns over other baseline methods.
Masoud Makrehchi, Sameena Shah, Wenhui Liao
Web Intelligence2
2013 Convergence of the dynamic load balancing problem to Nash equilibrium using distributed local interactions
Sameena Shah, Ravi Kothari
Inf. Sci.1