EDBT 2026 Demo / reviewers in the wild / expert
Yi Yang 0042
dblp:33/4854-42
· DBLP profile ↗
39ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0001-8863-112XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 4 first-author · 14 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Theory of computation · 3 · 1 first-author · 3 since 2021Computer networks · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic duality in hypergraphs: Uncertainty-aware bipolar evidence aggregation for temporal knowledge graph reasoning
Bin Chen 0030, Yi Yang 0042, Zhangtao Cheng, Xueting Liu 0005, Yicheng Xin, Kunpeng Zhang 0001, Fan Zhou 0002 |
Expert Syst. Appl. | 2 |
| 2026 | Divide and Contrast: A Text-Based Method for Firm Market Risk PredictionabstractForecasting the market risk for publicly traded companies is a critical task for market participants. Financial economics research demonstrates that the textual information contained in corporate disclosures, such as earnings conference call transcripts, can effectively predict a firm’s future risk. This finding has inspired a growing body of research focused specifically on transcript-based approaches to risk forecasting. However, earnings transcripts are typically long documents with thousands of words. Prior transcript-based risk forecasting studies that represent the entire transcript as one text sequence often fail to capture risk-relevant information and fall short in risk forecasting. In this work, we propose a novel divide-and-contrast machine learning method for predicting risks from earnings conference call transcripts. We exploit the semistructured nature of an earnings transcript and decompose it into several semantically coherent conversation units, ranging from the finest grained question–answer pair level to the coarsest grained transcript level. We then propose contrastive learning objectives as an auxiliary task to the risk forecasting objective, facilitating the learning of risk-relevant information from the earnings transcripts. We conduct experiments on a data set of U.S. market earnings call transcripts. The experimental results show that our proposed divide-and-contrast method substantially outperforms state-of-the-art methods by significantly reducing errors in risk forecasting. This paper sheds light on extracting informative insights from lengthy financial documents to support informed decision making. History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0195 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0195 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Yi Yang 0042, Defu Lian, Kunpeng Zhang 0001 |
INFORMS J. Comput. | 2 |
| 2025 | Evaluating and Aligning Human Economic Risk Preferences in LLMsabstractLarge Language Models (LLMs) are increasingly used in decision-making scenarios that involve risk assessment, yet their alignment with human economic rationality remains unclear.In this study, we investigate whether LLMs exhibit risk preferences consistent with human expectations across different personas.Specifically, we propose an evaluation metric called Risk Disparity Score (RDS) and assess whether LLM-generated responses reflect appropriate levels of risk aversion or riskseeking behavior based on individual's persona.Our results reveal that while LLMs make reasonable decisions in simplified, personalized risk contexts, their performance declines in more complex economic decision-making tasks.To address this, we test whether current state-of-art alignment methods such as Direct Preference Optimization(DPO) and In Context Learning(ICL) can enhance LLM adherence to persona-specific risk preferences.We find DPO can improve the economic rationality of LLMs in loss-related parameters, offering a step toward more human-aligned AI decision-making. Yixuan Tang 0001, Yi Yang 0042, Kar Yan Tam |
EMNLP | 3 |
| 2025 | FinMTEB: Finance Massive Text Embedding BenchmarkabstractThe efficacy of text embedding models in representing and retrieving information is crucial for many NLP applications, with performance significantly advanced by Large Language Models (LLMs).Despite this progress, existing benchmarks predominantly use generalpurpose datasets, inadequately addressing the nuanced requirements of specialized domains like finance.To bridge this gap, we introduce the Finance Massive Text Embedding Benchmark (FinMTEB), a comprehensive evaluation suite specifically designed for the financial domain.FinMTEB encompasses 64 datasets across 7 task types, including classification, clustering, retrieval, pair classification, reranking, summarization, and semantic textual similarity (STS) in English and Chinese.Alongside this benchmark, we introduce Fin-E5, a state-of-the-art finance-adapted embedding model, ranking first on FinMTEB.Fin-E5 is developed by fine-tuning e5-Mistral-7B-Instruct on a novel persona-based synthetic dataset tailored for diverse financial embedding tasks.Evaluating 15 prominent embedding models on FinMTEB, we derive three key findings: (1) domain-specific models, including our Fin-E5, significantly outperform general-purpose models; (2) performance on general benchmarks is a poor predictor of success on financial tasks; and (3) surprisingly, traditional Bag-of-Words (BoW) models surpass dense embedding models on financial STS tasks.This work provides a robust benchmark for financial NLP and offers actionable insights for developing future domain-adapted embedding solutions.Both FinMTEB and Fin-E5 will be open-sourced for the research community. Yixuan Tang 0001, Yi Yang 0042 |
EMNLP | 2 |
| 2025 | TDDBench: A Benchmark for Training data detectionabstractTraining Data Detection (TDD) is a task aimed at determining whether a specific data instance is used to train a machine learning model. In the computer security literature, TDD is also referred to as Membership Inference Attack (MIA). Given its potential to assess the risks of training data breaches, ensure copyright authentication, and verify model unlearning, TDD has garnered significant attention in recent years, leading to the development of numerous methods. Despite these advancements, there is no comprehensive benchmark to thoroughly evaluate the effectiveness of TDD methods.
In this work, we introduce TDDBench, which consists of 13 datasets spanning three data modalities: image, tabular, and text. We benchmark 21 different TDD methods across four detection paradigms and evaluate their performance from five perspectives: average detection performance, best detection performance, memory consumption, and computational efficiency in both time and memory. With TDDBench, researchers can identify bottlenecks and areas for improvement in TDD algorithms, while practitioners can make informed trade-offs between effectiveness and efficiency when selecting TDD algorithms for specific use cases. Our extensive experiments also reveal the generally unsatisfactory performance of TDD algorithms across different datasets. To enhance accessibility and reproducibility, we open-source TDDBench for the research community at https://github.com/zzh9568/TDDBench. Zhihao Zhu 0002, Yi Yang 0042, Defu Lian |
ICLR | 2 |
| 2025 | Adversarial Mixup UnlearningabstractMachine unlearning is a critical area of research aimed at safeguarding data privacy by enabling the removal of sensitive information from machine learning models. One unique challenge in this field is catastrophic unlearning, where erasing specific data from a well-trained model unintentionally removes essential knowledge, causing the model to deviate significantly from a retrained one. To address this, we introduce a novel approach that regularizes the unlearning process by utilizing synthesized mixup samples, which simulate the data susceptible to catastrophic effects. At the core of our approach is a generator-unlearner framework, MixUnlearn, where a generator adversarially produces challenging mixup examples, and the unlearner effectively forgets target information based on these synthesized data. Specifically, we first introduce a novel contrastive objective to train the generator in an adversarial direction: generating examples that prompt the unlearner to reveal information that should be forgotten, while losing essential knowledge. Then the unlearner, guided by two other contrastive loss terms, processes the synthesized and real data jointly to ensure accurate unlearning without losing critical knowledge, overcoming catastrophic effects. Extensive evaluations across benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches, offering a robust solution to machine unlearning. This work not only deepens understanding of unlearning mechanisms but also lays the foundation for effective machine unlearning with mixup augmentation. Zhuoyi Peng, Yixuan Tang 0001, Yi Yang 0042 |
ICLR | 3 |
| 2025 | Efficient Multi-Expert Tabular Language Model for BankingabstractPre-training large Tabular Language Models (TaLMs) on tabular data has shown effectiveness for table understanding tasks. However, training proprietary large TaLMs on a company's private databases requires substantial computational resources. This paper presents an efficient multi-expert TaLM architecture and training method tailored for multi-domain databases and modest infrastructure. This architecture leverages a divide-and-conquer pretraining approach and a sparsely activated fine-tuning paradigm to reduce computation. Using this architecture, we pre-train and fine-tune a TaLM with 10 billion parameters on a banking database under simple computational infrastructures. We apply our TaLM to support various important banking applications, including risk assessment, information prediction, and profit assessment. Compared with previous baselines, our model achieves +29.3% in [email protected]% on risk assessment and +16.5% in accuracy on information prediction, showing great effectiveness and profitability of our model. This model is successfully deployed in WeBank and now supports many real business scenarios. Yue Guo 0009, Vincent Wenchen Zheng, Yi Yang 0042 |
KDD (1) | 5 |
| 2025 | Hierarchical Deep Document ModelabstractTopic modeling is a commonly used text analysis tool for discovering latent topics in a text corpus. However, while topics in a text corpus often exhibit a hierarchical structure (e.g., cellphone is a sub-topic of electronics), most topic modeling methods assume a flat topic structure that ignores the hierarchical dependency among topics, or utilize a predefined topic hierarchy. In this work, we present a novel Hierarchical Deep Document Model (HDDM) to learn topic hierarchies using a variational autoencoder framework. We propose a novel objective function, sum of log likelihood, instead of the widely used evidence lower bound, to facilitate the learning of hierarchical latent topic structure. The proposed objective function can directly model and optimize the hierarchical topic-word distributions at all topic levels. We conduct experiments on four real-world text datasets to evaluate the topic modeling capability of the proposed HDDM method compared to state-of-the-art hierarchical topic modeling benchmarks. Experimental results show that HDDM achieves considerable improvement over benchmarks and is capable of learning meaningful topics and topic hierarchies. To further demonstrate the practical utility of HDDM, we apply it to a real-world medical notes dataset for clinical prediction. Experimental results show that HDDM can better summarize topics in medical notes, resulting in more accurate clinical predictions. Yi Yang 0042, John Lalor, Ahmed Abbasi, Daniel Dajun Zeng |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Self-Explainable Next POI RecommendationabstractPoint-of-Interest (POI) recommendation involves predicting users' next preferred POI and is becoming increasingly significant in location-based social networks. However, users are often reluctant to trust recommended results due to the lack of transparency in these systems. While recent work on explaining recommender systems has gained attention, prevailing methods only provide post-hoc explanations based on results or rudimentary explanations according to attention scores. Such limitations hinder reliability and applicability in risk-sensitive scenarios. Inspired by the information theory, we propose a self-explainable framework with an ante-hoc view called \M~for next POI recommendation aimed at overcoming these limitations. Specifically, we endow self-explainability to POI recommender systems through compact representation learning using a variational information bottleneck approach. The learned representation further improves accuracy by reducing redundancy behind massive spatial-temporal trajectories, which, in turn, boosts the recommendation performance. Experiments on three real-world datasets show significant improvements in both model explainability and recommendation performance. Yi Yang 0042, Qiang Gao 0003, Ting Zhong, Yong Wang 0046, Fan Zhou 0002 |
SIGIR | 2 |
| 2024 | Federated Meta Embedding Concept Stock RecommendationabstractMining relevant stocks given a trending topic/concept in capital markets is an application with significant economic and societal impacts. Previous concept stock recommendation system mines concept stocks only from public social media like financial news. On stock forums, investors discuss emerging concepts and stocks by using forum comments, which are unneglectable resources to capture trending concept stocks accurately and timely. However, the comment data from a single forum is insufficient to build a high-quality recommendation system. The forums are data silos protected by privacy regulations, and their comments are still underutilized. In this paper, we propose a federated concept stock recommendation baseline and an optimized method that both leverage the private forum comments and public social media without compromising privacy regulations. Our baseline,i.e., Federated Meta Embedding (FedME), is built upon the federated learning framework and learns a concept-stock embedding jointly from private and public data. Our optimized method, Federated Graph Meta Embedding (FedGME), improves FedME by using a graph to combine two sources of embeddings and additional human experts' concept-stock knowledge. Empirically, the experiments on two concept stock datasets show that FedME and FedGME substantially improve the performance of recommendation. Our methods provide practical guidance on privacy-preserving FinTech applications. Zhuoyi Peng, Yi Yang 0042, Liu Yang 0008, Kai Chen 0005 |
IEEE Trans. Big Data | 2 |
| 2024 | Should Fairness be a Metric or a Model? A Model-based Framework for Assessing Bias in Machine Learning PipelinesabstractFairness measurement is crucial for assessing algorithmic bias in various types of machine learning (ML) models, including ones used for search relevance, recommendation, personalization, talent analytics, and natural language processing. However, the fairness measurement paradigm is currently dominated by fairness metrics that examine disparities in allocation and/or prediction error as univariate key performance indicators (KPIs) for a protected attribute or group. Although important and effective in assessing ML bias in certain contexts such as recidivism, existing metrics don’t work well in many real-world applications of ML characterized by imperfect models applied to an array of instances encompassing a multivariate mixture of protected attributes, that are part of a broader process pipeline. Consequently, the upstream representational harm quantified by existing metrics based on how the model represents protected groups doesn’t necessarily relate to allocational harm in the application of such models in downstream policy/decision contexts. We propose FAIR-Frame, a model-based framework for parsimoniously modeling fairness across multiple protected attributes in regard to the representational and allocational harm associated with the upstream design/development and downstream usage of ML models. We evaluate the efficacy of our proposed framework on two testbeds pertaining to text classification using pretrained language models. The upstream testbeds encompass over fifty thousand documents associated with twenty-eight thousand users, seven protected attributes and five different classification tasks. The downstream testbeds span three policy outcomes and over 5.41 million total observations. Results in comparison with several existing metrics show that the upstream representational harm measures produced by FAIR-Frame and other metrics are significantly different from one another, and that FAIR-Frame’s representational fairness measures have the highest percentage alignment and lowest error with allocational harm observed in downstream applications. Our findings have important implications for various ML contexts, including information retrieval, user modeling, digital platforms, and text classification, where responsible and trustworthy AI is becoming an imperative. John Lalor, Ahmed Abbasi, Kezia Oketch, Yi Yang 0042, Nicole Forsgren |
ACM Trans. Inf. Syst. | 4 |
| 2023 | Exploring Hypergraph of Earnings Call for Risk Prediction (Student Abstract)abstractIn financial economics, studies have shown that the textual content in the earnings conference call transcript has predictive power for a firm's future risk. However, the conference call transcript is very long and contains diverse non-relevant content, which poses challenges for the text-based risk forecast. This study investigates the structural dependency within a conference call transcript by explicitly modeling the dialogue between managers and analysts. Specifically, we utilize TextRank to extract information and exploit the semantic correlation within a discussion using hypergraph learning. This novel design can improve the transcript representation performance and reduce the risk of forecast errors. Experimental results on a large-scale dataset show that our approach can significantly improve prediction performance compared to state-of-the-art text-based models. Wenxin Tai, Fan Zhou 0002, Yi Yang 0042 |
AAAI | 4 |
| 2023 | Debiasing Intrinsic Bias and Application Bias Jointly via Invariant Risk Minimization (Student Abstract)abstractDemographic biases and social stereotypes are common in pretrained language models (PLMs), while the fine-tuning in downstream applications can also produce new biases or amplify the impact of the original biases. Existing works separate the debiasing from the fine-tuning procedure, which results in a gap between intrinsic bias and application bias. In this work, we propose a debiasing framework CauDebias to eliminate both biases, which directly combines debiasing with fine-tuning and can be applied for any PLMs in downstream tasks. We distinguish the bias-relevant (non-causal factors) and label-relevant (causal factors) parts in sentences from a causal invariant perspective. Specifically, we perform intervention on non-causal factors in different demographic groups, and then devise an invariant risk minimization loss to trade-off performance between bias mitigation and task accuracy. Experimental results on three downstream tasks show that our CauDebias can remarkably reduce biases in PLMs while minimizing the impact on downstream tasks. Yuzhou Mao, Liu Yu 0001, Yi Yang 0042, Fan Zhou 0002, Ting Zhong |
AAAI | 3 |
| 2023 | Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant LearningabstractDemographic biases and social stereotypes are common in pretrained language models (PLMs), and a burgeoning body of literature focuses on removing the unwanted stereotypical associations from PLMs.However, when fine-tuning these bias-mitigated PLMs in downstream natural language processing (NLP) applications, such as sentiment classification, the unwanted stereotypical associations resurface or even get amplified.Since pretrain&fine-tune is a major paradigm in NLP applications, separating the debiasing procedure of PLMs from fine-tuning would eventually harm the actual downstream utility.In this paper, we propose a unified debiasing framework Causal-Debias to remove unwanted stereotypical associations in PLMs during fine-tuning.Specifically, Causal-Debias mitigates bias from a causal invariant perspective by leveraging the specific downstream task to identify bias-relevant and labelrelevant factors.We propose that bias-relevant factors are non-causal as they should have little impact on downstream tasks, while labelrelevant factors are causal.We perform interventions on non-causal factors in different demographic groups and design an invariant risk minimization loss to mitigate bias while maintaining task performance.Experimental results on three downstream tasks show that our proposed method can remarkably reduce unwanted stereotypical associations after PLMs are finetuned, while simultaneously minimizing the impact on PLMs and downstream applications. *Corresponding author unwanted stereotypical associations in PLMs.For example, some works (Zmigrod et al., 2019) pretrain a language model using original and counterfactual corpus in order to cancel-out biased associations, some works (Liang et al., 2020) focus on debiasing post-hoc sentence representations, and others (Guo et al., 2022;Cheng et al., 2021) design bias-equalizing objectives to fine-tune PLM's parameters. Fan Zhou 0002, Yuzhou Mao, Liu Yu 0001, Yi Yang 0042, Ting Zhong |
ACL (1) | 4 |
| 2023 | Predict the Future from the Past? On the Temporal Data Distribution Shift in Financial Sentiment ClassificationsabstractTemporal data distribution shift is prevalent in the financial text.How can a financial sentiment analysis system be trained in a volatile market environment that can accurately infer sentiment and be robust to temporal data distribution shifts?In this paper, we conduct an empirical study on the financial sentiment analysis system under temporal data distribution shifts using a real-world financial social media dataset that spans three years.We find that the fine-tuned models suffer from general performance degradation in the presence of temporal distribution shifts.Furthermore, motivated by the unique temporal nature of the financial text, we propose a novel method that combines out-of-distribution detection with time series modeling for temporal financial sentiment analysis.Experimental results show that the proposed method enhances the model's capability to adapt to evolving temporal shifts in a volatile financial market. Yue Guo 0009, Yi Yang 0042 |
EMNLP | 3 |
| 2023 | FinEntity: Entity-level Sentiment Classification for Financial TextsabstractIn the financial domain, conducting entity-level sentiment analysis is crucial for accurately assessing the sentiment directed toward a specific financial entity. To our knowledge, no publicly available dataset currently exists for this purpose. In this work, we introduce an entity-level sentiment classification dataset, called FinEntity, that annotates financial entity spans and their sentiment (positive, neutral, and negative) in financial news. We document the dataset construction process in the paper. Additionally, we benchmark several pre-trained models (BERT, FinBERT, etc.) and ChatGPT on entity-level sentiment classification. In a case study, we demonstrate the practical utility of using FinEntity in monitoring cryptocurrency markets. The data and code of FinEntity is available at https://github.com/yixuantt/FinEntity. ©2023 Association for Computational Linguistics. Yixuan Tang 0001, Yi Yang 0042, Allen Huang, Andy Tam, Justin Z. Tang |
EMNLP | 2 |
| 2023 | Deep Cross-Attention Network for Crowdfunding Success PredictionabstractCrowdfunding creates opportunities for entrepre- neurs. It allows startup companies to reach a large audience for fundraising and bring their creative ideas to life. In this work, we are concerned with crowdfunding project success prediction problem,i.e., to predict whether a project will successfully reach its funding goal by using its project profiles. This is important for startup companies to refine their project profiles and achieve their goals. Crowdfunding project success prediction is a typical classification problem but with a few critical challenges. On the one hand, with only coarse-grained project status as weak supervision, it is hard for a deep learning network to learn the relationship between project profiles and explain why it makes this prediction. On the other hand, on the project homepage, there are various modalities of description, including metadata, textual description, images, and videos. Among those, videos play an important role in the success of a crowdfunding project, however, were ignored in previous works, due to the difficulty in extracting useful semantic and authentic information from videos, especially for the crowdfunding project where information in different modalities are unaligned. To this end, we propose a novel framework called Deep Cross-Attention Network to learn and fuse information from introduction videos and textual descriptions of project profiles. More specifically, we develop a cross-attention block to align and represent mismatched textual description and untrimmed introduction videos and fuse the information from these two modalities, which effectively remedies the lack of supervised information caused by project status as weak supervision. More importantly, with our cross-attention mechanism, the model is able to interpret how it makes such predictions and show which keywords and keyframes it depends on. We conduct extensive experiments on two crowdfunding datasets (collected from Kickstarter and Indiegogo) and show that our method achieves superior performance over existing state-of-the-art baselines. Yi Yang 0042, Wen Li 0001, Defu Lian, Lixin Duan |
IEEE Trans. Multim. | 2 |
| 2022 | Auto-Debias: Debiasing Masked Language Models with Automated Biased PromptsabstractHuman-like biases and undesired social stereotypes exist in large pretrained language models.Given the wide adoption of these models in real-world applications, mitigating such biases has become an emerging and important task.In this paper, we propose an automatic method to mitigate the biases in pretrained language models.Different from previous debiasing work that uses external corpora to finetune the pretrained models, we instead directly probe the biases encoded in pretrained models through prompts.Specifically, we propose a variant of the beam search method to automatically search for biased prompts such that the cloze-style completions are the most different with respect to different demographic groups.Given the identified biased prompts, we then propose a distribution alignment loss to mitigate the biases.Experiment results on standard datasets and metrics show that our proposed Auto-Debias approach can significantly reduce biases, including gender and racial bias, in pretrained language models such as BERT, RoBERTa and ALBERT.Moreover, the improvement in fairness does not decrease the language models' understanding abilities, as shown using the GLUE benchmark. Yue Guo 0009, Yi Yang 0042, Ahmed Abbasi |
ACL (1) | 2 |
| 2022 | Benchmarking Intersectional Biases in NLPabstractJohn Lalor, Yi Yang, Kendall Smith, Nicole Forsgren, Ahmed Abbasi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. John Lalor, Yi Yang 0042, Kendall Smith, Nicole Forsgren, Ahmed Abbasi |
NAACL-HLT | 2 |
| 2022 | Analyzing Firm Reports for Volatility Prediction: A Knowledge-Driven Text-Embedding ApproachabstractPredicting stock return volatility is the key to investment and risk management. Traditional volatility-forecasting methods primarily rely on stochastic models. More recently, many machine-learning approaches, particularly text-mining techniques, have been implemented to predict stock return volatility, thus taking advantage of the availability of large amounts of unstructured data such as firm financial reports. Most existing studies develop simple but effective models to analyze text, such as dictionary-based matching algorithms that use a set of manually constructed keywords. However, the latent and deep semantics encoded in text are usually neglected. In this study, we build on recent progress in representation learning and propose a novel word-embedding method that incorporates external knowledge from a well-known finance-domain lexicon (the Loughran and McDonald (2011) word list), which helps us learn semantic relationships among words in firm reports for better volatility prediction. Using over 10 years of annual reports from Russell 3000 firms, we empirically show that, compared with cutting-edge benchmarks, our proposed method achieves significant improvement in terms of prediction error, for example, a 28.4% reduction on average. We also discuss the practical and methodological implications of our findings. Our financial-specific word-embedding program is available as open-source information so that researchers can use it to analyze financial reports and assess financial risks. Summary of Contribution: Predicting stock return volatility is the key to investment and risk management. Traditional volatility-forecasting methods primarily rely on stochastic models. More recently, many machine-learning, especially text-mining, techniques have been developed to predict stock return volatility given the availability of a large amount of unstructured data, such as firm annual reports. Most existing research develops simple but effective approaches, for example, manually constructing a set of keywords to analyze texts. However, the latent and deep semantics encoded in texts are usually ignored. In this research, we build on recent progress in representation learning and propose a novel word-embedding method that incorporates external knowledge from the finance-domain lexicon of Loughran and McDonald (2011), which helps us learn the semantic relationships among words in firm annual reports for better volatility prediction. In this study, we make the following contributions. First, methodologically, we are among the first to incorporate finance-specific lexicon into representation learning for stock volatility prediction. We propose a novel knowledge-driven text-embedding model that is trained on a large amount of unstructured textual data to learn high quality word embedding. Our proposed approach is effective in predicting stock return volatility, and the approach can potentially have broader applications. Second, substantively, we empirically show that the domain lexicon enhanced text representation learning can indeed significantly improve the performance, compared with bag-of-words models and generic word embedding for volatility prediction. Domain knowledge combined with text learning plays a critical enabling role in understanding financial reports. Third, our method adds on to existing literature on designing financial information systems by incorporating ontology knowledge, common-sense knowledge, and general prior knowledge. Yi Yang 0042, Kunpeng Zhang 0001, Yangyang Fan |
INFORMS J. Comput. | 1 |
| 2021 | Constructing a Psychometric Testbed for Fair Natural Language ProcessingabstractPsychometric measures of ability, attitudes, perceptions, and beliefs are crucial for understanding user behavior in various contexts including health, security, e-commerce, and finance.Traditionally, psychometric dimensions have been measured and collected using survey-based methods.Inferring such constructs from user-generated text could allow timely, unobtrusive collection and analysis.In this work we construct a corpus for psychometric natural language processing (NLP) related to important dimensions such as trust, anxiety, numeracy, and literacy, in the health domain.We discuss our multi-step process to align user text with their survey-based response items and provide an overview of the resulting testbed, which encompasses surveybased psychometric measures and accompanying user-generated text from 8,502 respondents.Our testbed also encompasses selfreported demographic information, including race, sex, age, income, and education, allowing for measuring bias and benchmarking fairness of text classification methods.We report preliminary results on use of the text to predict/categorize users' survey response labels and on the fairness of these models.We also discuss the important implications of our work and resulting testbed for future NLP research on psychometrics and fairness. Ahmed Abbasi, David G. Dobolyi, John Lalor, Richard G. Netemeyer, Kendall Smith, Yi Yang 0042 |
EMNLP (1) | 6 |
| 2021 | Rumor Detection on Social Media with Event AugmentationsabstractWith the rapid growth of digital data on the Internet, rumor detection on social media has been vital. Existing deep learning-based methods have achieved promising results due to their ability to learn high-level representations of rumors. Despite the success, we argue that these approaches require large reliable labeled data to train, which is time-consuming and data-inefficient. To address this challenge, we present a new solution, Rumor Detection on social media with Event Augmentations (RDEA), which innovatively integrates three augmentation strategies by modifying both reply attributes and event structure to extract meaningful rumor propagation patterns and to learn intrinsic representations of user engagement. Moreover, we introduce contrastive self-supervised learning for the efficient implementation of event augmentations and alleviate limited data issues. Extensive experiments conducted on two public datasets demonstrate that RDEA achieves state-of-the-art performance over existing baselines. Besides, we empirically show the robustness of RDEA when labeled data are limited. Zhenyu He 0008, Ce Li 0003, Fan Zhou 0002, Yi Yang 0042 |
SIGIR | 4 |
| 2021 | Unifying Online and Offline Preference for Social Link PredictionabstractRecent advances in network representation learning have enabled significant improvement in the link prediction task, which is at the core of many downstream applications. As an increasing amount of mobility data become available because of the development of location-based technologies, we argue that this resourceful mobility data can be used to improve link prediction tasks. In this paper, we propose a novel link prediction framework that utilizes user offline check-in behavior combined with user online social relations. We model user offline location preference via a probabilistic factor model and represent user social relations using neural network representation learning. To capture the interrelationship of these two sources, we develop an anchor link method to align these two different user latent representations. Furthermore, we employ locality-sensitive hashing to project the aggregated user representation into a binary matrix, which not only preserves the data structure but also improves the efficiency of convolutional network learning. By comparing with several baseline methods that solely rely on social networks or mobility data, we show that our unified approach significantly improves the link prediction performance. Summary of Contribution: This paper proposes a novel framework that utilizes both user offline and online behavior for social link prediction by developing several machine learning algorithms, such as probabilistic factor model, neural network embedding, anchor link model, and locality-sensitive hashing. The scope and mission has the following aspects: (1) We develop a data and knowledge modeling approach that demonstrates significant performance improvement. (2) Our method can efficiently manage large-scale data. (3) We conduct rigorous experiments on real-world data sets and empirically show the effectiveness and the efficiency of our proposed method. Overall, our paper can contribute to the advancement of social link prediction, which can spur many downstream applications in information systems and computer science. Fan Zhou 0002, Kunpeng Zhang 0001, Bangying Wu, Yi Yang 0042, Harry J. Wang |
INFORMS J. Comput. | 4 |
| 2020 | Interpreting Twitter User GeolocationabstractIdentifying user geolocation in online social networks is an essential task in many locationbased applications.Existing methods rely on the similarity of text and network structure, however, they suffer from a lack of interpretability on the corresponding results, which is crucial for understanding model behavior.In this work, we adopt influence functions to interpret the behavior of GNN-based models by identifying the importance of training users when predicting the locations of the testing users.This methodology helps with providing meaningful explanations on prediction results.Furthermore, it also initiates an attempt to uncover the so-called "black-box" GNN-based models by investigating the effect of individual nodes. Ting Zhong, Tianliang Wang, Fan Zhou 0002, Goce Trajcevski, Kunpeng Zhang 0001, Yi Yang 0042 |
ACL | 6 |
| 2020 | Interpretable Operational Risk Classification with Semi-Supervised Variational AutoencoderabstractOperational risk management is one of the biggest challenges nowadays faced by financial institutions.There are several major challenges of building a text classification system for automatic operational risk prediction, including imbalanced labeled/unlabeled data and lacking interpretability.To tackle these challenges, we present a semi-supervised text classification framework that integrates multi-head attention mechanism with Semisupervised variational inference for Operational Risk Classification (SemiORC).We empirically evaluate the framework on a realworld dataset.The results demonstrate that our method can better utilize unlabeled data and learn visually interpretable document representations.SemiORC also outperforms other baseline methods on operational risk classification. Fan Zhou 0002, Yi Yang 0042 |
ACL | 3 |
| 2020 | Neural Topic Model with Attention for Supervised LearningabstractTopic modeling utilizing neural variational inference has shown promising results recently. Unlike traditional Bayesian topic models, neural topic models use deep neural network to approximate the intractable marginal distribution and thus gain strong generalisation ability. However, neural topic models are unsupervised model. Directly using the document-specific topic proportions in downstream prediction tasks could lead to sub-optimal performance. This paper presents Topic Attention Model (TAM), a supervised neural topic model that integrates an attention recurrent neural network (RNN) model. We design a novel way to utilize document-specific topic proportions and global topic vectors learned from neural topic model in the attention mechanism. We also develop backpropagation inference method that allows for joint model optimisation. Experimental results on three public datasets show that TAM not only significantly improves supervised learning tasks, including classification and regression, but also achieves lower perplexity for the document modeling. Xinyi Wang 0003, Yi Yang 0042 |
AISTATS | 2 |
| 2020 | Generating Plausible Counterfactual Explanations for Deep Transformers in Financial Text ClassificationabstractCorporate mergers and acquisitions (M&A) account for billions of dollars of investment globally every year, and offer an interesting and challenging domain for artificial intelligence.However, in these highly sensitive domains, it is crucial to not only have a highly robust and accurate model, but be able to generate useful explanations to garner a user's trust in the automated system.Regrettably, the recent research regarding eXplainable AI (XAI) in financial text classification has received little to no attention, and many current methods for generating textual-based explanations result in highly implausible explanations, which damage a user's trust in the system.To address these issues, this paper proposes a novel methodology for producing plausible counterfactual explanations, whilst exploring the regularization benefits of adversarial training on language models in the domain of FinTech.Exhaustive quantitative experiments demonstrate that not only does this approach improve the model accuracy when compared to the current stateof-the-art and human performance, but it also generates counterfactual explanations which are significantly more plausible based on human trials. Linyi Yang, Eoin M. Kenny, Tin Lok James Ng, Yi Yang 0042, Barry Smyth, Ruihai Dong |
COLING | 4 |
| 2019 | What You Say and How You Say It Matters: Predicting Stock Volatility Using Verbal and Vocal CuesabstractPredicting financial risk is an essential task in financial market.Prior research has shown that textual information in a firm's financial statement can be used to predict its stock's risk level.Nowadays, firm CEOs communicate information not only verbally through press releases and financial reports, but also nonverbally through investor meetings and earnings conference calls.There are anecdotal evidences that CEO's vocal features, such as emotions and voice tones, can reveal the firm's performance.However, how vocal features can be used to predict risk levels, and to what extent, is still unknown.To fill the gap, we obtain earnings call audio recordings and textual transcripts for S&P 500 companies in recent years.We propose a multimodal deep regression model (MDRM) that jointly model CEO's verbal (from text) and vocal (from audio) information in a conference call.Empirical results show that our model that jointly considers verbal and vocal features achieves significant and substantial prediction error reduction.We also discuss several interesting findings and the implications to financial markets.The processed earnings conference calls data (text and audio) are released for readers who are interested in reproducing the results or designing trading strategy. Yi Yang 0042 |
ACL (1) | 2 |
| 2018 | vec2Link: Unifying Heterogeneous Data for Social Link PredictionabstractRecent advances in network representation learning have enabled significant improvements in the link prediction task, which is at the core of many downstream applications. As an increasing amount of mobility data becoming available due to the development of location technologies, we argue that this resourceful user mobility data can be used to improve link prediction performance. In this paper, we propose a novel link prediction framework that utilizes user offline check-in behavior combined with user online social relations. We model user offline location preference via probabilistic factor model and represent user social relations using neural network embedding. Furthermore, we employ locality-sensitive hashing to project the aggregated user representation into a binary matrix, which not only preserves the data structure but also speeds up the followed convolutional network learning. By comparing with several baseline methods that solely rely on social network or mobility data, we show that our unified approach significantly improves the performance. Fan Zhou 0002, Bangying Wu, Yi Yang 0042, Goce Trajcevski, Kunpeng Zhang 0001, Ting Zhong |
CIKM | 3 |
| 2016 | Improving Topic Model Stability for Effective Document Exploration
Yi Yang 0042, Shimei Pan, Yangqiu Song, Jie Lu 0002, Mercan Topkara |
IJCAI | 1 |
| 2016 | The Stability and Usability of Statistical Topic ModelsabstractStatistical topic models have become a useful and ubiquitous tool for analyzing large text corpora. One common application of statistical topic models is to support topic-centric navigation and exploration of document collections. Existing work on topic modeling focuses on the inference of model parameters so the resulting model fits the input data. Since the exact inference is intractable, statistical inference methods, such as Gibbs Sampling, are commonly used to solve the problem. However, most of the existing work ignores an important aspect that is closely related to the end user experience: topic model stability. When the model is either re-trained with the same input data or updated with new documents, the topic previously assigned to a document may change under the new model, which may result in a disruption of end users’ mental maps about the relations between documents and topics, thus undermining the usability of the applications. In this article, we propose a novel user-directed non-disruptive topic model update method that balances the tradeoff between finding the model that fits the data and maintaining the stability of the model from end users’ perspective. It employs a novel constrained LDA algorithm to incorporate pairwise document constraints, which are converted from user feedback about topics, to achieve topic model stability. Evaluation results demonstrate the advantages of our approach over previous methods. Yi Yang 0042, Shimei Pan, Jie Lu 0002, Mercan Topkara, Yangqiu Song |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2016 | Beating the Artificial Chaos: Fighting OSN Spam Using Its Own TemplatesabstractOnline social networks (OSNs) are extremely popular among Internet users. However, spam originating from friends and acquaintances not only reduces the joy of Internet surfing but also causes damage to less security-savvy users. Prior countermeasures combat OSN spam from different angles. Due to the diversity of spam, there is hardly any existing method that can independently detect the majority or most of OSN spam. In this paper, we empirically analyze the textual pattern of a large collection of OSN spam. An inspiring finding is that the majority (e.g., 76.4% in 2015) of the collected spam is generated with underlying templates. Based on the analysis, we propose tangram, an OSN spam filtering system that performs online inspection on the stream of user-generated messages. Tangram extracts the templates of spam detected by existing methods and then matching messages against the templates toward the accurate and the fast spam detection. It automatically divides the OSN spam into segments and uses the segments to construct templates to filter future spam. Experimental results on Twitter and Facebook data sets show that tangram is highly accurate and can rapidly generate templates to throttle newly emerged campaigns. Furthermore, we analyze the behavior of detected OSN spammers. We find a series of spammer properties-such as spamming accounts are created in bursts and a single active organization orchestrates more spam than all other spammers combined-that promise more comprehensive spam countermeasures. Tiantian Zhu 0001, Yi Yang 0042, Kai Bu, Yan Chen 0004, Doug Downey, Kathy Lee, Alok N. Choudhary |
IEEE/ACM Trans. Netw. | 3 |
| 2015 | Efficient Methods for Inferring Large Sparse Topic HierarchiesabstractDoug Downey, Chandra Bhagavatula, Yi Yang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Doug Downey, Chandra Bhagavatula, Yi Yang 0042 |
ACL (1) | 3 |
| 2015 | Efficient Methods for Incorporating Knowledge into Topic ModelsabstractLatent Dirichlet allocation (LDA) is a popular topic modeling technique for exploring hidden topics in text corpora.Increasingly, topic modeling needs to scale to larger topic spaces and use richer forms of prior knowledge, such as word correlations or document labels.However, inference is cumbersome for LDA models with prior knowledge.As a result, LDA models that use prior knowledge only work in small-scale scenarios.In this work, we propose a factor graph framework, Sparse Constrained LDA (SC-LDA), for efficiently incorporating prior knowledge into LDA.We evaluate SC-LDA's ability to incorporate word correlation knowledge and document label knowledge on three benchmark datasets.Compared to several baseline methods, SC-LDA achieves comparable performance but is significantly faster. Yi Yang 0042, Doug Downey, Jordan L. Boyd-Graber |
EMNLP | 1 |
| 2015 | User-directed Non-Disruptive Topic Model Update for Effective Exploration of Dynamic ContentabstractStatistical topic models have become a useful and ubiquitous text analysis tool for large corpora. One common application of statistical topic models is to support topic-centric navigation and exploration of document collections at the user interface by automatically grouping documents into coherent topics. For today's constantly expanding document collections, topic models need to be updated when new documents become available. Existing work on topic model update focuses on how to best fit the model to the data, and ignores an important aspect that is closely related to the end user experience: topic model stability. When the model is updated with new documents, the topics previously assigned to old documents may change, which may result in a disruption of end users' mental maps between documents and topics, thus undermining the usability of the applications. In this paper, we describe a user-directed non-disruptive topic model update system, nTMU, that balances the tradeoff between finding the model that fits the data and maintaining the stability of the model from end users' perspective. It employs a novel constrained LDA algorithm (cLDA) to incorporate pair-wise document constraints, which are converted from user feedback about topics, to achieve topic model stability. Evaluation results demonstrate advantages of our approach over previous methods. Yi Yang 0042, Shimei Pan, Yangqiu Song, Jie Lu 0002, Mercan Topkara |
IUI | 1 |
| 2014 | Spam ain't as diverse as it seems: throttling OSN spam with templates underneathabstractIn online social networks (OSNs), spam originating from friends and acquaintances not only reduces the joy of Internet surfing but also causes damage to less security-savvy users. Prior countermeasures combat OSN spam from different angles. Due to the diversity of spam, there is hardly any existing method that can independently detect the majority or most of OSN spam. In this paper, we empirically analyze the textual pattern of a large collection of OSN spam. An inspiring finding is that the majority (63.0%) of the collected spam is generated with underlying templates. We therefore propose extracting templates of spam detected by existing methods and then matching messages against the templates toward accurate and fast spam detection. We implement this insight through Tangram, an OSN spam filtering system that performs online inspection on the stream of user-generated messages. Tangram automatically divides OSN spam into segments and uses the segments to construct templates to filter future spam. Experimental results show that Tangram is highly accurate and can rapidly generate templates to throttle newly emerged campaigns. Specifically, Tangram detects the most prevalent template-based spam with 95.7% true positive rate, whereas the existing template generation approach detects only 32.3%. The integration of Tangram and its auxiliary spam filter achieves an overall accuracy of 85.4% true positive rate and 0.33% false positive rate. Yi Yang 0042, Kai Bu, Yan Chen 0004, Doug Downey, Kathy Lee, Alok N. Choudhary |
ACSAC | 2 |
| 2014 | Learning Representations for Weakly Supervised Natural Language Processing TasksabstractFinding the right representations for words is critical for building accurate NLP systems when domain-specific labeled data for the task is scarce. This article investigates novel techniques for extracting features from n-gram models, Hidden Markov Models, and other statistical language models, including a novel Partial Lattice Markov Random Field model. Experiments on part-of-speech tagging and information extraction, among other tasks, indicate that features taken from statistical language models, in combination with more traditional features, outperform traditional representations alone, and that graphical model representations outperform n-gram models, especially on sparse and polysemous words. Fei Huang 0008, Arun Ahuja, Doug Downey, Yi Yang 0042, Yuhong Guo, Alexander Yates |
Comput. Linguistics | 4 |
| 2014 | Incorporating conditional random fields and active learning to improve sentiment identification
Kunpeng Zhang 0001, Yusheng Xie, Yi Yang 0042, Aaron Sun, Hengchang Liu, Alok N. Choudhary |
Neural Networks | 3 |
| 2013 | Overcoming the Memory Bottleneck in Distributed Training of Latent Variable Models of Text
Yi Yang 0042, Alexander Yates, Doug Downey |
HLT-NAACL | 1 |