Chung-Chi Chen 0001

dblp:177/6602 · DBLP profile ↗
← Back
23ranked-venue papers in the field
15as first author
18since 2021 · last 2026
0000-0003-3680-9277ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 19 (12 first)Data Mining & Knowledge Discovery · 3 (2 first)Other / Interdisciplinary · 1 (1 first)
YearPublicationVenuePosition
2026 The First Workshop on Information Retrieval for Accountability and Integrity (IRAI)
Chung-Chi Chen 0001, Juyeon Kang, Anaïs Lhuissier, Dittaya Wanvarie, Min-Yuh Day, Hiroya Takamura, Yohei Seki
ECIR (3)1
2025 Advances in Financial AI: Innovations, Risk, and Responsibility in the Era of LLMs
abstract
The finance sector is seeing a rapid increase in the application of machine learning and AI, with Large Language Models (LLMs), ESG (Environmental, Social, and Governance) investing, and AI Safety significantly reshaping the field. This workshop focuses on how these advancements intersect with core financial AI applications. We will foster interdisciplinary discussion on applying LLMs to finance, addressing challenges in multilingual and non-English markets like Korea. The event will also highlight the integration of ESG signals into algorithmic decision-making and explore AI Safety, emphasizing reliability, fairness, and explainability for AI systems in regulated financial environments. By bringing together experts from academia, industry, and regulatory bodies, the workshop aims to stimulate discussions on practical issues, ethical dilemmas, and cutting-edge research shaping financial AI's future. We welcome submissions that combine technical rigor with societal relevance in AI-driven financial decisions.
Nazanin Mehrasa, Chanyeol Choi, Chung-Chi Chen 0001, Dhagash Mehta, Stefan Zohren, Chulheum Lee, Yeonhee Lee, Eunsook Oh
CIKM4
2025 Company-Specific Knowledge Matters: Retrieval-Augmented Generation for Earnings Call Answer Rehearsal
abstract
Retrieval-augmented generation (RAG) has long been used to guide generative models in producing more accurate answers, with most discussions focusing on reading comprehension-based question answering (QA). However, their role in real-world answer rehearsal scenarios remains underexplored. The rise of large language models (LLMs) presents new opportunities to develop systems that assist professionals, making this discussion both timely and essential. This paper explores how to better support corporate executives in answering questions from professional analysts during earnings calls. We compare the impact of two external knowledge sources: large-scale causal knowledge graphs (KGs) and historical Q&A records-retrieved either from a global pool or company-specific archives. Our findings suggest that a company's historical Q&A records are more influential than causal KGs in improving response quality. To the best of our knowledge, this is the first study to systematically compare and analyze different knowledge resources in answer rehearsal. Our findings show the potential of inspiring future research on the interplay between KG and historical QA pairs for answer rehearsal.
Yung-Yu Shih, Yun-Nung Chen, Chung-Chi Chen 0001
CIKM3
2025 Information Retrieval in Finance: Industry and Academic Perspectives on Innovation
abstract
Information retrieval (IR) plays a critical role in financial decision-making across investment research, trading, risk management, and reporting. With the rise of large language models (LLMs), IR systems have evolved to support more natural, context-aware workflows. In this tutorial, we survey recent advances in applying IR and LLM technologies in finance, covering agent-based simulations, investor recommender systems, retrieval-augmented research management, and LLM-driven portfolio construction. We highlight practical challenges and propose future research directions at the intersection of IR, LLMs, and financial innovation. More materials can be found at http://irfin.nlpfin.com/.
Chung-Chi Chen 0001, Alejandro Lopez-Lira, Chanyeol Choi, Richard McCreadie, Javier Sanz-Cruzado
SIGIR1
2024 Professionalism-Aware Pre-Finetuning for Profitability Ranking
abstract
Opinion mining, specifically in the investment sector, has experienced a significant increase in interest over recent years. This paper presents a novel approach to overcome current limitations in assessing and ranking investor opinions based on profitability. The study introduces a pre-finetuning scheme to improve language models' capacity to distinguish professionalism, thus enabling ranking of all available opinions. Furthermore, the paper evaluates ranking results using traditional metrics and suggests the use of a pairwise setting for better performances over a regression setting. Lastly, our method is shown to be effective across various investor opinion tasks, encompassing both professional and amateur investors. The results indicate that this approach significantly enhances the efficiency and accuracy of opinion mining in the investment sector.
Chung-Chi Chen 0001, Hiroya Takamura, Ichiro Kobayashi 0001, Yusuke Miyao
CIKM1
2024 Automation of Text-Based Economic Indicator Construction: A Pilot Exploration on Economic Policy Uncertainty Index
abstract
The growing popularity of text-as-data in various domain-specific applications and research has often relied on manually selected keywords or annotations. Although labor-intensive, expensive and time-consuming, the effectiveness of these efforts is not always guaranteed, especially in the early stages of research. This predicament raises the question of the extent to which large language models (LLMs) can aid in verifying the potential of a nascent research idea. This paper seeks to explore the reliability of LLM-suggested keywords in the automatic construction of the Economic Policy Uncertainty (EPU) index. Our findings confirm that LLMs can effectively automate the construction of EPU index. Furthermore, we delve into the potential of LLMs in enhancing the indicator construction process.
Hsiu-Hsuan Yeh, Yu-Lieh Huang, Ziho Park, Chung-Chi Chen 0001
CIKM4
2023 DynamicESG: A Dataset for Dynamically Unearthing ESG Ratings from News Articles
abstract
This paper introduces the DynamicESG dataset, a unique resource for dynamically extracting ESG ratings from news articles. The ESG rating, a novel metric employed annually to gauge a company's sustainability, relies heavily on corporate disclosure and other external information, especially news narratives. Our dataset, comprising a wide spectrum of news over a twelve-year span, annotates articles in accordance with MSCI ESG ratings methodology and SASB standards, with relevance to ESG issues. DynamicESG provides a comprehensive means of investigating the relationship between public discourse, ESG-related events, and subsequent ESG rating adjustments. We detail our data collection, curation, annotation procedure, and inter-rater agreement, ensuring high data quality and usability. Importantly, our dataset includes a temporal dimension, enabling the analysis of longitudinal trends in ESG ratings and their correlation with news coverage. Moreover, the dataset incorporates an opportunity/risk tendency, thus permitting analysis from diverse perspectives to discern if the news is beneficial or detrimental to the company. We believe this dataset will serve as a valuable resource for researchers in fields such as corporate social responsibility, sustainable investing, machine learning, and natural language processing. Initial analysis using the dataset underscores its potential to facilitate new insights into the dynamics of ESG ratings and the influence of news media on these ratings.
Yu-Min Tseng, Chung-Chi Chen 0001, Hen-Hsen Huang, Hsin-Hsi Chen
CIKM2
2023 Personalized Dynamic Recommender System for Investors
abstract
With the development of online platforms, people can share and obtain opinions quickly. It also makes individuals' preferences change dynamically and rapidly because they may change their minds when getting convincing opinions from other users. Unlike representative areas of recommendation research such as e-commerce platforms where items' features are fixed, in investment scenarios financial instruments' features such as stock price, also change dynamically over time. To capture these dynamic features and provide a better-personalized recommendation for amateur investors, this study proposes a Personalized Dynamic Recommender System for Investors, PDRSI. The proposed PDRSI considers two investor's personal features: dynamic preferences and historical interests, and two temporal environmental properties: recent discussions on the social media platform and the latest market information. The experimental results support the usefulness of the proposed PDRSI, and the ablation studies show the effect of each module. For reproduction, we follow Twitter's developer policy to share our dataset for future work.
Takehiro Takayanagi, Chung-Chi Chen 0001, Kiyoshi Izumi
SIGIR2
2023 FinTech on the Web: An Overview
abstract
In this article, we provide an overview of ACM TWEB’s special issue, Financial Technology on the Web . This special issue covers diverse topics: (1) a new architecture for leveraging online news to investment and risk management, (2) a cross-platform analysis of the post quality and users’ behaviors, and (3) an empirical study on disentangling decentralized finance compositions. In addition to a guide for the special issue, we also share a brief opinion on the future of financial technology on the Web.
Chung-Chi Chen 0001, Hen-Hsen Huang, Hiroya Takamura, Makoto P. Kato, Yu-Lieh Huang
ACM Trans. Web1
2021 Which kind of rumors may undermine society: perspectives from court orders
abstract
Freedom of speech is one of the principles in the constitution of most countries. However, in the 2020 United States presidential election, Donald Trump's Twitter account is suspended due to the risk of further incitement of violence. That leads to the question: Which kind of rumors may undermine society? In this paper, we discuss this question based on the case studies of real-world court orders, which are the judges' official proclamations. We point out the possible research directions that NLP researchers may need to consider before applying our systems to society.
Chung-Chi Chen 0001, Hen-Hsen Huang, Hsin-Hsi Chen
ASONAM1
2021 NQuAD: 70, 000+ Questions for Machine Comprehension of the Numerals in Text
abstract
Numeral information plays an important role in narratives of several domains such as medicine, engineering, and finance. Previous works focus on the foundation exploration toward numeracy and show that fine-grained numeracy is a challenging task. In machine reading comprehension, our statistics show that only a few numeral-related questions appear in previous datasets. It indicates that few benchmark datasets are designed for numeracy learning. In this paper, we present a Numeral-related Question Answering Dataset, NQuAD, for fine-grained numeracy, and propose several baselines for future works. We compare NQuAD with three machine reading comprehension datasets and show that NQuAD is more challenging than the numeral-related questions in other datasets. NQuAD is published under the CC BY-NC-SA 4.0 license for academic purposes.
Chung-Chi Chen 0001, Hen-Hsen Huang, Hsin-Hsi Chen
CIKM1
2021 Constructing Noise Free Economic Policy Uncertainty Index
abstract
The economic policy uncertainty (EPU) index is one of the important text-based indexes in finance and economics fields. The EPU indexes of more than 26 countries have been constructed to reflect the policy uncertainty on country-level economic environments and serve as an important economic leading indicator. The EPU indexes are calculated based on the number of news articles with some manually-selected keywords related to economic, uncertainty, and policy. We find that the keyword-based EPU indexes contain noise, which will influence their explainability and predictability. In our experimental dataset, over 40% of news articles with the selected keywords are not related to the EPU. Instead of using keywords only, our proposed models take contextual information into account and get good performance on identifying the articles unrelated to EPU. The noise free EPU index performs better than the keyword-based EPU index in both explainability and predictability.
Chung-Chi Chen 0001, Hen-Hsen Huang, Yu-Lieh Huang, Hsin-Hsi Chen
CIKM1
2021 Distilling Numeral Information for Volatility Forecasting
abstract
The volatility of stock price reflects the risk of stock and influences the risk of investor's portfolio. It is also a crucial part of pricing derivative securities. Researchers have paid their attention to predict the stock volatility with different kinds of textual data. However, most of them focus on using word information only. Few touch on capturing the numeral information in textual data, providing fine-grained clues for financial document understanding. In this paper, we present a novel dataset, ECNum, for understanding the numerals in the transcript of earnings conference calls. We propose a simple but efficient method, Numeral-Aware Model (NAM), for enhancing the capacity of numeral understanding of neural network models. We employ the distilled information in the stock volatility forecasting task and achieve the best performance compared to the previous works in short-term scenarios.
Chung-Chi Chen 0001, Hen-Hsen Huang, Yu-Lieh Huang, Hsin-Hsi Chen
CIKM1
2021 A Research Agenda for Financial Opinion Mining
Chung-Chi Chen 0001, Hen-Hsen Huang, Hsin-Hsi Chen
ICWSM1
2021 Risk-aware Regularization for Opinion-based Portfolio Selection
Ting-Wei Hsu, Chung-Chi Chen 0001, Hen-Hsen Huang, Hsin-Hsi Chen
ICWSM2
2021 Retrieving Implicit Information for Stock Movement Prediction
abstract
Previous studies on the financial news focus mainly on the news articles explicitly mentioning the target financial instruments, and may suffer from data sparsity. As taking into consideration other related news, e.g., sector-related news, is a crucial part of real-world decision-making, we explore the use of news without explicit target mentions to enrich the information for the prediction model. We develop a neural network framework that jointly learns with a news selection mechanism to extract implicit information from the chaotic daily news pool. Our proposed model, called the news distilling network (NDN), takes advantage of neural representation learning and collaborative filtering to capture the relationship between stocks and news. With NDN, we learn latent stock and news representations to facilitate similarity measurements, and apply a gating mechanism to prevent noisy news representations from flowing to a higher level encoding stage, which encodes the selected news representation of each day. Extensive experiments on real-world stock market data demonstrate the effectiveness of our framework and show improvements over previous techniques.
Tsun-Hsien Tang, Chung-Chi Chen 0001, Hen-Hsen Huang, Hsin-Hsi Chen
SIGIR2
2021 FinSense: An Assistant System for Financial Journalists and Investors
abstract
This paper demonstrates FinSense, a system that improves the working efficiency of financial information processing. Given the draft of a financial news story, FinSense extracts the explicit-mentioned stocks and further infers the implicit stocks, providing insightful information for decision making. We propose a novel graph convolutional network model that performs implicit financial instrument inference toward the in-domain data. In addition, FinSense generates candidate headlines for the draft, reducing a significant amount of time in journalism production. The proposed system also provides assistance to investors to sort out the information in the financial news articles.
Yi-Ting Liou, Chung-Chi Chen 0001, Tsun-Hsien Tang, Hen-Hsen Huang, Hsin-Hsi Chen
WSDM2
2021 Evaluating the Rationales of Amateur Investors
abstract
Social media’s rise in popularity has demonstrated the usefulness of the wisdom of the crowd. Most previous works take into account the law of large numbers and simply average the results extracted from tasks such as opinion mining and sentiment analysis. Few attempt to identify high-quality opinions from the mined results. In this paper, we propose an approach for capturing expert-like rationales from social media platforms without the requirement of the annotated data. By leveraging stylistic and semantic features, our approach achieves an F1-score of 90.81%. The comparison between the rationales of experts and those of the crowd is done from stylistic and semantic perspectives, revealing that stylistic and semantic information provides complementary cues for professional rationales. We further show the advantage of using these superlative analysis results in the financial market, and find that top-ranked opinions identified by our approach increase potential returns by up to 90.31% and reduce downside risk by up to 71.69%, compared with opinions ranked by feedback from social media users. Moreover, the performance of our method on downside risk control is comparable with that of professional analysts.
Chung-Chi Chen 0001, Hen-Hsen Huang, Hsin-Hsi Chen
WWW1
2020 NumClaim: Investor's Fine-grained Claim Detection
abstract
The goal of claim detection in argument mining is to sort out the key points from a long narrative. In this paper, we design a novel task for argument mining in the financial domain, and provide an expert-annotated dataset, NumClaim, for the proposed task. Based on the statistics, we discuss the differences between the claims in other datasets and the claims of the investors in NumClaim. With the ablation analysis, we show that encoding numeral and co-training with the auxiliary task of the numeral understanding, i.e., the category classification task, can improve the performance of the proposed task under different neural network architectures. The annotations in the NumClaim is published for academic usage under the CC BY-NC-SA 4.0 license.
Chung-Chi Chen 0001, Hen-Hsen Huang, Hsin-Hsi Chen
CIKM1
2019 Next cashtag prediction on social trading platforms with auxiliary tasks
abstract
Social trading platforms provide a forum for investors to share their analysis and opinions. Posts on these platforms are characterized by narrative styles which are much different from posts on general social platforms, for instance tweets. As a result, recommendation systems for social trading platforms should leverage tailor-made latent features. This paper presents a representation for these latent features in both textual data and market information. A real-world dataset is adopted to conduct experiments involving a novel task called next cashtag prediction. We propose a joint learning model with an attentive capsule network. Experimental results show positive results with the proposed methods and the corresponding auxiliary tasks.
Chung-Chi Chen 0001, Hen-Hsen Huang, Hsin-Hsi Chen
ASONAM1
2019 Numeral Attachment with Auxiliary Tasks
abstract
In this paper we propose the task of numeral attachment to detect the attached target of a numeral. Compared with other kinds of named entities, numerals provide richer and more crucial information in some domains. Fine-grained understanding of the information embedded in numerals is a fundamental challenge. We develop NumAttach, a pilot dataset for the proposed task based on tweets. Two main challenges of this task include the informal writing style in tweets and the representation of numerals. To address these challenges, we present an embedding technique that considers word and numeral information simultaneously. Furthermore, we design a joint learning model with the capsule network to accomplish the proposed task. We also release NumAttach to the research community as a resource.
Chung-Chi Chen 0001, Hen-Hsen Huang, Hsin-Hsi Chen
SIGIR1
2019 CrowdPT: Summarizing Crowd Opinions as Professional Analyst
abstract
This paper demonstrates a novel analytics service, CrowdPT, for capturing the key information, price target (PT), of individual investors on social media. PT, which is mentioned as a conclusion in most of analysts' reports, indicates not only the market sentiment (bullish/bearish) of investors, but also the analysis results. In order to provide the latest opinions of individual investors, we monitor Twitter in real time and update the information in price chart daily. For all component stocks in Dow Jones Industrial Average, textual information from numerous tweets is summarized into a single number, PT, in CrowdPT. Case studies confirm the effectiveness of our analytics service in the financial domain, and show that capturing the PT of individual investors is promising for stock price prediction. The Web API of CrowdPT is also provided for academic purpose.
Chung-Chi Chen 0001, Hen-Hsen Huang, Chia-Wen Tsai, Hsin-Hsi Chen
WWW1
2018 Numeral Understanding in Financial Tweets for Fine-Grained Crowd-Based Forecasting
abstract
Numerals that contain much information in financial documents are crucial for financial decision making. They play different roles in financial analysis processes. This paper is aimed at understanding the meanings of numerals in financial tweets for fine-grained crowd-based forecasting. We propose a taxonomy that classifies the numerals in financial tweets into 7 categories, and further extend some of these categories into several subcategories. Neural network-based models with word and character-level encoders are proposed for 7-way classification and 17-way classification. We perform backtest to confirm the effectiveness of the numeric opinions made by the crowd. This work is the first attempt to understand numerals in financial social media data, and we provide the first comparison of fine-grained opinion of individual investors and analysts based on their forecast price. The numeral corpus used in our experiments, called FinNum 1.0, is available for research purposes.
Chung-Chi Chen 0001, Hen-Hsen Huang, Yow-Ting Shiue, Hsin-Hsi Chen
WI1