EDBT 2026 Demo / reviewers in the wild / expert
Chuan-Ju Wang
dblp:03/5904
· DBLP profile ↗
25ranked-venue papers in the field
1as first author
15since 2021 · last 2026
0000-0002-5281-2962ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 19 (1 first)Data Mining & Knowledge Discovery · 4Database Systems & Data Management · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | G-TRAC: Graph-textual Representations Alignment for Cold-start RecommendationsabstractThe cold-start problem remains a significant challenge in recommendation systems, particularly for new users or unseen items with little to no historical data. Existing methods, including graph neural networks, often struggle in such scenarios. Inspired by the success of transformer models in natural language processing, we propose G-TRAC (Graph-Textual Representations Alignment for Cold-start Recommendations), a novel approach that integrates transformer-based textual modeling with graph neural networks. By effectively leveraging both textual and structural information, G-TRAC addresses cold-start challenges more effectively. Extensive experiments demonstrate its ability to enhance recommendation quality and generalize well across diverse scenarios. Li-Yang Chang, Yuan Fang 0001, Ming-Feng Tsai, Chuan-Ju Wang |
WSDM | 4 |
| 2025 | Dynamic Margin-based Contrastive Learning for Robust Negative Sampling in Information RetrievalabstractModern information retrieval (IR) systems, powered by bi-encoder architectures and pretrained language models, rely on effective negative sampling for contrastive learning. While easy negatives are computationally simple, they fail to challenge the model, whereas hard negatives-selected via methods like BM25, ANCE, or ADORE-can be overly difficult and misleading. To this end, this paper proposes dynamic margin-based contrastive learning (DMCL), which adaptively adjusts the decision boundary based on query-negative similarity, ensuring consistent exposure to moderately hard negatives. Experiments across diverse datasets and models show that DMCL outperforms traditional methods, achieving state-of-the-art retrieval performance with minimal computational cost. Tsai-Tsung Chen, Chuan-Ju Wang, Ming-Feng Tsai |
SIGIR | 2 |
| 2025 | BETag: Behavior-enhanced Item Tagging with Finetuned Large Language ModelsabstractTags play a critical role in enhancing product discoverability, optimizing search results, and enriching recommendation systems on e-commerce platforms. Despite the recent advancements in large language models (LLMs), which have shown proficiency in processing and understanding textual information, their application in tag generation remains an under-explored yet complex challenge. To this end, we introduce a novel method for automatic product tagging using LLMs to create behavior-enhanced tags (BETags). Specifically, our approach begins by generating base tags using an LLM. These base tags are then refined into BETags by incorporating user behavior data. This method aligns the tags with users' actual browsing and purchasing behavior, enhancing the accuracy and relevance of tags to user preferences. By personalizing the base tags with user behavior data, BETags are able to capture deeper behavioral insights, which is essential for understanding nuanced user interests and preferences in e-commerce environments. Moreover, since BETags are generated offline, they do not impose real-time computational overhead and can be seamlessly integrated into downstream tasks commonly associated with recommendation systems and search optimization. Our evaluation of BETag across three datasets--- Amazon (Scientific), MovieLens-1M, and FreshFood---shows that our approach significantly outperforms both human-annotated tags and other automated methods. These results highlight BETag as a scalable and efficient solution for personalized automated tagging, advancing e-commerce platforms by creating more tailored and engaging user experiences. Shao-En Lin, Miao-Chen Chiang, Ming-Yi Hong 0002, Yu-Shiang Huang, Chuan-Ju Wang, Che Lin |
WWW | 6 |
| 2024 | Accurate Neural Network Option Pricing Methods with Control Variate Techniques and Data Synthesis/Cleaning with Financial RationalityabstractThis paper enhances option pricing accuracy by incorporating financial expertise into a neural network (NN) design and optimizing data sample quality through cleaning and synthesis. Instead of directly estimating option values (OVs) with NNs, we leverage the concept of control variate by decomposing OVs as time values (TVs) estimated by NNs, plus the analytically solvable intrinsic values (IVs). TV surface can be decomposed into two scenarios with very different properties, and we design two NNs according to our derived no-arbitrage constraints for these two scenarios. To alleviate learning inaccuracy due to the kink of the TV surface along the scenario boundary, we synthesize training samples based on our derived constraints to smoothly extend the surface for each scenario. On the other hand, irrational option quotes commonly found in illiquid markets incur uneven surfaces, significantly deteriorating NN predictability. We develop a learnable data-cleaning method to remove potentially irrational quotes spotted by no-arbitrage constraints properly. Besides, unnecessary data syntheses proposed in previous literature can also be removed by incorporating corresponding constraints into our NN to enhance training efficiency. Comprehensive experiments on liquid S&P 500 and illiquid TAIEX option markets examine the superiority of our approach. Chia-Wei Hsu, Tian-Shyr Dai, Chuan-Ju Wang, Ying-Ping Chen |
CIKM | 3 |
| 2023 | CPR: Cross-Domain Preference Ranking with User Transformation
Hsien-Hao Chen, Tung-Lin Wu, Chia-Yu Yeh, Jing-Kai Lou, Ming-Feng Tsai, Chuan-Ju Wang |
ECIR (2) | 7 |
| 2023 | Improving Conversational Passage Re-ranking with View EnsembleabstractThis paper presents ConvRerank, a conversational passage re-ranker that employs a newly developed pseudo-labeling approach. Our proposed view-ensemble method enhances the quality of pseudo-labeled data, thus improving the fine-tuning of ConvRerank. Our experimental evaluation on benchmark datasets shows that combining ConvRerank with a conversational dense retriever in a cascaded manner achieves a good balance between effectiveness and efficiency. Compared to baseline methods, our cascaded pipeline demonstrates lower latency and higher top-ranking effectiveness. Furthermore, the in-depth analysis confirms the potential of our approach to improving the effectiveness of conversational search. Jia-Huei Ju, Sheng-Chieh Lin, Ming-Feng Tsai, Chuan-Ju Wang |
SIGIR | 4 |
| 2022 | INForex: Interactive News Digest for Forex Investors
Chih-Hen Lee, Yi-Shyuan Chiang, Chuan-Ju Wang |
ECIR (2) | 3 |
| 2022 | Multiperiod Corporate Default Prediction Through Neural Parametric Family LearningabstractDefault analysis plays an essential role in financial markets because it narrows the information gap between borrowers and lenders. Of late, machine learning-based methods have found their way to default analysis and typically view it as a risk classification task by slotting obligors into risk categories. The quality of such an approach is assessed by its prediction accuracy in risk rankings. Rarely considered but important are issues on the predicted numbers of default occurrences and the term structure of cumulative default probabilities for which classification tools are by nature silent. In this paper, we depart from the typical practice of risk classification and focus on employing machine learning to estimate the term structure of cumulative default probabilities—a structured estimation that contains default probabilities from short-term to long-term periods. To this end, we formulate the task as a problem of parametric family learning via a neural model consisting of two segments: parameter generation and parametric family determination. The proposed neural approach offers added flexibility in improving long-term default predictions. Moreover, the carefully designed model successfully maintains vital economic characteristics of its predictions. Experiments on a US corporate default dataset show that our approach achieves measurably better prediction performance in both risk classification and matching the predicted numbers of default occurrences with the actual ones. Wei-Lun Luo, Yu-Ming Lu, Jheng-Hong Yang, Jin-Chuan Duan, Chuan-Ju Wang |
SDM | 5 |
| 2022 | RecDelta: An Interactive Dashboard on Top-k Recommendation for Cross-model EvaluationabstractIn this demonstration, we present RecDelta, an interactive tool for the cross-model evaluation of top-k recommendation. RecDelta is a web-based information system where people visually compare the performance of various recommendation algorithms and their recommended items. In the proposed system, we visualize the distribution of the δ scores between algorithms--a distance metric measuring the intersection between recommendation lists. Such visualization allows for rapid identification of users for whom the items recommended by different algorithms diverge or vice versa; then, one can further select the desired user to present the relationship between recommended items and his/her historical behavior. RecDelta benefits both academics and practitioners by enhancing model explainability as they develop recommendation algorithms with their newly gained insights. Note that while the system is now online at https://cfda.csie.org/recdelta, we also provide a video recording at https://tinyurl.com/RecDelta to introduce the concept and the usage of our system. Yi-Shyuan Chiang, Yu-Ze Liu, Chen-Feng Tsai, Jing-Kai Lou, Ming-Feng Tsai, Chuan-Ju Wang |
SIGIR | 6 |
| 2022 | IPR: Interaction-level Preference Ranking for Explicit feedbackabstractExplicit feedback---user input regarding their interest in an item---is the most helpful information for recommendation as it comes directly from the user and shows their direct interest in the item. Most approaches either treat the recommendation given such feedback as a typical regression problem or regard such data as implicit and then directly adopt approaches for implicit feedback; both methods, however,tend to yield unsatisfactory performance in top-k recommendation. In this paper, we propose interaction-level preference ranking(IPR), a novel pairwise ranking embedding learning approach to better utilize explicit feedback for recommendation. Experiments conducted on three real-world datasets show that IPR yields the best results compared to six strong baselines. Shih-Yang Liu, Hsien-Hao Chen, Chih-Ming Chen 0003, Ming-Feng Tsai, Chuan-Ju Wang |
SIGIR | 5 |
| 2022 | Item Concept Network: Towards Concept-Based Item Representation LearningabstractItem concept modeling is commonly achieved by leveraging textual information. However, many existing models do not leverage the inferential property of concepts to capture word meanings, which therefore ignores the relatedness between correlated concepts, a phenomenon which we term conceptual “correlation sparsity.” In this paper, we distinguish between word modeling and concept modeling and propose an item concept modeling framework centering around the item concept network (ICN). ICN models and further enriches item concepts by leveraging the inferential property of concepts and thus addresses the correlation sparsity issue. Specifically, there are two stages in the proposed framework: ICN construction and embedding learning. In the first stage, we propose a generalized network construction method to build ICN, a structured network which infers expanded concepts for items via matrix operations. The second stage leverages neighborhood proximity to learn item and concept embeddings. With the proposed ICN, the resulting embedding facilitates both homogeneous and heterogeneous tasks, such as item-to-item and concept-to-item retrieval, and delivers related results which are more diverse than traditional keyword-matching-based approaches. As our experiments on two real-world datasets show, the framework encodes useful conceptual information and thus outperforms traditional methods in various item classification and retrieval tasks. Ting-Hsiang Wang, Hsiu-Wei Yang, Chih-Ming Chen 0003, Ming-Feng Tsai, Chuan-Ju Wang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2021 | XRR: Explainable Risk Ranking for Financial Reports
Ting-Wei Lin, Ruei-Yao Sun, Hsuan-Ling Chang, Chuan-Ju Wang, Ming-Feng Tsai |
ECML/PKDD (4) | 4 |
| 2021 | Text-to-Text Multi-view Learning for Passage Re-rankingabstractRecently, much progress in natural language processing has been driven by deep contextualized representations pretrained on large corpora. Typically, the fine-tuning on these pretrained models for a specific downstream task is based on single-view learning, which is however inadequate as a sentence can be interpreted differently from different perspectives. Therefore, in this work, we propose a text-to-text multi-view learning framework by incorporating an additional view---the text generation view---into a typical single-view passage ranking model. Empirically, the proposed approach is of help to the ranking performance compared to its single-view counterpart. Component analysis is also reported in the paper. Jia-Huei Ju, Jheng-Hong Yang, Chuan-Ju Wang |
SIGIR | 3 |
| 2021 | LSTPR: Graph-based Matrix Factorization with Long Short-term Preference RankingabstractConsidering the temporal order of user-item interactions for recommendation forms a novel class of recommendation algorithms in recent years, among which sequential recommendation models are the most popular approaches. Although, theoretically, such fine-grained modeling should be beneficial to the recommendation performance, these sequential models in practice greatly suffer from the issue of data sparsity as there are a huge number of combinations for item sequences. To address the issue, we propose LSTPR, a graph-based matrix factorization model that incorporates both high-order graph information and long short-term user preferences into the modeling process. LSTPR explicitly distinguishes long-term and short-term user preferences and enriches the sparse interactions via random surfing on the user-item graph. Experiments on three recommendation datasets with temporal user-item information demonstrate that the proposed LSTPR model achieves significantly better performance than the seven baseline methods. Chih-Hen Lee, Jun-En Ding, Chih-Ming Chen 0003, Jing-Kai Lou, Ming-Feng Tsai, Chuan-Ju Wang |
SIGIR | 6 |
| 2021 | Multi-Stage Conversational Passage Retrieval: An Approach to Fusing Term Importance Estimation and Neural Query RewritingabstractConversational search plays a vital role in conversational information seeking. As queries in information seeking dialogues are ambiguous for traditional ad hoc information retrieval (IR) systems due to the coreference and omission resolution problems inherent in natural language dialogue, resolving these ambiguities is crucial. In this article, we tackle conversational passage retrieval, an important component of conversational search, by addressing query ambiguities with query reformulation integrated into a multi-stage ad hoc IR system. Specifically, we propose two conversational query reformulation (CQR) methods: (1) term importance estimation and (2) neural query rewriting. For the former, we expand conversational queries using important terms extracted from the conversational context with frequency-based signals. For the latter, we reformulate conversational queries into natural, stand-alone, human-understandable queries with a pretrained sequence-to-sequence model. Detailed analyses of the two CQR methods are provided quantitatively and qualitatively, explaining their advantages, disadvantages, and distinct behaviors. Moreover, to leverage the strengths of both CQR methods, we propose combining their output with reciprocal rank fusion, yielding state-of-the-art retrieval effectiveness, 30% improvement in terms of NDCG@3 compared to the best submission of Text REtrieval Conference (TREC) Conversational Assistant Track (CAsT) 2019. Sheng-Chieh Lin, Jheng-Hong Yang, Rodrigo Nogueira 0001, Ming-Feng Tsai, Chuan-Ju Wang, Jimmy Lin |
ACM Trans. Inf. Syst. | 5 |
| 2020 | TPR: Text-aware Preference Ranking for Recommender SystemsabstractTextual data is common and informative auxiliary information for recommender systems. Most prior art utilizes text for rating prediction, but rare work connects it to top-recommendation. Moreover, although advanced recommendation models capable of incorporating auxiliary information have been developed, none of these are specifically designed to model textual information, yielding a limited usage scenario for typical user-to-item recommendation. In this work, we present a framework of text-aware preference ranking (TPR) for top- recommendation, in which we comprehensively model the joint association of user-item interaction and relations between items and associated text. Using the TPR framework, we construct a joint likelihood function that explicitly describes two ranking structures: 1) item preference ranking (IPR) and 2) word relatedness ranking (WRR), where the former captures the item preference of each user and the latter captures the word relatedness of each item. As these two explicit structures are by nature mutually dependent, we propose TPR-OPT, a simple yet effective learning criterion that additionally includes implicit structures, such as relatedness between items and relatedness between words for each user for model optimization. Such a design not only successfully describes the joint association among users, words, and text comprehensively but also naturally yields powerful representations that are suitable for a range of recommendation tasks, including user-to-item, item-to-item, and user-to-word recommendation, as well as item-to-word reconstruction. In this paper, extensive experiments have been conducted on eight recommendation datasets, the results of which demonstrate that by including textual information from item descriptions, the proposed TPR model consistently outperforms state-of-the-art baselines on various recommendation tasks. Yu-Neng Chuang, Chih-Ming Chen 0003, Chuan-Ju Wang, Ming-Feng Tsai, Yuan Fang 0001, Ee-Peng Lim |
CIKM | 3 |
| 2019 | Keyword Extraction with Character-Level Convolutional Neural Tensor Networks
Zhe-Li Lin, Chuan-Ju Wang |
PAKDD (1) | 2 |
| 2019 | SMORe: modularize graph embedding for recommendationabstractIn the Age of Big Data, graph embedding has received increasing attention for its ability to accommodate the explosion in data volume and diversity, which challenge the foundation of modern recommender systems. Respectively, graph facilitates fusing complex systems of interactions into a unified structure and distributed embedding enables efficient retrieval of entities, as in the case of approximate nearest neighbor (ANN) search. When combined, graph embedding captures relational information beyond entity interaction and towards a problem's underlying structure, as epitomized by struct2vec [20] and PinSage [26]. This session will start by brushing up on the basics about graphs and embedding methods and discussing their merits. We then quickly dive into using the mathematical formulation of graph embedding to derive the modular framework: Sampler-Mapper-Optimizer for Recommendation, or SMORe. We demonstrate existing models used for recommendation, such as MF and BPR, can all be assembled using three basic components: sampler, mapper, and optimizer. The tutorial is accompanied by a hands-on session, where we show how graph embedding can model complex systems through the multi-task learning and the cross-platform data sparsity alleviation tasks. Chih-Ming Chen 0003, Ting-Hsiang Wang, Chuan-Ju Wang, Ming-Feng Tsai |
RecSys | 3 |
| 2019 | Collaborative Similarity Embedding for Recommender SystemsabstractWe present collaborative similarity embedding (CSE), a unified framework that exploits comprehensive collaborative relations available in a user-item bipartite graph for representation learning and recommendation. In the proposed framework, we differentiate two types of proximity relations: direct proximity and k-th order neighborhood proximity. While learning from the former exploits direct user-item associations observable from the graph, learning from the latter makes use of implicit associations such as user-user similarities and item-item similarities, which can provide valuable information especially when the graph is sparse. Moreover, for improving scalability and flexibility, we propose a sampling technique that is specifically designed to capture the two types of proximity relations. Extensive experiments on eight benchmark datasets show that CSE yields significantly better performance than state-of-the-art recommendation methods. Chih-Ming Chen 0003, Chuan-Ju Wang, Ming-Feng Tsai, Yi-Hsuan Yang |
WWW | 2 |
| 2018 | HOP-rec: high-order proximity for implicit recommendationabstractRecommender systems are vital ingredients for many e-commerce services. In the literature, two of the most popular approaches are based on factorization and graph-based models; the former approach captures user preferences by factorizing the observed direct interactions between users and items, and the latter extracts indirect preferences from the graphs constructed by user-item interactions. In this paper we present HOP-Rec, a unified and efficient method that incorporates the two approaches. The proposed method involves random surfing on a graph to harvest high-order information among neighborhood items for each user. Instead of factorizing a transition matrix, our method introduces a confidence weighting parameter to simulate all high-order information simultaneously, for which we maintain a sparse user-item interaction matrix and enrich the matrix for each user using random walks. Experimental results show that our approach significantly outperforms the state of the art on a range of large-scale real-world datasets. Jheng-Hong Yang, Chih-Ming Chen 0003, Chuan-Ju Wang, Ming-Feng Tsai |
RecSys | 3 |
| 2018 | NavWalker: Information Augmented Network EmbeddingabstractWe present NavWalker, a flexible random walk-based approach for learning the representations of vertices in an information network. The proposed method enables us to incorporate different walk strategies into the sampling process of random walks, in order to further boost the network embedding techniques. Specifically, we formulate the proposed method by integrating the adjacency matrix of a network with a pre-defined information augmentation matrix. In contrast to SkipGram-based network embedding methods such as DeepWalk and Node2vec, which use only local network information to learn the representations, our method is flexible to further incorporate global or other auxiliary network information to guide the sampling process. Experiments on six real-world datasets demonstrate the advantages of the flexibility and its superior performance as compared to other state-of-the-art network embedding algorithms for the tasks of classification and recommendation. Kwei-Herng Lai, Chih-Ming Chen 0003, Ming-Feng Tsai, Chuan-Ju Wang |
WI | 4 |
| 2017 | Text Embedding for Sub-Entity Ranking from User ReviewsabstractThis paper attempts to conduct analysis for one certain type of user reviews; that is, the reviews on a super-entity (e.g., restaurant) involve descriptions for many sub-entities (e.g., dishes). To deal with such analysis, we propose a text embedding framework for ranking sub-entities from user reviews of a given super-entity. Experiments on two real-world datasets show that our method outperforms three baselines by a statistically significant amount. Intriguing cases from the experiments are discussed in the paper. Chih-Yu Chao, Yi-Fan Chu, Hsiu-Wei Yang, Chuan-Ju Wang, Ming-Feng Tsai |
CIKM | 4 |
| 2017 | ICE: Item Concept Embedding via Textual InformationabstractThis paper proposes an item concept embedding (ICE) framework to model item concepts via textual information. Specifically, in the proposed framework there are two stages: graph construction and embedding learning. In the first stage, we propose a generalized network construction method to build a network involving heterogeneous nodes and a mixture of both homogeneous and heterogeneous relations. The second stage leverages the concept of neighborhood proximity to learn the embeddings of both items and words. With the proposed carefully designed ICE networks, the resulting embedding facilitates both homogeneous and heterogeneous retrieval, including item-to-item and word-to-item retrieval. Moreover, as a distributed embedding approach, the proposed ICE approach not only generates related retrieval results but also delivers more diverse results than traditional keyword-matching-based approaches. As our experiments on two real-world datasets show, ICE encodes useful textual information and thus outperforms traditional methods in various item classification and retrieval tasks. Chuan-Ju Wang, Ting-Hsiang Wang, Hsiu-Wei Yang, Bo-Sin Chang, Ming-Feng Tsai |
SIGIR | 1 |
| 2016 | FIN10K: A Web-based Information System for Financial Report Analysis and VisualizationabstractIn this demonstration, we present FIN10K, a web-based information system that facilitates the analysis of textual information in financial reports. The proposed system has three main components: (1) a 10-K Corpus, including an inverted index of financial reports on Form 10-K, several numerical finance measures, and pre-trained word embeddings; (2) an information retrieval system; and (3) two data visualizations of the analyzed results. The system can be of great help in revealing valuable insights within large amounts of textual information. The system is now online available at http: //clip.csie.org/10K/. Yu-Wen Liu, Liang-Chih Liu, Chuan-Ju Wang, Ming-Feng Tsai |
CIKM | 3 |
| 2013 | Risk Ranking from Financial Reports
Ming-Feng Tsai, Chuan-Ju Wang |
ECIR | 2 |