Yu-Che Tsai

dblp:65/2921 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
5since 2021 · last 2024
0009-0008-1053-7609ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model Queries
abstract
This study investigates the privacy risks associated with text embeddings, focusing on the scenario where attackers cannot access the original embedding model.Contrary to previous research requiring direct model access, we explore a more realistic threat model by developing a transfer attack method.This approach uses a surrogate model to mimic the victim model's behavior, allowing the attacker to infer sensitive information from text embeddings without direct access.Our experiments across various embedding models and a clinical dataset demonstrate that our transfer attack significantly outperforms traditional methods, revealing the potential privacy vulnerabilities in embedding technologies and emphasizing the need for enhanced security measures.
Yu-Hsiang Huang, Yu-Che Tsai, Hsiang Hsiao, Hong-Yi Lin, Shou-De Lin
ACL (1)2
2023 GraphFC: Customs Fraud Detection with Label Scarcity
abstract
Customs officials across the world encounter huge volumes of transactions. Associated with customs transactions is customs fraud-the intentional manipulation of goods declarations to avoid taxes and duties. Due to limited manpower, the customs offices can only manually inspect a small number of declarations, necessitating the automation of customs fraud detection by machine learning techniques. The limited availability of manually inspected ground truth data makes it essential for the ML approach to generalize well on unseen data. However, current customs fraud detection models are not well suited or designed for this setting. In this work, we propose GraphFC (Graph Neural networks for Customs Fraud), a model-agnostic, domain-specific, graph neural network based customs fraud detection model that is designed to work in a real-world setting with limited ground truth data. Extensive experimentation using real customs data from two countries demonstrates that GraphFC generalizes well over unseen data and outperforms various baselines and other models by a large margin.
Karandeep Singh, Yu-Che Tsai, Cheng-Te Li, Meeyoung Cha, Shou-De Lin
CIKM2
2023 Graph Neural Networks for Tabular Data Learning
abstract
Deep learning-based approaches to Tabular Data Learning (TDL) have shown promising performance compared to their conventional counterparts. However, these methods often fail to account for the latent correlation among data instances and feature values. Recently, graph neural networks (GNNs) have gained attention across various application domains, including TDL, for their ability to model relations and interactions between different data entities. By creating appropriate graph structures from the input tabular data and employing GNNs for learning, the performance of TDL can be improved significantly. In this tutorial, we systematically introduce the methodologies of designing and applying GNNs to TDL. Our discussion covers the foundations and overview of GNN-based TDL methods, with a focus on formulating TDL as different graph structures. We also provide a comprehensive taxonomy of constructing graph structures and representation learning in GNN-based TDL methods. We describe the TDL model training framework, which includes different auxiliary tasks and supports open-world learning. Additionally, we discuss how to apply GNNs to various TDL application scenarios and tasks. Finally, we outline the limitations of current research and future directions for this field.
Cheng-Te Li, Yu-Che Tsai, Jay Chiehen Liao
ICDE2
2023 FinGAT: Financial Graph Attention Networks for Recommending Top-$K$K Profitable Stocks
abstract
Financial technology (FinTech) has drawn much attention among investors and companies. While conventional stock analysis in FinTech targets at predicting stock prices, less effort is made for profitable stock recommendation. Besides, in existing approaches on modeling time series of stock prices, the relationships among stocks and sectors (i.e., categories of stocks) are either neglected or pre-defined. Ignoring stock relationships will miss the information shared between stocks while using pre-defined relationships cannot depict the latent interactions or influence of stock prices between stocks. In this work, we aim at recommending the top-K profitable stocks in terms of return ratio using time series of stock prices and sector information. We propose a novel deep learning-based model, Financial Graph Attention Networks (FinGAT), to tackle the task under the setting that no pre-defined relationships between stocks are given. The idea of FinGAT is three-fold. First, we devise a hierarchical learning component to learn short-term and long-term sequential patterns from stock time series. Second, a fully-connected graph between stocks and a fully-connected graph between sectors are constructed, along with graph attention networks, to learn the latent interactions among stocks and sectors. Third, a multi-task objective is devised to jointly recommend the profitable stocks and predict the stock movement. Experiments conducted on Taiwan Stock, S&P 500, and NASDAQ datasets exhibit remarkable recommendation performance of our FinGAT, comparing to state-of-the-art methods.
Yi-Ling Hsu, Yu-Che Tsai, Cheng-Te Li
IEEE Trans. Knowl. Data Eng.2
2021 CARE: learning convolutional attentional recurrent embedding for sequential recommendation
abstract
Top-N sequential recommendation is to predict the next few items based on user's sequential interactions with past items. This paper aims at boosting the performance of top-N sequential recommendation based on a state-of-the-art model, Caser. We point out three insufficiencies of Caser - do not model variant-sized sequential patterns, treating the impact of each past time step equally, and cannot learn cumulative features. Then we propose a novel Convolutional Attentional Recurrent Embedding (CARE) learning model. Experiments conducted on a large-scale user-location check-in dataset exhibit promising performance, comparing to Caser.
Yu-Che Tsai, Cheng-Te Li
ASONAM1
2020 DATE: Dual Attentive Tree-aware Embedding for Customs Fraud Detection
abstract
Intentional manipulation of invoices that lead to undervaluation of trade goods is the most common type of customs fraud to avoid ad valorem duties and taxes. To secure government revenue without interrupting legitimate trade flows, customs administrations around the world strive to develop ways to detect illicit trades. This paper proposes DATE, a model of Dual-task Attentive Tree-aware Embedding, to classify and rank illegal trade flows that contribute the most to the overall customs revenue when caught. The strength of DATE comes from combining a tree-based model for interpretability and transaction-level embeddings with dual attention mechanisms. To accurately identify illicit transactions and predict tax revenue, DATE learns simultaneously from illicitness and surtax of each transaction. With a five-year amount of customs import data with a test illicit ratio of 2.24%, DATE shows a remarkable precision of 92.7% on illegal cases and a recall of 49.3% on revenue after inspecting only 1% of all trade flows. We also discuss issues on deploying DATE in Nigeria Customs Service, in collaboration with the World Customs Organization.
Sundong Kim, Yu-Che Tsai, Karandeep Singh, Yeonsoo Choi, Etim Ibok, Cheng-Te Li, Meeyoung Cha
KDD2
2019 Query Embedding Learning for Context-based Social Search
abstract
Recommending individuals through keywords is an essential and common search task in online social platforms such as Facebook and LinkedIn. However, it is often that one has only the impression about the desired targets, depicted by labels of social contexts (e.g. gender, interests, skills, visited locations, employment, etc). Assume each user is associated a set of labels, we propose a novel task, Search by Social Contexts (SSC), in online social networks. SSC is a kind of query-based people recommendation, recommending the desired target based on a set of user-specified query labels. We develop the method Social Query Embedding Learning (SQEL) to deal with SSC. SQEL aims to learn the feature representation (i.e., embedding vector) of the query, along with user feature vectors derived from graph embedding, and use the learned query vectors to find the targets via similarity. Experiments conducted on Facebook and Twitter datasets exhibit satisfying accuracy and encourage more advanced efforts on search by social contexts.
Yu-Che Tsai, Cheng-Te Li
CIKM2
2019 FineNet: a joint convolutional and recurrent neural network model to forecast and recommend anomalous financial items
abstract
Financial technology (FinTech) draws much attention in these years, with the advances of machine learning and deep learning. In this work, given historical time series of stock prices of companies, we aim at forecasting upcoming anomalous financial items, i.e., abrupt soaring or diving stocks, in financial time series, and recommending the corresponding stocks to support financial operations. We propose a novel joint convolutional and recurrent neural network model, Financial Event Neural Network (FineNet), to forecast and recommend anomalous stocks. Experiments conducted on the time series of stock prices of 300 well-known companies exhibit the promising performance of FineNet in terms of precision and recall. We build FineNet as a Web platform for live demonstration.
Yu-Che Tsai, Chih-Yao Chen, Shao-Lun Ma, Pei-Chi Wang, You-Jia Chen, Yu-Chieh Chang, Cheng-Te Li
RecSys1
2009 An Optimal Solution for the Heterogeneous Multiprocessor Single-Level Voltage-Setup Problem
abstract
A heterogeneous multiprocessor (HeMP) system consists of several heterogeneous processors, each of which is specially designed to deliver the best energy-saving performance for a particular category of applications. A low-power real-time scheduling algorithm is required to schedule tasks on such a system to minimize its energy consumption and complete all tasks by their deadlines. Existing works assume that processor speeds are known asa prioriand cannot deliver the optimal energy-saving performance. The problem of determining the optimal voltage for each processor to minimize the total energy consumption is called a voltage-setup problem. To the best of our knowledge, this is the first paper to propose the optimal solution for the HeMP single-level voltage-setup problem. This paper provides an optimal solution for the HeMP single-level voltage-setup problem. We first formulate the problem as a nonlinear generalized assignment problem that has been proved to be nondeterministic polynomial-time hard (NP-hard). We next develop a pruning-based algorithm to obtain the optimal solution. A heuristic algorithm is also proposed to derive an approximate solution. After obtaining the optimal partition, each processor's speed is determined by its final workload. In our simulations, we model more than a couple dozens of off-the-shelf embedded processors including ARM processor and TI DSP. The results show that the pruning-based algorithm reduces the time needed to derive the optimal solution by at least 98%, compared with the exhaustive search. Also, our heuristic algorithm achieves the minimum energy consumption over existing works.
Edward T.-H. Chu, Yu-Che Tsai
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2008 An Efficient Real-Time Disk-Scheduling Framework with Adaptive Quality Guarantee
abstract
A multimedia server requires a real-time disk-scheduling algorithm to deliver isochronous data for real-time streams. Traditional disk-scheduling algorithms focus on providing good quality in a best effort manner. In this paper, we propose a novel real-time disk-scheduling algorithm called WRR-SCAN (weighted-round-robin-SCAN) to provide quality guarantees for all in-service streams encoded at variable bit rates and bounded response times for aperiodic jobs. WRR-SCAN divides a real-time stream into guaranteed jobs and optional jobs. The admission control admits a stream as long as its guaranteed jobs are satisfied. Such a decision is made in 0(1) time as WRR-SCAN reserves a fixed weight for each stream. WRR-SCAN incorporates an aggressive policy to dynamically reclaim unused bandwidth during runtime. The reclaimed bandwidth is used to serve optional jobs or more aperiodic jobs. We conducted a set of simulations to compare WRR-SCAN with a set of referred disk-scheduling algorithms. The evaluations are conducted on a commonly used disk simulator with traces from a real multimedia server. The experimental results show that WRR-SCAN provides significantly better quality for real-time streams and yields considerably shorter response times for aperiodic jobs.
Cheng-Han Tsai, Edward T.-H. Chu, Chun-Hang Wei, Yu-Che Tsai
IEEE Trans. Computers5
2007 A Near-optimal Solution for the Heterogeneous Multi-processor Single-level Voltage Setup Problem
abstract
A heterogeneous multi-processor (HeMP) system consists of several heterogeneous processors, each of which is specially designed to deliver the best energy-saving performance for a particular category of applications. A low-power real-time scheduling algorithm is required to schedule tasks on such a system to minimize its energy consumption and complete all tasks by their deadline. The problem of determining the optimal speed for each processor to minimize the total energy consumption is called the voltage setup problem. This paper provides a near-optimal solution for the HeMP single-level voltage setup problem. To our best knowledge, we are the first work that addresses this problem. Initially, each task is assigned to a processor in a local-optimal manner. We next propose a couple of solutions to reduce energy by migrating tasks between processors. Finally, we determine each processor's speed by its final workload and the deadline. We conducted a series of simulations to evaluate our algorithms. The results show that the local-optimal partition leads to a considerably better energy-saving schedule than a commonly-used homogeneous multi-processor scheduling algorithm. Furthermore, at all measurable configurations, our energy consumption is at most 3% more than the optimal value obtained by an exhaustive iteration of all possible task-to-processor assignments. In summary, our work is shown to provide a near-optimal solution at its polynomial-time complexity.
Yu-Che Tsai, Edward T.-H. Chu
IPDPS2