Yunda Tsai

dblp:164/8914 · also Yun-Da Tsai · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
9since 2021 · last 2026
0009-0009-5385-4660ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
abstract
Pei-Fu Guo, Yun-Da Tsai, Chun-Chia Hsu, Kai-Xin Chen, Ya An Tsai, Kai-Wei Chang, Nanyun Peng, Mi-Yen Yeh, Shou-De Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pei-Fu Guo, Yunda Tsai, Chun-Chia Hsu, Kai-Xin Chen, Ya-An Tsai, Kai-Wei Chang 0001, Nanyun Peng 0001, Mi-Yen Yeh, Shou-De Lin
ACL (1)2
2025 Enhance Modality Robustness in Text-Centric Multimodal Alignment with Adversarial Prompting
abstract
Converting different modalities into generalized text, which then serves as input prompts for large language models (LLMs), is a common approach for aligning multimodal models, particularly when pairwise data is limited. Text-centric alignment method leverages the unique properties of text as a modality space, transforming diverse inputs into a unified textual representation, thereby enabling downstream models to effectively interpret various modal inputs. This study evaluates the quality and robustness of multimodal representations in the face of noise imperfections, dynamic input order permutations, and missing modalities, revealing that current text-centric alignment methods can compromise downstream robustness. To address this issue, we propose a new text-centric adversarial training approach that significantly enhances robustness compared to traditional robust training methods and pre-trained multimodal foundation models. Our findings underscore the potential of this approach to improve the robustness and adaptability of multimodal representations, offering a promising solution for dynamic and real-world applications.
Yunda Tsai, Ting-Yu Yen, Keng-Te Liao 0001, Shou-De Lin
AAAI1
2025 CraftRTL: High-quality Synthetic Data Generation for Verilog Code Models with Correct-by-Construction Non-Textual Representations and Targeted Code Repair
abstract
Despite the significant progress made in code generation with large language models, challenges persist, especially with hardware description languages such as Verilog. This paper first presents an analysis of fine-tuned LLMs on Verilog coding, with synthetic data from prior methods. We identify two main issues: difficulties in handling non-textual representations (Karnaugh maps, state-transition diagrams and waveforms) and significant variability during training with models randomly making ''minor'' mistakes. To address these limitations, we enhance data curation by creating correct-by-construction data targeting non-textual representations. Additionally, we introduce an automated framework that generates error reports from various model checkpoints and injects these errors into open-source code to create targeted code repair data. Our fine-tuned Starcoder2-15B outperforms prior state-of-the-art results by 3.8\%, 10.9\%, 6.6\% for pass@1 on VerilogEval-Machine, VerilogEval-Human, and RTLLM.
Yunda Tsai, Wenfei Zhou, Haoxing Ren
ICLR2
2025 Differentiable Good Arm Identification
Yunda Tsai, Tzu-Hsien Tsai, Shou-De Lin
PAKDD (1)1
2025 LinearAPT: An Adaptive Algorithm for the Fixed-Budget Thresholding Linear Bandit Problem
Yun-Ang Wu, Yunda Tsai, Shou-De Lin
PAKDD (1)2
2024 Toward More Generalized Malicious URL Detection Models
abstract
This paper reveals a data bias issue that can profoundly hinder the performance of machine learning models in malicious URL detection. We describe how such bias can be diagnosed using interpretable machine learning techniques and further argue that such biases naturally exist in the real world security data for training a classification model. To counteract these challenges, we propose a debiased training strategy that can be applied to most deep-learning based models to alleviate the negative effects of the biased features. The solution is based on the technique of adversarial training to train deep neural networks learning invariant embedding from biased data. Through extensive experimentation, we substantiate that our innovative strategy fosters superior generalization capabilities across both CNN-based and RNN-based detection models. The findings presented in this work not only expose a latent issue in the field but also provide an actionable remedy, marking a significant step forward in the pursuit of more reliable and robust malicious URL detection.
Yunda Tsai, Cayon Liow, Yin Sheng Siang, Shou-De Lin
AAAI1
2024 RTLFixer: Automatically Fixing RTL Syntax Errors with Large Language Model
abstract
This paper presents RTLFixer, a novel framework enabling automatic syntax errors fixing for Verilog code with Large Language Models (LLMs). Despite LLM's promising capabilities, our analysis indicates that approximately 55% of errors in LLM-generated Verilog are syntax-related, leading to compilation failures. To tackle this issue, we introduce a novel debugging framework that employs Retrieval-Augmented Generation (RAG) and ReAct prompting, enabling LLMs to act as autonomous agents in interactively debugging the code with feedback. This framework demonstrates exceptional proficiency in resolving syntax errors, successfully correcting about 98.5% of compilation errors in our debugging dataset, comprising 212 erroneous implementations derived from the VerilogEval benchmark. Our method leads to 32.3% and 10.1% increase in pass@1 success rates in the VerilogEval-Machine and VerilogEval-Human benchmarks, respectively. The source code and benchmark are available at https://github.com/NVlabs/RTLFixer.
Yunda Tsai, Haoxing Ren
DAC1
2024 lil'HDoC: An Algorithm for Good Arm Identification Under Small Threshold Gap
Tzu-Hsien Tsai, Yunda Tsai, Shou-De Lin
PAKDD (5)2
2021 Toward an Effective Black-Box Adversarial Attack on Functional JavaScript Malware against Commercial Anti-Virus
abstract
Machine learning has been a rising technique in signatureless malware detection and is popular in the anti-virus industry. Despite the powerful ability of machine learning, it is known to be vulnerable to attack by injecting specially crafted input noise (adversarial example). In this paper, we develop a systematic attack method that is effective, general and also efficient which automatically generates functional malware. Experiment results showed that such adversarial malware could deceive commercial anti-virus and completely defeat learning-based malware detector provided by a well-known anti-virus vendor. We further examine the effectiveness of our approach on multiple anti-virus engines on VirusTotal and investigate the transferability of our proposed method between different features and classification algorithms. Finally, we show how our attack could resist JavaScript de-obfuscation techniques.
Yunda Tsai, Cheng-Kuan Chen, Shou-De Lin
CIKM1
2016 Minimizing Radio Resource Usage for Machine-to-Machine Communications through Data-Centric Clustering
abstract
While clustered communication has been considered as one key technology for wireless sensor networks, existing work on cluster formation predominantly takes a pure graph-theoretic approach with the goal of optimizing the performance of individual machines. Since the radio resource available for M2M communications is typically limited yet the amount of data to transport is large, such “resource-agnostic” and “data-agnostic” clustering techniques could lead to sub-optimal performance. To address this problem, we propose “data-centric” clustering in a resource-constrained M2M network by prioritizing the quality of overall data over the performance of individual machines. We first formulate an optimization problem to minimize the amount of radio resource needed for supporting two-tier clustered communications. We then partition the formulated problem into the inner power control and outer cluster formation sub-problems and propose algorithms for solving the problems. While power control can be optimally solved for any given cluster structure by the proposed algorithm, cluster formation is an NP-hard problem. Hence, we propose an anytime, guided, stochastic search algorithm to find a reasonably good cluster structure without incurring prohibitive computation complexity. Compared with baseline approaches, our evaluation results show that data-centric clustering can achieve noticeable performance gain by selecting only important machines and forming a cluster structure that can balance the radio resource usage of the two tiers. We therefore motivate data-centric clustering as a promising communication model for resource-constrained M2M networks.
Hung-Yun Hsieh, Tzu-Chuan Juan, Yunda Tsai, Hong-Chen Huang
IEEE Trans. Mob. Comput.3
2015 Correlation-aware machine selection for M2M data gathering in cellular networks
abstract
In machine-to-machine communications, several machines located close to one another can form a cluster to leverage spatial reuse gains and energy savings. We investigate a machine selection algorithm where only a machine in each cluster transfers its data to the cellular user via device-to-device communication and then the cellular user forwards it to the base station. Unlike the previous works that focus on data rate or channel condition as the selection metrics, this paper proposes a correlation-aware selection algorithm which maximizes the joint entropy extracted from the data delivered from the selected machines. Our evaluation results show the proposed selection scheme reaps a significant gain and is near-optimal.
Hojin Song, Hung-Yun Hsieh, Yunda Tsai, Wan Choi 0001
PIMRC3
2015 Joint Optimization of Clustering and Scheduling for Machine-to-Machine Communications in Cellular Wireless Networks
abstract
Most related work on cluster formation has modeled the cluster structure from a purely graph-theoretic perspective with predefined communication links and data rates among the set of nodes under consideration. Failure to accurately model inter- cluster interference among concurrent transmissions, however, results in sub-optimal performance for supporting M2M communications in cellular networks with limited radio resource and tight interference control. In this paper, we investigate the problem of energy-efficient clustering by jointly considering cluster formation, transmission scheduling, and power control. We specifically take the communication constraint into consideration to ensure the subset of machines scheduled and the transmission powers used allow concurrent, reliable transmissions in the cluster structure. Based on the evaluation for a semi- realistic camera surveillance network, the proposed approach can result in minimum powers for tier-1 and tier-2 transmissions while collecting the required amount of data for M2M applications.
Yunda Tsai, Chang-Yu Song, Hung-Yun Hsieh
VTC Spring1