Ying-Jia Lin

dblp:257/6587 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0003-4347-0232ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 MAPLE: Enhancing Review Generation with Multi-Aspect Prompt LEarning in Explainable Recommendation
abstract
Ching-Wen Yang, Zhi-Quan Feng, Ying-Jia Lin, Che Wei Chen, Kun-da Wu, Hao Xu, Yao Jui-Feng, Hung-Yu Kao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ching-Wen Yang, Zhi-Quan Feng, Ying-Jia Lin, Che Wei Chen, Kun-da Wu, Jui-Feng Yao, Hung-Yu Kao
ACL (1)3
2025 From Persona to Person: Enhancing the Naturalness with Multiple Discourse Relations Graph Learning in Personalized Dialogue Generation
Chih-Hao Hsu, Ying-Jia Lin, Hung-Yu Kao
PAKDD (2)2
2025 Exploring the Effectiveness of Pre-training Language Models with Incorporation of Diglossia for Hong Kong Content
abstract
In this article, we present our works to create the first Hong Kong content-based public pre-training dataset and the experiments which resulted in the creation of ELECTRA-based models for commonly used languages in Hong Kong. The creation of pre-training dataset is required for us to study the effect of diglossia on Hong Kong language model, and this is the first ever study on the effect starting all the way from dataset creation phase. Our experiment shows that removing diglossia from pre-training data hurts model performance. We will release our data and models to encourage future studies in Hong Kong languages. 1
Yiu Cheong Yung, Ying-Jia Lin, Hung-Yu Kao
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2024 CFEVER: A Chinese Fact Extraction and VERification Dataset
abstract
We present CFEVER, a Chinese dataset designed for Fact Extraction and VERification. CFEVER comprises 30,012 manually created claims based on content in Chinese Wikipedia. Each claim in CFEVER is labeled as “Supports”, “Refutes”, or “Not Enough Info” to depict its degree of factualness. Similar to the FEVER dataset, claims in the “Supports” and “Refutes” categories are also annotated with corresponding evidence sentences sourced from single or multiple pages in Chinese Wikipedia. Our labeled dataset holds a Fleiss’ kappa value of 0.7934 for five-way inter-annotator agreement. In addition, through the experiments with the state-of-the-art approaches developed on the FEVER dataset and a simple baseline for CFEVER, we demonstrate that our dataset is a new rigorous benchmark for factual extraction and verification, which can be further used for developing automated systems to alleviate human fact-checking efforts. CFEVER is available at https://ikmlab.github.io/CFEVER.
Ying-Jia Lin, Chia-Jen Yeh, Yi-Ting Li, Yun-Yu Hu, Chih-Hao Hsu, Mei-Feng Lee, Hung-Yu Kao
AAAI1
2024 Contrastive Learning for Unsupervised Sentence Embedding with False Negative Calibration
Chi-Min Chiu, Ying-Jia Lin, Hung-Yu Kao
PAKDD (3)2
2024 GViG: Generative Visual Grounding Using Prompt-Based Language Modeling for Visual Question Answering
Yi-Ting Li, Ying-Jia Lin, Chia-Jen Yeh, Hung-Yu Kao
PAKDD (6)2
2023 Improved Unsupervised Chinese Word Segmentation Using Pre-trained Knowledge and Pseudo-labeling Transfer
abstract
Unsupervised Chinese word segmentation (UCWS) has made progress by incorporating linguistic knowledge from pre-trained language models using parameter-free probing techniques.However, such approaches suffer from increased training time due to the need for multiple inferences using a pre-trained language model to perform word segmentation.This work introduces a novel way to enhance UCWS performance while maintaining training efficiency.Our proposed method integrates the segmentation signal from the unsupervised segmental language model to the pre-trained BERT classifier under a pseudo-labeling framework.Experimental results demonstrate that our approach achieves state-of-the-art performance on the seven out of eight UCWS tasks while considerably reducing the training time compared to previous approaches.
Hsiu-Wen Li, Ying-Jia Lin, Yi-Ting Li, Chun Lin, Hung-Yu Kao
EMNLP2