VLDB 2026 Research / reviewers in the wild / expert
Ying-Jia Lin
dblp:257/6587
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0003-4347-0232ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MAPLE: Enhancing Review Generation with Multi-Aspect Prompt LEarning in Explainable RecommendationabstractChing-Wen Yang, Zhi-Quan Feng, Ying-Jia Lin, Che Wei Chen, Kun-da Wu, Hao Xu, Yao Jui-Feng, Hung-Yu Kao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ching-Wen Yang, Zhi-Quan Feng, Ying-Jia Lin, Che Wei Chen, Kun-da Wu, Jui-Feng Yao, Hung-Yu Kao |
ACL (1) | 3 |
| 2025 | From Persona to Person: Enhancing the Naturalness with Multiple Discourse Relations Graph Learning in Personalized Dialogue Generation
Chih-Hao Hsu, Ying-Jia Lin, Hung-Yu Kao |
PAKDD (2) | 2 |
| 2025 | Exploring the Effectiveness of Pre-training Language Models with Incorporation of Diglossia for Hong Kong ContentabstractIn this article, we present our works to create the first Hong Kong content-based public pre-training dataset and the experiments which resulted in the creation of ELECTRA-based models for commonly used languages in Hong Kong. The creation of pre-training dataset is required for us to study the effect of diglossia on Hong Kong language model, and this is the first ever study on the effect starting all the way from dataset creation phase. Our experiment shows that removing diglossia from pre-training data hurts model performance. We will release our data and models to encourage future studies in Hong Kong languages. 1 Yiu Cheong Yung, Ying-Jia Lin, Hung-Yu Kao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | CFEVER: A Chinese Fact Extraction and VERification DatasetabstractWe present CFEVER, a Chinese dataset designed for Fact Extraction and VERification. CFEVER comprises 30,012 manually created claims based on content in Chinese Wikipedia. Each claim in CFEVER is labeled as “Supports”, “Refutes”, or “Not Enough Info” to depict its degree of factualness. Similar to the FEVER dataset, claims in the “Supports” and “Refutes” categories are also annotated with corresponding evidence sentences sourced from single or multiple pages in Chinese Wikipedia. Our labeled dataset holds a Fleiss’ kappa value of 0.7934 for five-way inter-annotator agreement. In addition, through the experiments with the state-of-the-art approaches developed on the FEVER dataset and a simple baseline for CFEVER, we demonstrate that our dataset is a new rigorous benchmark for factual extraction and verification, which can be further used for developing automated systems to alleviate human fact-checking efforts. CFEVER is available at https://ikmlab.github.io/CFEVER. Ying-Jia Lin, Chia-Jen Yeh, Yi-Ting Li, Yun-Yu Hu, Chih-Hao Hsu, Mei-Feng Lee, Hung-Yu Kao |
AAAI | 1 |
| 2024 | Contrastive Learning for Unsupervised Sentence Embedding with False Negative Calibration
Chi-Min Chiu, Ying-Jia Lin, Hung-Yu Kao |
PAKDD (3) | 2 |
| 2024 | GViG: Generative Visual Grounding Using Prompt-Based Language Modeling for Visual Question Answering
Yi-Ting Li, Ying-Jia Lin, Chia-Jen Yeh, Hung-Yu Kao |
PAKDD (6) | 2 |
| 2023 | Improved Unsupervised Chinese Word Segmentation Using Pre-trained Knowledge and Pseudo-labeling TransferabstractUnsupervised Chinese word segmentation (UCWS) has made progress by incorporating linguistic knowledge from pre-trained language models using parameter-free probing techniques.However, such approaches suffer from increased training time due to the need for multiple inferences using a pre-trained language model to perform word segmentation.This work introduces a novel way to enhance UCWS performance while maintaining training efficiency.Our proposed method integrates the segmentation signal from the unsupervised segmental language model to the pre-trained BERT classifier under a pseudo-labeling framework.Experimental results demonstrate that our approach achieves state-of-the-art performance on the seven out of eight UCWS tasks while considerably reducing the training time compared to previous approaches. Hsiu-Wen Li, Ying-Jia Lin, Yi-Ting Li, Chun Lin, Hung-Yu Kao |
EMNLP | 2 |