VLDB 2026 Research / reviewers in the wild / expert
Chao-Lin Liu
dblp:23/2735
· DBLP profile ↗
45ranked-venue papers
24as first author
10since 2021 · last 2025
0000-0002-4093-1497ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 16 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 9 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning from the Judicial Journey: A Predictive Model for Case DisputabilityabstractTo aid litigants’ appeal strategies, this study introduces Disputability (δ), a novel metric quantifying a case’s contentiousness by the total number of judicial instances it undergoes. Using a dataset of 52,993 Taiwanese tax judgments, we test the hypothesis that similar cases share similar disputability levels. We compare Judgment-Level and Sentence-Level prediction models, finding that a Judgment-Level approach with multilingual-e5-base embeddings provides the strongest and most practical baseline. While the Sentence-Level model suffered from significant label noise, a high-precision kNN voting strategy proved particularly effective for identifying high-risk, disputable cases. Our work establishes an adaptable and extensible baseline for quantifying case disputability, offering a practical tool for computational law and legal analytics. Ho-Chien Huang, Chao-Lin Liu |
JURIX | 2 |
| 2025 | Detecting reading orders with block-based models for OCR for classical Chinese documents: concepts and demonstrationsabstractTexts, printed or written, are the most important form of cultural preservation. This is true for Chinese studies and many cultures. For this reason, optical character recognition (OCR) is one of the most important tools for heritage preservation. With OCR and reading order detection (ROD), we convert the information in their hard-copy forms to digitized documents, allowing both lasting storage and further analysis. Given a page image, the step of ROD helps us determine the order of the texts on the page. An effective ROD reduces the cost of post-OCR human labor for correction. Rule-based and n-gram models are conceivable and popular techniques for implementing ROD, considering the coordinates and the alternative orders of the individually detected characters. In this paper, we report a block-based approach that does not rely on the n-gram techniques for ROD. Instead, we identified the text blocks, and ordered the blocks based on the ideas of learning-to-rank, and concatenated the texts within the blocks by their orders. We evaluated our approach on the MTHv2 dataset. Experimental results show that our approach achieved impressive results, reducing the position error rate (PER) to 5.4%, compared to 61.7% for heuristic methods and 56.5% for MLP, representing a significant improvement of 91.3% and 90.4%, respectively. Additional experiments show that we could have used less data than one might expect to achieve a satisfactory quality of ROD. We implemented and demonstrated a fully functioning prototype at a recent AAAI conference. $$^{*}$$ The actual performance of the prototype provides solid evidence of the effectiveness and extensibility of our design to more general contexts, e.g., Japanese texts. Hsing-Yuan Ma, Chao-Lin Liu, Hen-Hsen Huang |
Multim. Tools Appl. | 2 |
| 2024 | Reading between the Lines: Image-Based Order Detection in OCR for Chinese Historical DocumentsabstractChinese historical documents, with their unique layouts and reading patterns, pose significant challenges for traditional Optical Character Recognition (OCR) systems. This paper introduces a tailored OCR system designed to address these complexities, particularly emphasizing the crucial aspect of Reading Order Detection(ROD). Our system operates through a threefold process: text detection using the Differential Binarization++ model, text recognition with the SVTR Net, and a novel ROD approach harnessing raw image features. This innovative method for ROD, inspired by human perception, utilizes visual cues present in raw images to deduce the inherent sequence of ancient texts. Preliminary results show promising reductions in page error rates. By preserving both content and context, our system contributes meaningfully to the accurate and contextual digitization of Chinese historical manuscripts. Hsing-Yuan Ma, Hen-Hsen Huang, Chao-Lin Liu |
AAAI | 3 |
| 2024 | Similar Phrases for Cause of Actions of Civil CasesabstractIn the Taiwanese judicial system, inconsistent Cause of Action (COA) titles complicate case classification. This paper presents a method to detect similar COAs using the Dice coefficient for cited legal acts and DBSCAN for clustering. We analyzed civil judgments with over 200 cases from district courts and improved classification accuracy by merging similar COAs. Our approach enhances case retrieval efficiency and can be adapted across different legal domains. Ho-Chien Huang, Chao-Lin Liu |
JURIX | 2 |
| 2023 | Audio-Driven Facial Landmark Generation in Violin Performance using 3DCNN Network with Self Attention ModelabstractIn a music scenario, both auditory and visual elements are essential to achieve an outstanding performance. Recent research has focused on the generation of body movements or fingering from audio in music performance. The audio-driven face generation technique in music performance is still deficient. In this paper, we compile a violin soundtrack and facial expression dataset (VSFE) for modeling facial expressions in violin performance. To our knowledge, this is the first dataset mapping the relationship between violin performance audio and musicians’ facial expressions. We then propose a 3DCNN network with self-attention and residual blocks for audio-driven facial expression generation. In the experiments, we compare our methods with three baselines on talking face generation. The codes and dataset are available on the Github (https://github.com/kevinlin91/icassp_music2face). Ting-Wei Lin, Chao-Lin Liu |
ICASSP | 2 |
| 2022 | Model AI Assignments 2022
Todd W. Neller, Jazmin Collins, Yim Register, Chia-Wei Tang, Chao-Lin Liu, Roozbeh Aliabadi, Annabel Hasty, Sultan Albarakati, Haotian Fang, Harvey Yin, Joel Wilson |
AAAI | 7 |
| 2022 | Using Wordle for Learning to Design and Compare StrategiesabstractWordle has become a very popular online game since November 2021. We designed and evaluated several strategies for solving Wordle in this paper. Our strategies achieved impressive performances in realistic evaluations that aimed to guess all of the known answers of the current Wordle. On average, we may solve a Wordle game with about 3.67 guesses, solve a Wordle game with six or fewer guesses higher than 98% of the time, and hit the answer with 2 or fewer guesses more than 5% of the time. In fact, our strategies are applicable to the word guessing games that are more general than the current Wordle. More importantly, we present our work in ways that our experiences may be used as classroom examples for learning to design strategies for computer games. Chao-Lin Liu |
CoG | 1 |
| 2022 | Toward an Integrated Annotation and Inference Platform for Enhancing Justifications for Algorithmically Generated Legal Recommendations and DecisionsabstractWe introduce our workflow that integrates the steps of annotation and classification, and hope that the end products are helpful for improving the justifications for legal reasoning and for recommending similar civil cases. Yi-Tang Huang, Hong-Ren Lin, Chao-Lin Liu |
JURIX | 3 |
| 2022 | Functional Classification of Statements of Chinese Judgment Documents of Civil CasesabstractEnabling the inference systems for assisting legal decisions to identify the functions of sentences and paragraphs in documents of legal judgments can enhance the justifiability of their algorithmic recommendations. The information about the functions of larger linguistic constituents complements the information at the word level like NER, and provides more clues about the arguments for the legal decisions. We explore this venue for the civil cases, which is a relatively uncommon choice in legal informatics and more challenging than working on the criminal cases. Current experimental results are promising. Chao-Lin Liu, Hong-Ren Lin, Wei-Zhi Liu, Chieh Yang |
JURIX | 1 |
| 2021 | HRRegionNet: Chinese Character Segmentation in Historical Documents with Regional Awareness
Chia-Wei Tang, Chao-Lin Liu, Po-Sen Chiu |
ICDAR (4) | 2 |
| 2020 | Multi-label Classification of Chinese Judicial Documents based on BERTabstractJudicial decisions are an important part of modern democratic societies. In this paper, I present results of multi-label classification of Chinese judicial documents. The experiments employ the same corpus that was used in Chinese AI & Law Challenge(CAIL) 2018. Mian Dai, Chao-Lin Liu |
IEEE BigData | 2 |
| 2020 | HRCenterNet: An Anchorless Approach to Chinese Character Segmentation in Historical DocumentsabstractThe information provided by historical documents has always been indispensable in the transmission of human civilization, but it has also made these books susceptible to damage due to various factors. Thanks to recent technology, the automatic digitization of these documents are one of the quickest and most effective means of preservation. The main steps of automatic text digitization can be divided into two stages, mainly: character segmentation and character recognition, where the recognition results depend largely on the accuracy of segmentation. Therefore, in this study, we will only focus on the character segmentation of historical Chinese documents. In this research, we propose a model named HRCenterNet, which is combined with an anchorless object detection method and parallelized architecture. The MTHv2 dataset consists of over 3000 Chinese historical document images and over 1 million individual Chinese characters; with these enormous data, the segmentation capability of our model achieves IoU 0.81 on average with the best speed-accuracy trade-off compared to the others. Our source code is available at https://github.com/Tverous/HRCenterNet. Chia-Wei Tang, Chao-Lin Liu, Po-Sen Chiu |
IEEE BigData | 2 |
| 2019 | Extracting the Gist of Chinese Judgments of the Supreme CourtabstractThe gist of judgement documents encodes important experience and viewpoints of the Supreme Court, and provides instrumental and educational information for judges, lawyers, practitioners, and students. The Supreme Court in Taiwan appoints senior members to produce the gist for selected judgments of the Supreme Court, but is unable to offer the gist for all judgment documents. Based on our observation of the existing gist statements, we can treat the generation of the gist as a sentence classification problem. We apply machine-learning methods, including gradient boosting, multilayer perceptrons, and deep learning methods with long short-term memory units; and consider legal, linguistic, statistical information, and different word embedding methods to build several classifiers. By using more sophisticated classifiers and more relevant features, we gradually achieved better results, and the best result was 0.9372 in F1 measure. Chao-Lin Liu, Kuan-Chun Chen |
ICAIL | 1 |
| 2019 | Synthesizing electronic health records using improved generative adversarial networksabstractObjective: The aim of this study was to generate synthetic electronic health records (EHRs). The generated EHR data will be more realistic than those generated using the existing medical Generative Adversarial Network (medGAN) method. Materials and Methods: We modified medGAN to obtain two synthetic data generation models-designated as medical Wasserstein GAN with gradient penalty (medWGAN) and medical boundary-seeking GAN (medBGAN)-and compared the results obtained using the three models. We used 2 databases: MIMIC-III and National Health Insurance Research Database (NHIRD), Taiwan. First, we trained the models and generated synthetic EHRs by using these three 3 models. We then analyzed and compared the models' performance by using a few statistical methods (Kolmogorov-Smirnov test, dimension-wise probability for binary data, and dimension-wise average count for count data) and 2 machine learning tasks (association rule mining and prediction). Results: We conducted a comprehensive analysis and found our models were adequately efficient for generating synthetic EHR data. The proposed models outperformed medGAN in all cases, and among the 3 models, boundary-seeking GAN (medBGAN) performed the best. Discussion: To generate realistic synthetic EHR data, the proposed models will be effective in the medical industry and related research from the viewpoint of providing better services. Moreover, they will eliminate barriers including limited access to EHR data and thus accelerate research on medical informatics. Conclusion: The proposed models can adequately learn the data distribution of real EHRs and efficiently generate realistic synthetic EHRs. The results show the superiority of our models over the existing model. Mrinal Kanti Baowaly, Chia-Ching Lin, Chao-Lin Liu, Kuan-Ta Chen |
J. Am. Medical Informatics Assoc. | 3 |
| 2017 | Exploring lexical, syntactic, and semantic features for Chinese textual entailment in NTCIR RITE evaluation tasks
Wei-Jie Huang, Chao-Lin Liu |
Soft Comput. | 2 |
| 2015 | Mining local gazetteers of literary Chinese with CRF and pattern based methods for biographical information in Chinese historyabstractPerson names and location names are essential building blocks for identifying events and social networks in historical documents that were written in literary Chinese. We take the lead to explore the research on algorithmically recognizing named entities in literary Chinese for historical studies with language-model based and conditional-random-field based methods, and extend our work to mining the document structures in historical documents. Practical evaluations were conducted with texts that were extracted from more than 220 volumes of local gazetteers (Difangzhi, $$$). Difangzhi is a huge and the single most important collection that contains information about officers who served in local government in Chinese history. Our methods performed very well on these realistic tests. Thousands of names and addresses were identified from the texts. A good portion of the extracted names match the biographical information currently recorded in the China Biographical Database (CBDB) of Harvard University, and many others can be verified by historians and will become as new additions to CBDB.1 Chao-Lin Liu, Chih-Kai Huang 0002, Hongsu Wang, Peter K. Bol |
IEEE BigData | 1 |
| 2015 | Toward Algorithmic Discovery of Biographical Information in Local Gazetteers of Ancient China
Chao-Lin Liu, Chih-Kai Huang 0002, Hongsu Wang, Peter K. Bol |
PACLIC | 1 |
| 2015 | Color Aesthetics and Social Networks in Complete Tang Poems: Explorations and Discoveries
Chao-Lin Liu, Hongsu Wang, Wen-Huei Cheng, Chu-Ting Hsu, Wei-Yun Chiu |
PACLIC | 1 |
| 2015 | Introduction to the Special Issue on Chinese Spell CheckingabstractThis special issue contains four articles based on and expanded from systems presented at the SIGHAN-7 Chinese Spelling Check Bakeoff. We provide an overview of the approaches and designs for Chinese spelling checkers presented in these articles. We conclude this introductory article with a summary of possible future directions. Lung-Hao Lee, Gina-Anne Levow, Shih-Hung Wu, Chao-Lin Liu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2015 | Intelligent interaction, reasoning, and applications
Chao-Lin Liu, Mitsunori Matsushita, Yasufumi Takama, Min-Yuh Day, Vincent S. Tseng |
Web Intell. | 1 |
| 2014 | Semantical Clustering of Morphologically Related Chinese WordsabstractA Chinese character embedded in different compound words may carry different meanings. In this paper, we aim at semantical clustering of a given family of morphologically related Chinese words. In Experiment 1, we employed linguistic features at the word, syntactic, semantic, and contextual levels in aggregated computational linguistics methods to handle the clustering task. In Experiment 2, we recruited adults and children to perform the clustering task. Experimental results indicate that our computational model achieved a similar level of performance as children. Chia-Ling Lee, Ya-Ning Chang, Chao-Lin Liu, Chia-Ying Lee, Yung-Jen Hsu 0001 |
AAAI | 3 |
| 2014 | Unsupervised Clustering of Morphologically Related Chinese Words
Chia-Ling Lee, Ya-Ning Chang, Chao-Lin Liu, Chia-Ying Lee, Yung-Jen Hsu 0001 |
CogSci | 3 |
| 2012 | A Cognition-Based Game Platform and its Authoring Environment for Learning Chinese Characters
Chao-Lin Liu, Chia-Ying Lee, Wei-Jie Huang, Yu-Lin Tzeng, Chia-Ru Chou |
ITS | 1 |
| 2011 | A Construction Grammar Approach to Prepositional Phrase Attachment: Semantic Feature Analysis of V NP1 into NP2 Construction
Liyin Chen, Siaw-Fong Chung, Chao-Lin Liu |
PACLIC | 3 |
| 2011 | Translating Common English and Chinese Verb-Noun Pairs in Technical Documents with Collocational and Bilingual Information
Yi-Hsuan Chuang, Chao-Lin Liu, Jing-Shin Chang |
PACLIC | 2 |
| 2011 | Visually and Phonologically Similar Characters in Incorrect Chinese Words: Analyses, Identification, and ApplicationsabstractInformation about students’ mistakes opens a window to an understanding of their learning processes, and helps us design effective course work to help students avoid replication of the same errors. Learning from mistakes is important not just in human learning activities; it is also a crucial ingredient in techniques for the developments of student models. In this article, we report findings of our study on 4,100 erroneous Chinese words. Seventy-six percent of these errors were related to the phonological similarity between the correct and the incorrect characters, 46% were due to visual similarity, and 29% involved both factors. We propose a computing algorithm that aims at replication of incorrect Chinese words. The algorithm extends the principles of decomposing Chinese characters with the Cangjie codes to judge the visual similarity between Chinese characters. The algorithm also employs empirical rules to determine the degree of similarity between Chinese phonemes. To show its effectiveness, we ran the algorithm to select and rank a list of about 100 candidate characters, from more than 5,100 characters, for the incorrectly written character in each of the 4,100 errors. We inspected whether the incorrect character was indeed included in the candidate list and analyzed whether the incorrect character was ranked at the top of the candidate list. Experimental results show that our algorithm captured 97% of incorrect characters for the 4,100 errors, when the average length of the candidate lists was 104. Further analyses showed that the incorrect characters ranked among the top 10 candidates in 89% of the phonologically similar errors and in 80% of the visually similar errors. Chao-Lin Liu, Min-Hua Lai, Kan-Wen Tien, Yi-Hsuan Chuang, Shih-Hung Wu |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2009 | Two Applications of Lexical Information to Computer-Assisted Item Authoring for Elementary Chinese
Chao-Lin Liu, Kan-Wen Tien, Yi-Hsuan Chuang, Chih-Bin Huang, Juei-Yu Weng |
IEA/AIE | 1 |
| 2006 | An Experience in Learning about Learning Composite ConceptsabstractStudents need to integrate multiple basic concepts to become competent in the activities that require the knowledge of the composite concept. Traditionally, we rely on experts' judgments to build models for this integration process. In this paper, we explore computational methods for unveiling how students learn composite concepts, and compare effects of applying mutual information-based and hierarchical search-based techniques for guessing the unobservable processes, which were simulated by Bayesian networks. Experimental results show that computational methods can be useful in assisting this student modelling task Chao-Lin Liu |
ICALT | 1 |
| 2006 | Learning Students' Learning Patterns with Support Vector Machines
Chao-Lin Liu |
ISMIS | 1 |
| 2006 | Exploring Phrase-Based Classification of Judicial Documents for Criminal Charges in Chinese
Chao-Lin Liu, Chwen-Dar Hsieh |
ISMIS | 1 |
| 2006 | Learning How Students Learn with Bayes Nets
Chao-Lin Liu |
Intelligent Tutoring Systems | 1 |
| 2006 | Using Genetic Algorithms for Feature Selection in Predicting Financial Distresses with Support Vector MachinesabstractFinancial distresses in corporations harm both individual investors and financial institutions, and can cause social problems due to the cascading effects. Government agencies, fund managers, and even small investors need to arm themselves with devices for predicting financial distresses in corporations, even when the financial problems are shadowed by malignant window dressing. In previous work, we explored the effectiveness of making predictions based on both financial ratios, including those proposed by Altman for the Z-score models and those used in the common-size analysis, and qualitative indicators, such as corporate governance. In this paper, we report results of our attempt to select the best features from the previously proposed features with genetic algorithms and gain ratio-based methods. Experimental results indicate that the selected features outperform the features used in the Z-score models. Not surprisingly, the genetic algorithms surpass the gain ratio-based methods in the task of feature selection. Pei-Wen Huang, Chao-Lin Liu |
SMC | 2 |
| 2006 | Learning Students' Learning Patterns with Neural ComputingabstractStudent modeling is a key ingredient for intelligent interaction with students, and understanding how students learn composite concepts can help us build models of higher quality. An interesting question is how machines can help us infer about students' learning processes of composite concepts from students' external performance. Since Bayesian networks have been used in student modeling in many research projects, we employ them for simulating students' performance in taking tests, and seek methods for learning the original networks based on the simulated students' performance. The problem is not easy because the data we can observe have only indirect and uncertain relationship with the variables for mastery levels, and we would like to know the relationships among these hidden variables. We applied mutual information and artificial neural networks for this learning problem, and we achieved 75% in accuracy even when the item responses have really uncertain relationships with the actual mastery levels. Chao-Lin Liu |
SMC | 1 |
| 2005 | Classifying Criminal Charges in Chinese for Web-Based Legal Services
Chao-Lin Liu, Ting-Ming Liao |
APWeb | 1 |
| 2005 | Computer-Assisted Item Generation for Listening Cloze Tests in EnglishabstractAs the steps of globalization accelerate, learning foreign languages has become a modern challenge for everyone. Obtaining a broad range of learning and practice material will boost the efficiency of language learning, and the Web serves as a rich source of text material. We offer methods for algorithmically creating test items that may meet needs of individual learners and instructors of English. At the current stage, we explore the generation of test items for students' practicing listening cloze in English, using text material obtained from the Web. Relying on the text corpus, a voice-synthesizer software, and linguistics-based criteria, our system identifies candidate sentences and selects distractors for composing test items for listening cloze. Teachers can select and compose the machine-generated items as they wish, and allow students to practice the composed items. In addition, the current system records histories of the performance of individual student, so the resulting system paves our way to adoptively supporting students' activities for polishing their competence in listening English. Shang-Ming Huang, Chao-Lin Liu, Zhao-Ming Gao |
ICALT | 2 |
| 2005 | Some Theoretical Properties of Mutual Information for Student Assessments in Intelligent Tutoring Systems
Chao-Lin Liu |
ISMIS | 1 |
| 2004 | Using Mutual Information for Adaptive Student AssessmentsabstractWhen we have an item bank available for assessing students, we would like to select the test items that can reveal students' knowledge levels as effectively and accurately as possible. This may not be a trivial task when we consider the fact that students' item-response patterns may not reflect their competence exactly. Theoretical and experimental results reported in this paper indicate that mutual information between test items and educational targets provides a principled and instrumental basis for adaptively selecting test items for student assessments. Chao-Lin Liu |
ICALT | 1 |
| 2004 | Bounding probabilistic relationships in Bayesian networks using qualitative influences: methods and applications
Chao-Lin Liu, Michael P. Wellman |
Int. J. Approx. Reason. | 1 |
| 2003 | Classification and Clustering for Case-Based Criminal Summary JudgementabstractWe investigate the effectiveness of machine-generated criteria for classification problems related to criminal summary judgments. Our system utilizes documents of closed lawsuits as training data for generating keyword-based and case-based classification criteria, and applies these machine-generated criteria for the classification tasks. To construct databases of the classification criteria, we employ different levels of lexical knowledge in extracting information from legal documents in Chinese, and build a case instance for each closed lawsuit. Experimental results indicate that case-based classification outperforms keyword-based classification, and that machine-generated cases may offer performance accuracy that is about 7% below that of human-provided cases. Hoping to boost inference efficiency of our classifiers, we also design methods that merge the machine-generated criteria. Empirical results show that our methods can maintain the classification quality within 20% of the quality achieved by human-provided cases, even when we aggressively reduce the number of previously machine-generated cases by about seventy percents. Chao-Lin Liu, Cheng-Tsung Chang, Jim-How Ho |
ICAIL | 1 |
| 2003 | Some Case-Refinement Strategies for Case-Based Criminal Summary Judgments
Chao-Lin Liu, Tseng-Chung Chang |
ISMIS | 1 |
| 2003 | Traffic Sign Recognition in Disturbing Environments
Hsiu-Ming Yang, Chao-Lin Liu, Kun-Hao Liu, Shang-Ming Huang |
ISMIS | 2 |
| 2002 | Evaluation of Bayesian networks with flexible state-space abstraction methods
Chao-Lin Liu, Michael P. Wellman |
Int. J. Approx. Reason. | 1 |
| 1998 | Incremental Tradeoff Resolution in Qualitative Probabilistic Networks
Chao-Lin Liu, Michael P. Wellman |
UAI | 1 |
| 1998 | Using Qualitative Relationships for Bounding Probability Distributions
Chao-Lin Liu, Michael P. Wellman |
UAI | 1 |
| 1994 | State-Space Abstraction for Anytime Evaluation of Probabilistic Networks
Michael P. Wellman, Chao-Lin Liu |
UAI | 2 |