VLDB 2026 Research / reviewers in the wild / expert
Tien-Hsuan Wu
dblp:228/2551
· DBLP profile ↗
11ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0003-0212-8107ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GenAI for Social Work Field Education: Client Simulation with Real-Time Feedback
James Sungarda, Hongkai Liu, Tien-Hsuan Wu, Johnson Chun-Sing Cheung, Ben Kao |
IEEE Big Data | 4 |
| 2025 | QBR - A Question-Bank-Based Approach to Fine-Grained Legal Knowledge Retrieval for the General PublicabstractRetrieval of legal knowledge by the general public is a challenging problem due to the technicality of the professional knowledge and the lack of fundamental understanding by laypersons on the subject. Traditional information retrieval techniques assume that users are capable of formulating succinct and precise queries for effective document retrieval. In practice, however, the wide gap between the highly technical contents and untrained users makes legal knowledge retrieval very difficult. We propose a methodology, called QBR, which employs a Questions Bank (QB) as an effective medium for bridging the knowledge gap. We show how the QB is used to derive training samples to enhance the embedding of knowledge units within documents, which leads to effective fine-grained knowledge retrieval. We discuss and evaluate through experiments various advantages of QBR over traditional methods. These include more accurate, efficient, and explainable document retrieval, better comprehension of retrieval results, and highly effective fine-grained knowledge retrieval. We also present some case studies and show that QBR achieves social impact by assisting citizens to resolve everyday legal concerns. Mingruo Yuan, Ben Kao, Tien-Hsuan Wu |
IJCAI | 3 |
| 2024 | A Multi-Stage Prompting and RAG Approach to Generating Legal Analysis in Common Law SystemsabstractWe explore the use of multi-stage prompting with retrieval-augmented generation (RAG) in generating legal analysis. We break the complex legal problems into small steps, including initial analysis, legal query formulation, legal sources retrieval, legal reasoning, result consolidation, and output generation. The design of this framework emulates the process a legal professional in a common law system approaches the problem. Titus T. H. Ng, Tien-Hsuan Wu, Benjamin Minhao Chen, Yongxi Chen, Ben Kao |
JURIX | 2 |
| 2023 | CEMA - Cost-Efficient Machine-Assisted Document AnnotationsabstractWe study the problem of semantically annotating textual documents that are complex in the sense that the documents are long, feature rich, and domain specific. Due to their complexity, such annotation tasks require trained human workers, which are very expensive in both time and money. We propose CEMA, a method for deploying machine learning to assist humans in complex document annotation. CEMA estimates the human cost of annotating each document and selects the set of documents to be annotated that strike the best balance between model accuracy and human cost. We conduct experiments on complex annotation tasks in which we compare CEMA against other document selection and annotation strategies. Our results show that CEMA is the most cost-efficient solution for those tasks. Guowen Yuan, Ben Kao, Tien-Hsuan Wu |
AAAI | 3 |
| 2023 | Judgment Retrieval Made Easier Through Query AnalysisabstractThe Hong Kong Legal Information Institute (HKLII) provides a repository of legal documents in Hong Kong and such as ordinances and historical court judgments. HKLII provides a search facility through which users retrieve relevant documents. We perform statistical analysis on HKLII access log over a 5-year period categorizing user search queries and discovering interesting user access patterns. Based on this study, we propose enhancing user experience through identifying the search intent of the user. We classify a user query into one of several types and customize the search behavior based on the query type to enhance user experience. Our study provides an example of leveraging log analysis in legal information system design. Tien-Hsuan Wu, Ben Kao, Michael M. K. Cheung |
JURIX | 1 |
| 2022 | Judgment Tagging and Recommendation Using Pre-Trained Language Models and Legal TaxonomyabstractWe study the problem of machine comprehension of court judgments and generation of descriptive tags for judgments. Our approach makes use of a legal taxonomy D, which serves as a dictionary of canonicalized legal concepts. Given a court judgment J, our method identifies the key contents of J and then applies Word2Vec and BERT-based models to select a short list TJ of terms/phrases from the taxonomy D as descriptive tags of J. The tag set TJ suggests concepts that are relevant to or associative with J and provides a simple mechanism for readers of J to compose associative queries for effective judgment recommendation. Our prototype system implemented on the Hong Kong Legal Information Institute (HKLII) platform shows that our method provides a highly effective tool that assists users in exploring a judgment corpus and in obtaining relevant judgment recommendation. Tien-Hsuan Wu, Ben Kao, Henry W. H. Chan, Michael M. K. Cheung |
JURIX | 1 |
| 2021 | Semantic Search and Summarization of Judgments Using Topic ModelingabstractOnline legal document libraries, such as WorldLII, are indispensable tools for legal professionals to conduct legal research. We study how topic modeling techniques can be applied to such platforms to facilitate searching of court judgments. Specifically, we improve search effectiveness by matching judgments to queries at semantics level rather than at keyword level. Also, we design a system that summarizes a retrieved judgment by highlighting a small number of paragraphs that are semantically most relevant to the user query. This summary serves two purposes: (1) It explains to the user why the machine finds the retrieved judgment relevant to the user’s query, and (2) it helps the user quickly grasp the most salient points of the judgment, which significantly reduces the amount of time needed by the user to go through the returned search results. We further enhance our system by integrating domain knowledge provided by legal experts. The knowledge includes the features and aspects that are most important for a given category of judgments. Users can then view a judgement’s summary focusing on particular aspects only. We illustrate the effectiveness of our techniques with a user evaluation experiment on the HKLII platform. The results show that our methods are highly effective. Tien-Hsuan Wu, Ben Kao, Felix Chan, Anne S. Y. Cheung, Michael M. K. Cheung, Guowen Yuan, Yongxi Chen |
JURIX | 1 |
| 2020 | Integrating Domain Knowledge in AI-Assisted Criminal Sentencing of Drug Trafficking CasesabstractJudgment prediction is the task of predicting various outcomes of legal cases of which sentencing prediction is one of the most important yet difficult challenges. We study the applicability of machine learning (ML) techniques in predicting prison terms of drug trafficking cases. In particular, we study how legal domain knowledge can be integrated with ML models to construct highly accurate predictors. We illustrate how our criminal sentence predictors can be applied to address four important issues in legal knowledge management, which include (1) discovery of model drifts in legal rules, (2) identification of critical features in legal judgments, (3) fairness in machine predictions, and (4) explainability of machine predictions. Tien-Hsuan Wu, Ben Kao, Anne S. Y. Cheung, Michael M. K. Cheung, Yongxi Chen, Guowen Yuan, Reynold Cheng |
JURIX | 1 |
| 2020 | MULCE: Multi-level Canonicalization with Embeddings of Open Knowledge Bases
Tien-Hsuan Wu, Ben Kao, Zhiyong Wu 0003, Xiyang Feng, Qianli Song |
WISE (1) | 1 |
| 2020 | PERQ: Predicting, Explaining, and Rectifying Failed Questions in KB-QA SystemsabstractA knowledge-based question-answering (KB-QA) system is one that answers natural-language questions by accessing information stored in a knowledge base (KB). Existing KB-QA systems generally register an accuracy of 70-80% for simple questions and less for more complex ones. We observe that certain questions are intrinsically difficult to answer correctly with existing systems. We propose the PERQ framework to address this issue. Given a question q, we perform three steps to boost answer accuracy: (1) (Prediction) We predict if q can be answered correctly by a KB-QA system S. (2) (Explanation) If S is predicted to fail q, we analyze them to determine the most likely reasons of the failure. (3) (Rectification) We use the prediction and explanation results to rectify the answer. We put forward tools to achieve the three steps and analyze their effectiveness. Our experiments show that the PERQ framework can significantly improve KB-QA systems' accuracies over simple questions. Zhiyong Wu 0003, Ben Kao, Tien-Hsuan Wu, Qun Liu 0001 |
WSDM | 3 |
| 2018 | Towards Practical Open Knowledge Base CanonicalizationabstractAn Open Information Extraction (OIE) system processes textual data to extract assertions, which are structured data typically represented in the form of (subject;relation; object) triples. An Open Knowledge Base (OKB) is a collection of such assertions. We study the problem of canonicalizing an OKB, which is defined as the problem of mapping each name (a textual term such as "the rockies", "colorado rockies") to a canonical form (such as "rockies"). Galárraga et al. [18] proposed a hierarchical agglomerative clustering algorithm using canopy clustering to tackle the canonicalization problem. The algorithm was shown to be very effective. However, it is not efficient enough to practically handle large OKBs due to the large number of similarity score computations. We propose the FAC algorithm for solving the canonicalization problem. FAC employs pruning techniques to avoid unnecessary similarity computations, and bounding techniques to efficiently approximate and identify small similarities. In our experiments, FAC registers ordersof-magnitude speedups over other approaches. Tien-Hsuan Wu, Zhiyong Wu 0003, Ben Kao |
CIKM | 1 |