Jia-Ling Koh

dblp:68/5059 · DBLP profile ↗
← Back
28ranked-venue papers
18as first author
3since 2021 · last 2025
0000-0002-3223-6021ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 20 · 14 first-author · 3 since 2021Artificial intelligence and machine learning · 13 · 7 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2025 Improving Prompt-Based Learning Framework for Mental Health Aspect Detection from Social Media
Jia-Ling Koh, Hsiao-Ting Huang, Yin-Ju Lien
DEXA (1)1
2024 Category-Aware Sequential Recommendation with Time Intervals of Purchases
Jia-Ling Koh, Cheng-Wei Chen
DEXA (1)1
2021 Multimodal depression detection on instagram considering time interval of posts
Chun-Yueh Chiu, Hsien Yuan Lane, Jia-Ling Koh, Arbee L. P. Chen
J. Intell. Inf. Syst.3
2020 Question Generation Through Transfer Learning
Yin-Hsiang Liao, Jia-Ling Koh
IEA/AIE2
2018 Conditional Relationship Extraction for Diseases and Symptoms by a Web Search-Based Approach
abstract
This paper studies the strategies of automatically extracting the conditional relationships between diseases and symptoms from a Chinese encyclopedia site and the disease-related web pages searched from the Internet. At first, the seed symptoms of a disease are extracted from an online medical encyclopedia automatically. These seed symptoms are utilized as query keywords to automatically find more symptoms of a disease from the unstructured documents of the disease-related search results. Next, a jointly learning method is used to construct the embedded representations of the conditional terms and pattern terms. Besides, the semantic similarity matrix of conditional terms is computed through the co-clustering algorithm to discover the representative conditional terms of the clusters. The result of experiments shows that the proposed method, which discovers the semantically related symptoms of a disease associated with conditionals, achieves high accuracy. Besides, many unusually known symptoms considered by the medical experts are discovered, which may be noticeable symptoms needing further verification in the future.
Yi-Hui Lee, Jia-Ling Koh
WI2
2017 Timeline Summarization for Event-Related Discussions on a Chinese Social Media Platform
Han Wang 0031, Jia-Ling Koh
IEA/AIE (1)2
2017 MapReduce skyline query processing with partitioning and distributed dominance tests
Jia-Ling Koh, Chia-Ching Chen, Chih-Yu Chan, Arbee L. P. Chen
Inf. Sci.1
2016 A Maximum Dimension Partitioning Approach for Efficiently Finding All Similar Pairs
Jia-Ling Koh, Shao-Chun Peng
DaWaK1
2015 Dynamic Facet Hierarchy Constructing for Browsing Web Search Results Efficiently
Jia-Ling Koh
IEA/AIE2
2014 An Efficient Approach for Mining Top-k High Utility Specialized Query Expansions on Social Tagging Systems
Jia-Ling Koh, I-Chih Chiu
DASFAA (2)1
2014 Finding k most favorite products based on reverse top-t queries
Jia-Ling Koh, Chen-Yi Lin, Arbee L. P. Chen
VLDB J.1
2013 Determining $(k)$-Most Demanding Products with Maximum Expected Number of Total Customers
abstract
In this paper, a problem of production plans, named k-most demanding products (k-MDP) discovering, is formulated. Given a set of customers demanding a certain type of products with multiple attributes, a set of existing products of the type, a set of candidate products that can be offered by a company, and a positive integer k, we want to help the company to select k products from the candidate products such that the expected number of the total customers for the k products is maximized. We show the problem is NP-hard when the number of attributes for a product is 3 or more. One greedy algorithm is proposed to find approximate solution for the problem. We also attempt to find the optimal solution of the problem by estimating the upper bound of the expected number of the total customers for a set of k candidate products for reducing the search space of the optimal solution. An exact algorithm is then provided to find the optimal solution of the problem by using this pruning strategy. The experiment results demonstrate that both the efficiency and memory requirement of the exact algorithm are comparable to those for the greedy algorithm, and the greedy algorithm is well scalable with respect to k.
Chen-Yi Lin, Jia-Ling Koh, Arbee L. P. Chen
IEEE Trans. Knowl. Data Eng.2
2012 A multi-level hierarchical index structure for supporting efficient similarity search on tag sets
abstract
Social communication websites has been an emerging type of a Web service that helps users to share their resources. For providing efficient similarity search of tag set in a social tagging system, we propose a multi-level hierarchical index structure to group similar tag sets. Not only the algorithms of similarity searches of tag sets, but also the algorithms of deletion and updating of tag sets by using the constructed index structure are provided. Furthermore, we define a modified hamming distance function on tag sets, which consider the semantically relatedness when comparing the members for evaluating the similarity of two tag sets. This function is more applicable to evaluate the similarity search of two tag sets. A systematic performance study is performed to verify the effectiveness and the efficiency of the proposed strategies. The experiment results show that the proposed MHIB approach further improves the pruning effect of the previous work which constructs a two-level index structure. Especially, the MHIB approach is well scalable with respect to the three parameters when using either the hamming distance or the modified hamming distance for similarity measure. Although the insertion operation of the MHIB approach requires higher cost than the naïve method, with the assistant of the constructed inverted list of clusters, it performs faster than the previous work. Besides, the cost of performing deletion operation by using the MHIB approach is much less than the other two approaches and so is the update operation.
Jia-Ling Koh, Nonhlanhla Shongwe, Chung-Wen Cho
RCIS1
2011 Informative Sentence Retrieval for Domain Specific Terminologies
Jia-Ling Koh, Chin-Wei Cho
IEA/AIE (1)1
2010 Hierarchical Topic-Based Communities Construction for Authors in a Literature Database
Chien-Liang Wu, Jia-Ling Koh
IEA/AIE (2)2
2010 A Better Strategy of Discovering Link-Pattern Based Communities by Classical Clustering Methods
Chen-Yi Lin, Jia-Ling Koh, Arbee L. P. Chen
PAKDD (1)2
2010 A Tree-based Approach for Efficiently Mining Approximate Frequent Itemsets
abstract
The strategies for mining frequent itemsets, which is the essential part of discovering association rules, have been widely studied over the last decade. In real-world datasets, it is possible to discover multiple fragmented patterns but miss the longer true patterns due to random noise and errors in the data. Therefore, a number of methods have been proposed recently to discover approximate frequent itemsets. However, a challenge of providing an efficient algorithm for solving this problem is how to avoid costly candidate generation and test. In this paper, an algorithm, named FP-AFI (FP-tree based Approximate Frequent Itemsets mining algorithm), is developed to discover approximate frequent itemsets from a FP-tree-like structure. We define a recursive function for getting the set of transactions which fault-tolerant contain an itemset P. The patterns in the fault-tolerant supporting transactions of P are represented by the conditional AFP-trees of P. Moreover, to avoid re-constructing the tree structure in the mining process, two pseudo-projection operations on AFP-trees are provided to obtain the conditional AFP-trees of a candidate itemset systematically. Consequently, the approximate support of a candidate itemset and the item supports of each item in the candidate are obtained easily from the conditional AFP-trees. Hence, the constrain test of a candidate itemset is performed efficiently without additional database scan. The experimental results show that the FP-AFI algorithm performs much better than the FP-Apriori and the AFI algorithms in efficiency especially when the size of data set is large and the minimum threshold of approximate support is small. Moreover, the execution time of FP-AFI is scalable even when the error threshold parameters become large.
Jia-Ling Koh, Yi-Lang Tu
RCIS1
2006 An Efficient Approach for Mining Top-K Fault-Tolerant Repeating Patterns
Jia-Ling Koh, Yu-Ting Kung
DASFAA1
2006 An Approximate Approach for Mining Recently Frequent Itemsets from Data Streams
Jia-Ling Koh, Shu-Ning Shin
DaWaK1
2005 An Efficient Approach for Mining Fault-Tolerant Frequent Patterns Based on Bit Vector Representations
Jia-Ling Koh, Pei-Wy Yo
DASFAA1
2005 Improved Sequential Pattern Mining Using an Extended Bitmap Representation
Chien-Liang Wu, Jia-Ling Koh, Pao-Ying An
DEXA2
2004 An Efficient Approach for Maintaining Association Rules Based on Adjusting FP-Tree Structures1
Jia-Ling Koh, Shui-Feng Shieh
DASFAA1
2002 Efficient Query Processing in Integrated Multiple Object Databases with Maybe Result Certification
abstract
Within integrated multiple object databases, missing data occurs due to the missing attribute conflict as well as the existence of null values. A set of algorithms is provided in this paper to process the predicates of global queries with missing data. To provide more informative answers to users, the "maybe" results due to missing data are presented in addition to the "certain" results. The local "maybe" results may become "certain" results via the concept of object isomerism. One algorithm is designed based on the centralized approach in which data are forwarded to the same site for integration and processing. Furthermore, to reduce the response time, localized approaches evaluate the predicates within distinct component databases in parallel. The object signature is also applied in the design to further reduce the data transfer. These algorithms are compared and discussed according to the simulation results of both the total execution and response times. Alternately, the global schema may contain multi-valued attributes with values derived from attribute values in different component databases. Hence, the proposed approaches are also extended to process the global queries involving this kind of multi-valued attribute.
Jia-Ling Koh, Arbee L. P. Chen
IEEE Trans. Knowl. Data Eng.1
2001 Efficient Feature Mining in Music Objects
Jia-Ling Koh, William D. C. Yu
DEXA1
1997 A Query Language and Interface for Integrated Media and Alphanumeric Database Systems
Jia-Ling Koh, Arbee L. P. Chen, Paul C. M. Chang, James C. C. Chen
DEXA1
1996 Query Execution Strategies for Missing Data in Distributed Heterogeneous Object Databases
abstract
The problem of missing data arises in the distributed heterogeneous object databases because of the missing attribute conflict and the existence of null values. A set of algorithms are provided in this paper for query processing when the predicates in global queries involve missing data. For providing more informative answers to the users, the maybe results due to missing data are presented in addition to the certain results. One algorithm is designed based on the centralized approach in which the data are sent to the same site for integration and then processing. On the other hand, for reducing the response time, the localized approaches evaluate the predicates in different component databases in parallel. The proposed algorithms are compared and discussed by the simulation result on the total execution time and response time.
Jia-Ling Koh, Arbee L. P. Chen
ICDCS1
1996 Identifying Object Isomerism in Multidatabase Systems
Arbee L. P. Chen, Pauray S. M. Tsai, Jia-Ling Koh
Distributed Parallel Databases3
1993 Integration of Heterogeneous Object Schemas
Jia-Ling Koh, Arbee L. P. Chen
ER1