Tingting Wang 0009

dblp:75/4702-9 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0002-4912-7171ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Distinctiveness Maximization in Datasets Assemblage
abstract
In this paper, given a user's query set and budget, we aim to use the limited budget to help users assemble a set of datasets that can enrich a base dataset by introducing the maximum number of distinct tuples (i.e., maximizing distinctiveness). We prove this problem to be NP-hard. A greedy algorithm using exact distinctiveness computation attains an approximation ratio of (1-e-1 )/2, but it lacks efficiency and scalability due to its frequent computation of the exact distinctiveness marginal gain of any candidate dataset for selection. This requires scanning through every tuple in candidate datasets and thus is unaffordable in practice. To overcome this limitation, we propose an efficient machine learning (ML)-based method for estimating the distinctiveness marginal gain of any candidate dataset. This effectively eliminates the need to test each tuple individually. Estimating the distinctiveness marginal gain of a dataset involves estimating the number of distinct tuples in the tuple sets returned by each query in a query set across multiple datasets. This can be viewed as the cardinality estimation for a query set on a set of datasets, and the proposed method is the first to tackle this cardinality estimation problem. This is a significant advancement over prior methods that were limited to single-query cardinality estimation on a single dataset and struggled with identifying overlaps among tuple sets returned by each query in a query set across multiple datasets. Extensive experiments using five real-world data pools demonstrate that our algorithm, which utilizes ML-based distinctiveness estimation, outperforms all relevant baselines in effectiveness, efficiency, and scalability. A case study on two downstream ML tasks also highlights its potential to find datasets with more useful tuples to enhance the performance of ML tasks.
Tingting Wang 0009, Shixun Huang, Zhifeng Bao, J. Shane Culpepper, Volkan Dedeoglu, Reza Arablouei
WWW1
2024 Optimizing Data Acquisition to Enhance Machine Learning Performance
abstract
In this paper, we study how to acquire labeled data points from a large data pool to enrich a training set for enhancing supervised machine learning (ML) performance. The state-of-the-art solution is the clustering-based training set selection (CTS) algorithm, which initially clusters the data points in a data pool and subsequently selects new data points from clusters. The efficiency of CTS is constrained by its frequent retraining of the target ML model, and the effectiveness is limited by the selection criteria, which represent the state of data points within each cluster and impose a restriction of selecting only one cluster in each iteration. To overcome these limitations, we propose a new algorithm, called CTS with incremental estimation of adaptive score (IAS). IAS employs online learning, enabling incremental model updates by using new data, and eliminating the need to fully retrain the target model, and hence improves the efficiency. To enhance the effectiveness of IAS, we introduce adaptive score estimation, which serves as novel selection criteria to identify clusters and select new data points by balancing trade-offs between exploitation and exploration during data acquisition. To further enhance the effectiveness of IAS, we introduce a new adaptive mini-batch selection method that, in each iteration, selects data points from multiple clusters rather than a single cluster, hence eliminating the potential bias due to using only one cluster. By integrating this method into the IAS algorithm, we propose a novel algorithm termed IAS with adaptive mini-batch selection (IAS-AMS). Experimental results highlight the superior effectiveness of IAS-AMS, with IAS also outperforming other competing algorithms. In terms of efficiency, IAS takes the lead, while the efficiency of IAS-AMS is on par with that of the existing CTS algorithm.
Tingting Wang 0009, Shixun Huang, Zhifeng Bao, J. Shane Culpepper, Volkan Dedeoglu, Reza Arablouei
Proc. VLDB Endow.1
2023 An Integrative Disease Information Network Approach to Similar Disease Detection
abstract
Disease similarity analysis impacts significantly in pathogenesis revealing, treatment recommending, and disease-causing genes predicting. Previous works study the disease similarity based on the semantics obtaining from biomedical ontologies (e.g., disease ontology) or the function of disease-causing molecules. However, such methods almost focus on a single perspective for obtaining disease features, which may lead to biased results for similar disease detection. To address this issue, we propose a disease information network-based integrative approach named MISSION for detecting similar diseases. By leveraging the associations between diseases and other biomedical entities, the disease information network is established first. Then, the disease similarity features extracted from the aspects of disease taxonomy, attributes, literature, and annotations are integrated into the disease information network. Finally, the top-k similar disease query is performed based on the integrative disease information. The experiments conducted on real-world datasets demonstrate that MISSION is effective and useful in similar disease detection.
Wuli Xu, Lei Duan, Huiru Zheng, Jesse Li-Ling, Yidan Zhang 0001, Tingting Wang 0009, Ruiqi Qin 0001
IEEE ACM Trans. Comput. Biol. Bioinform.7
2023 Dynamic Ridesharing With Minimal Regret: Towards an Enhanced Engagement Among Three Stakeholders
abstract
In dynamic ridesharing, the platform serves as the mediator by tailoring the assignment result between workers and riders with a focus on a certain objective. Existing studies generally focus on either one or two stakeholders when modelling the problem while the wellbeing of the other parties may be ignored or even undermined. For example, purely maximizing the total revenue of the ridesharing platform may cause the loss of riders and in turn lead to a low served rate, because those expensive orders will be processed in priority. In this paper, we for the first time study how to incorporate the willingness of all stakeholders (i.e., the platform, workers and riders). Given a set of workers and a set of rider requests, we aim to return the matchable worker-rider pairs in order to minimize theregret. Specifically, two types of regret are defined: (i) theserved rate regret, which refers to the rate of unserved requests, catering for the reputation and profit of the platform and workers; (ii) therevenue regret, which considers the portion of revenue loss from unserved riders, catering for the focus of workers and riders in the trip schedule. We prove the NP-hardness of this problem. To tackle this problem, we first propose a dynamic programming insertion algorithm to improve the efficiency of inserting a rider request into a trip schedule of a worker. Furthermore, two kinds of heuristic algorithms are devised to match rider requests with workers effectively. Comprehensive experiments on two real-world datasets verify the effectiveness, efficiency and scalability of our solutions in dealing with different supply-demand relationships in practice.
Tingting Wang 0009, Hui Luo 0001, Zhifeng Bao, Lei Duan
IEEE Trans. Knowl. Data Eng.1
2022 Representative Routes Discovery from Massive Trajectories
abstract
In this work, we study how to find the k most representative routes over large scale trajectory data, which is a fundamental operation that benefits various real-world applications, such as traffic monitoring and public transportation planning. The operator is time-sensitive as it must be able to adapt the results as traffic conditions change. We first prove the NP-hardness of the problem, and then propose a range of effective approximate solutions that have rapid response times. Specifically, we first build a lookup table that stores the trajectories covered by each edge in a given road network. Rather than performing a depth-first search for all possible routes, we find a 1/η approximate solution by developing a maximum-weight algorithm. Since each edge in a route may be close to several trajectories, we further propose a coverage-first algorithm to locate the edges with the greatest coverage gain in the solution route set. By observing that in the real world each edge is connected to only a few other edges in a road network, we have developed a connect-first algorithm that finds consecutive edges for k representative routes by greedily selecting edges with the maximum marginal gain for each route. Finally, comprehensive experiments over two real-world datasets are conducted to verify the effectiveness and efficiency of our proposed algorithms, and provide evidence of the usefulness of our solution and rapid response times in traffic monitoring tasks.
Tingting Wang 0009, Shixun Huang, Zhifeng Bao, J. Shane Culpepper, Reza Arablouei
KDD1
2022 Efficient mining of concept-hierarchy aware distinguishing sequential patterns
Chengxin He, Lei Duan, Guozhu Dong, Jyrki Nummenmaa, Tingting Wang 0009, Tinghai Pang
Knowl. Based Syst.5
2022 Mining Similar Aspects for Gene Similarity Explanation Based on Gene Information Network
abstract
Analysis of gene similarity not only can provide information on the understanding of the biological roles and functions of a gene, but may also reveal the relationships among various genes. In this paper, we introduce a novel idea of mining similar aspects from a gene information network, i.e., for a given gene pair, we want to know in which aspects (meta paths) they are most similar from the perspective of the gene information network. We defined a similarity metric based on the set of meta paths connecting the query genes in the gene information network and used the rank of similarity of a gene pair in a meta path set to measure the similarity significance in that aspect. A minimal set of gene meta paths where the query gene pair ranks the highest is a similar aspect, and the similar aspect of a query gene pair is far from trivial. We proposed a novel method, SCENARIO, to investigate minimal similar aspects. Our empirical study on the gene information network, constructed from six public gene-related databases, verified that our proposed method is effective, efficient, and useful.
Yidan Zhang 0001, Lei Duan, Huiru Zheng, Jesse Li-Ling, Ruiqi Qin 0001, Chengxin He, Tingting Wang 0009
IEEE ACM Trans. Comput. Biol. Bioinform.8
2020 Efficient Mining of Outlying Sequence Patterns for Analyzing Outlierness of Sequence Data
abstract
Recently, a lot of research work has been proposed in different domains to detect outliers and analyze the outlierness of outliers for relational data. However, while sequence data is ubiquitous in real life, analyzing the outlierness for sequence data has not received enough attention. In this article, we study the problem of mining outlying sequence patterns in sequence data addressing the question: given a query sequence s in a sequence dataset D , the objective is to discover sequence patterns that will indicate the most unusualness (i.e., outlierness) of s compared against other sequences. Technically, we use the rank defined by the average probabilistic strength ( aps ) of a sequence pattern in a sequence to measure the outlierness of the sequence. Then a minimal sequence pattern where the query sequence is ranked the highest is defined as an outlying sequence pattern. To address the above problem, we present OSPMiner, a heuristic method that computes aps by incorporating several pruning techniques. Our empirical study using both real and synthetic data demonstrates that OSPMiner is effective and efficient.
Tingting Wang 0009, Lei Duan, Guozhu Dong, Zhifeng Bao
ACM Trans. Knowl. Discov. Data1
2019 ATOM: Construction of Anti-tumor Biomaterial Knowledge Graph by Biomedicine Literature
abstract
With the rapid development of anti-tumor biomaterials, biomedicine literature with respect to anti-tumor biomaterials has been leveraged for tumor treatment as it provides abundant and useful information. A large number of biomedicine literature contains unstructured data, making it difficult for researchers to obtain desired messages from it. Knowledge Graphs (KGs) provides structured relationships among entities and can be served as a solution. However, no existing tool can be found in constructing an anti-tumor biomaterial knowledge graph from biomedicine literature. To fill this gap, a novel approach, ATOM, was proposed to construct an anti-tumor biomaterial knowledge graph from biomedicine literature through a series of process including the recognition of anti-tumor entities, the simplification of sentences, the extraction of triples, and the predicate mapping. Experiments demonstrated that ATOM is able to effectively express the extracted anti-tumor entities and their relationships.
Tingting Wang 0009, Lei Duan, Chengxin He, Geng Deng, Ruiqi Qin 0001, Yidan Zhang 0001
BIBM1