Liwei Wang 0011

dblp:47/1798-11 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-5275-8786ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 16 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 LORE: Learning-Based Resource Recommendation for Big Data Queries
abstract
With the development of modern cloud platforms, an increasing number of users are migrating their data analysis tasks to the cloud. Cloud platforms offer a “pay-as-you-go” model, prompting users to focus on both performance and resource costs. Existing query optimization methods primarily address query performance while neglecting resource costs. Mapping queries to their resource consumption is a complex task. To tackle this challenge, we propose a novel learning-based query resource recommendation method called LORE. LORE efficiently and accurately estimates the optimal resources for queries by leveraging dual information from SQL query statements and query execution plans. We model SQL queries and execution plans as directed acyclic graphs and utilize graph neural networks to derive comprehensive representations. To capture the dependencies among all nodes involved in data transmission within an execution plan, we assign path weights to the dependency edges of each node. Our approach integrates data distribution information and captures both direct and indirect dependencies among plan nodes while avoiding unnecessary redundant computations. Experimental results demonstrate that, compared to traditional and other learning-based methods, the LORE model achieves higher accuracy in predicting the optimal resources for queries.
Yan Li 0161, Liwei Wang 0011, Bolong Zheng, Zhiyong Peng 0001
ICDE2
2025 Towards proactive rumor control: When a budget constraint meets impression counts
Zhiyong Peng 0001, Liwei Wang 0011
Comput. Commun.3
2025 Proactive Rumor Control Using Graph Neural Networks and Evolutionary Optimization
abstract
Abstract In the digital age, social networks have become critical platforms for information dissemination, but they also pose significant risks due to the rapid spread of rumors and misinformation. Existing approaches to rumor control often rely on models that assume a single exposure to anti-rumor information is sufficient to mitigate its impact, overlooking the necessity of multiple impressions for effective behavior change. In this work, we address the Rumor Control with Impression Counts problem by proposing the first-ever Machine Learning (ML)-based solution. Our approach leverages Graph Neural Networks combined with a greedy algorithm to efficiently manage large-scale social networks. To further enhance computational efficiency, we incorporate Evolutionary Optimization, resulting in a method that not only addresses the effectiveness challenges but also scales efficiently with network size. Extensive experiments on real-world datasets demonstrate that our approach outperforms existing methods in both effectiveness and scalability, improving computational efficiency by 1 to 2 orders of magnitude.
Liwei Wang 0011, Zhiyong Peng 0001
Data Sci. Eng.2
2024 A learned cost model for big data query processing
Yan Li 0161, Liwei Wang 0011, Sheng Wang 0007, Yuan Sun 0003, Bolong Zheng, Zhiyong Peng 0001
Inf. Sci.2
2023 Proactive Rumor Control: When Impression Counts
Zhiyong Peng 0001, Liwei Wang 0011
PAKDD (4)3
2022 A Resource-Aware Deep Cost Model for Big Data Query Processing
abstract
The efficiency of query processing is highly affected by execution plans and allocated resources in the Spark SQL big data processing engine. However, the cost models for Spark SQL are still based on hand-crafted rules. The learning-based cost models have been proposed for relational databases, but it does not consider the effect of the available resources. To address this, we propose a resource-aware deep learning model that can automatically predict the execution time of query plans based on historical data. To train our model, we embed the query execution plans based on the query plan tree and extract features from the allocated resources. A deep learning model with adaptive attention mechanisms is then trained to predict the execution time of query plans. The experiments show that our deep cost model can achieve higher accuracy in predicting the execution time of query plans compared to traditional rule-based methods and relational database learning-based optimizers.
Yan Li 0161, Liwei Wang 0011, Sheng Wang 0007, Yuan Sun 0003, Zhiyong Peng 0001
ICDE2
2020 Deep Reinforcement Learning-Based Approach to Tackle Topic-Aware Influence Maximization
abstract
Abstract Motivated by the application of viral marketing , the topic-aware influence maximization (TIM) problem has been proposed to identify the most influential users under given topics. In particular, it aims to find k seeds (users) in social network G , such that the seeds can maximize the influence on users under the specific query topics and diffusion model such as independent cascade (IC) or linear threshold (LT). This problem has been proved to be NP-hard, and most of the proposed techniques suffer from the efficiency issue due to the lack of generalization. Even worse, the design of these algorithms requires significant specialized knowledge which is hard to be understood and implemented. To overcome these issues, this paper aims to learn a generalized heuristic framework to solve TIM problems by meta-learning. To this end, we first propose two topic-aware social influence propagation models based on IC and LT model, respectively, which is conducive to better advertising injections. We then encode the feature of each node by a vector and introduce a model, called deep influence evaluation model , to evaluate the user influence under different circumstances. Based on this model, we can construct the solution according to the influence evaluations efficiently, rather than spending a high cost to compute the exact influence by considering the complex graph structure. We conducted experiments on generated graph instances and real-world social networks. The results show the superiority in performance and comparable quality of our framework.
Shan Tian, Songsong Mo, Liwei Wang 0011, Zhiyong Peng 0001
Data Sci. Eng.3
2019 SIRCS: Slope-intercept-residual Compression by Correlation Sequencing for Multi-stream High Variation Data
Zixin Ye, Wen Hua, Liwei Wang 0011, Xiaofang Zhou 0001
DASFAA (1)3
2017 Probabilistic object deputy model for uncertain data and lineage management
Liang Wang 0051, Liwei Wang 0011, Zhiyong Peng 0001
Data Knowl. Eng.2
2015 A Working Model for Uncertain Data with Lineage
Liang Wang 0051, Liwei Wang 0011, Zhiyong Peng 0001
ER2
2013 AML: Efficient Approximate Membership Localization within a Web-Based Join Framework
abstract
In this paper, we propose a new type of Dictionary-based Entity Recognition Problem, named Approximate Membership Localization (AML). The popular Approximate Membership Extraction (AME) provides a full coverage to the true matched substrings from a given document, but many redundancies cause a low efficiency of the AME process and deteriorate the performance of real-world applications using the extracted substrings. The AML problem targets at locating nonoverlapped substrings which is a better approximation to the true matched substrings without generating overlapped redundancies. In order to perform AML efficiently, we propose the optimized algorithm P-Prune that prunes a large part of overlapped redundant matched substrings before generating them. Our study using several real-word data sets demonstrates the efficiency of P-Prune over a baseline method. We also study the AML in application to a proposed web-based join framework scenario which is a search-based approach joining two tables using dictionary-based entity recognition from web documents. The results not only prove the advantage of AML over AME, but also demonstrate the effectiveness of our search-based approach.
Zhixu Li, Laurianne Sitbon, Liwei Wang 0011, Xiaofang Zhou 0001, Xiaoyong Du 0001
IEEE Trans. Knowl. Data Eng.3
2012 Efficient provenance storage for relational queries
abstract
Provenance information is vital in many application areas as it helps explain data lineage and derivation. However, storing fine-grained provenance information can be expensive. In this paper, we present a framework for storing provenance information relating to data derived via database queries. In particular, we first propose a provenance tree data structure which matches the query structure and thereby presents a possibility to avoid redundant storage of information regarding the derivation process. Then we investigate two approaches for reducing storage costs. The first approach utilizes two ingenious rules to achieve reduction on provenance trees. The second one is a dynamic programming solution, which provides a way of optimizing the selection of query tree nodes where provenance information should be stored. The optimization algorithm runs in polynomial time in the query size and is linear in the size of the provenance information, thus enabling provenance tracking and optimization without incurring large overheads. Experiments show that our approaches guarantee significantly lower storage costs than existing approaches.
Zhifeng Bao, Henning Köhler, Liwei Wang 0011, Xiaofang Zhou 0001, Shazia Sadiq
CIKM3
2011 Sorting of Search Results Based on Data Quality
abstract
With the rapid development of Web technology, more and more data on the web are considered as the source of information in current society. However, as the qualities of data fetched from different sources are different, it takes a lot of time to search the valuable data from tremendous web information. This paper proposed an approach to evaluate the data quality through integrating all dimensions of data representations. The evaluation of weight is estimated by the scores for data quality provided by users, which results in that the search results sorted by data quality can achieve the requirements of users. This method has been applied in a microbe system.
Liang Wang 0051, Liwei Wang 0011, Zufa Fu, Zeqian Huang, Zhiyong Peng 0001
WISA2
2011 Efficient Name Disambiguation in Digital Libraries
Jia Zhu 0003, Gabriel Pui Cheong Fung, Liwei Wang 0011
WAIM3
2010 Approximate membership localization (AML) for web-based join
abstract
In this paper, we propose a search-based approach to join two tables in the absence of clean join attributes. Non-structured documents from the web are used to express the correlations between a given query and a reference list. To implement this approach, a major challenge we meet is how to efficiently determine the number of times and the locations of each clean reference from the reference list that is approximately mentioned in the retrieved documents. We formalize the Approximate Membership Localization (AML) problem and propose an efficient partial pruning algorithm to solve it. A study using real-word data sets demonstrates the effectiveness of our search-based approach, and the efficiency of our AML algorithm.
Zhixu Li, Laurianne Sitbon, Liwei Wang 0011, Xiaofang Zhou 0001, Xiaoyong Du 0001
CIKM3
2010 Active Duplicate Detection
Liwei Wang 0011, Xiaofang Zhou 0001, Shazia Sadiq, Gabriel Pui Cheong Fung
DASFAA (1)2
2006 A Scientific Workflow Framework Integrated with Object Deputy Model for Data Provenance
Liwei Wang 0011, Zhiyong Peng 0001, Wenhao Ji, Zeqian Huang
WAIM1