Ning Wang 0024

dblp:46/2005-24 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-8903-8790ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 13 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author
YearPublicationVenuePosition
2026 CrossER: A robust and adaptable generalized entity resolution framework for diverse and heterogeneous datasets
Yunong Tian, Ning Wang 0024, Anshun Zhou
Inf. Syst.2
2025 A Robust Attention-based Cardinality Estimator with Domain Adaptation
Qixiong Zeng, Pengfei Xi, Ning Wang 0024
Expert Syst. Appl.5
2025 QueryAccelerator: A SQL converter capturing and transforming the slow plan
Shaocong Yang, Ning Wang 0024, Qixiong Zeng, Anshun Zhou
Neurocomputing2
2024 CrowdDA: Difficulty-aware crowdsourcing task optimization for cleaning web tables
Yihai Xi, Ning Wang 0024
Expert Syst. Appl.2
2023 Popularity sensitive and domain-aware summarization for web tables
Yihai Xi, Ning Wang 0024, Shuang Hao 0002
Inf. Sci.2
2023 An optimized task assignment framework based on crowdsourcing knowledge graph and prediction
Junyuan Quan, Ning Wang 0024
Knowl. Based Syst.2
2023 HOFD: An Outdated Fact Detector for Knowledge Bases
abstract
Knowledge bases (KBs), which store high-quality information, are crucial for many applications, such as enhancing search results and serving as external sources for data cleaning. Not surprisingly, there exist outdated facts in most KBs due to the rapid change of information. Naturally, it is important to keep KBs up-to-date. Traditional wisdom has investigated the problem of using reference data (such as new facts extracted from the news) to detect outdated facts in KBs. However, existing approaches can only cover a small percentage of facts in KBs. In this paper, we proposeHOFD, a novel human-in-the-loop approach for outdated fact detection in KBs.HOFDtrains a binary classifier using features such as historical update frequency and update time of a fact to compute the likelihood of a fact in a KB to be outdated. Then,HOFDinteracts with humans to verify whether a fact with high likelihood is indeed outdated. In addition,HOFDalso uses logical rules to detect more outdated facts based on human feedback. The outdated facts detected by the logical rules will also be fed back to train the ML model further fordata augmentation. Extensive experiments on real-world KBs, such as Yago and DBpedia, show the effectiveness of our solution.
Shuang Hao 0002, Chengliang Chai, Guoliang Li 0001, Nan Tang 0001, Ning Wang 0024
IEEE Trans. Knowl. Data Eng.5
2022 SmartIndex: An Index Advisor with Learned Cost Estimator
abstract
As an important part of database optimization, index selection problem remains a hot topic. Existing methods tend to use the cost estimated by the DBMS optimizer to measure the benefit of an index. However, due to the limitations of the cost estimation model in database management system (DBMS), these methods may not find the optimal index configuration. To address this problem, we present SmartIndex, an index advisor for relational database with learned cost estimator. We first design a graph convolutional network (GCN) based cost estimation model to predict a query's execution time on certain indexes. After that, we use a greedy method for index selection under certain constraints including number of indexes and storage cost of indexes, which can find better solutions for a given workload.
Jianling Gao, Ning Wang 0024, Shuang Hao 0002
CIKM3
2022 Automatic index selection with learned cost estimator
Jianling Gao, Ning Wang 0024, Shuang Hao 0002, Haoyan Wu
Inf. Sci.3
2022 EasyDR: A Human-in-the-loop Error Detection&Repair Platform for Holistic Table Cleaning
abstract
Many tables on the web suffer from multi-level and multi-type quality problems, but existing cleaning systems cannot provide a comprehensive quality improvement for them. Most of these systems are designed for solving a specific type of error, so that we need to resort to a number of different cleaning tools (one per error type) to get a high quality table. In this demonstration, we propose a human-in-the-loop cleaning platform EasyDR for detecting and repairing multi-level&multi-type errors in tables. The attendees will experience the following features of EasyDR: 1) Holistic error detection&repair. Users are able to perform a holistic table cleaning in EasyDR where machine algorithms are responsible for error detection while human intelligence is leveraged for error repairing. 2) Human-in-the-loop table cleaning. EasyDR performs an all-round quality diagnosis for the table, and automatically generates crowdsourcing cleaning tasks for the detected errors. To simplify cleaning tasks for crowdsourcing workers, EasyDR provides two task optimization techniques including domain-aware table summarization and difficulty-aware task order optimization. 3) Customizable cleaning mode. EasyDR provides a declarative language for users to customize cleaning tasks flexibly, e.g., selecting target errors, restricting the cleaning scope, defining the cooperation mode for machine and crowd.
Yihai Xi, Ning Wang 0024
Proc. VLDB Endow.2
2021 Mis-categorized entities detection
Shuang Hao 0002, Nan Tang 0001, Guoliang Li 0001, Jianhua Feng, Ning Wang 0024
VLDB J.5
2020 Outdated Fact Detection in Knowledge Bases
abstract
Knowledge bases (KBs), which store high-quality information, are crucial for many applications, such as enhancing search results and serving as external sources for data cleaning. Not surprisingly, there exist outdated facts in most KBs due to the rapid change of information. Naturally, it is important to keep KBs up-to-date. Traditional wisdom has investigated the problem of using reference data (such as new facts extracted from the news) to detect outdated facts in KBs. However, existing approaches can only cover a small percentage of facts in KBs. In this paper, we propose a novel human-in-the-loop approach for outdated fact detection in KBs. It trains a binary classifier using features such as historical update frequency and existence time of a fact to compute the likelihood of a fact in a KB to be outdated. Then, it interacts with humans to verify whether a fact with high likelihood is indeed outdated. In addition, it also uses logical rules to detect more outdated facts based on human feedback. The outdated facts detected by the logical rules will also be fed back to train the ML model further for data augmentation. Extensive experiments on real-world KBs, such as Yago and DBpedia, show the effectiveness of our solution.
Shuang Hao 0002, Chengliang Chai, Guoliang Li 0001, Nan Tang 0001, Ning Wang 0024
ICDE5
2020 PocketView: A Concise and Informative Data Summarizer
abstract
A data summarization for the large table can be of great help, which provides a concise and informative overview and assists the user to quickly figure out the subject of the data. However, a high quality summarization needs to have two desirable properties: presenting notable entities and achieving broad domain coverage. In this demonstration, we propose a summarizer system called PocketView that is able to create a data summarization through a pocket view of the table. The attendees will experience the following features of our system:(1) time-sensitive notability evaluation - PocketView can automatically identify notable entities according to their significance and popularity in user-defined time period; (2) broad-coverage pocket view - Our system will provide a pocket view for the table without losing any domain, which is much simpler and clearer for attendees to figure out the subject compared with the original table.
Yihai Xi, Ning Wang 0024, Shuang Hao 0002, Wenyang Yang
ICDE2
2018 Identifying Multiple Entity Columns in Web Tables
abstract
Unlike tables in relational database, web tables have no designated key attributes or entity columns, so it is difficult for computers to understand a table and associate with it a concept in the knowledge taxonomy. Existing techniques for entity column detection can only process tables with single entity column, discarding tables which describe multiple concepts. In this paper, we propose a framework for identifying multiple entity columns in a web table. At first, we annotate column labels for a web table with missing or noninformative labels based on external knowledge base Probase. By detecting concept-attribute relationships between table columns and calculating the credibility of attribute dependency, we construct a column dependency view for the table. Then, the column semantic intensity is calculated for each column in a web table, which depends on its connectivity in column dependency view and the dependency credibility of attribute dependency relationships related to it. We can identify all entity columns from the web table by iteratively selecting primary entity column with the highest column semantic intensity and accordingly separate columns describing the primary concept from present column dependency view. The results of a comprehensive set of experiments indicate that our entity detection method is more effective than existing methods for either single or multiple concept tables.
Ning Wang 0024, Xiangran Ren
Int. J. Softw. Eng. Knowl. Eng.1
2017 Building Top-k Consistent Results for Web Table Augmentation
abstract
Web table augmentation enables users to augment attributes based on key column and other known information. For table augmentation, most of systems return a single result which could not meet the users' needs of selection and validation. Furthermore, previous works only consider the entity-attribute binary tables with the first column corresponding to the entity name and the second to an attribute to be extended. When a table has multiple columns to be extended, the result table consolidated by binary tables will suffer from entity inconsistency. In this paper, we present a framework called TAT to build Top-k consistent results for web table augmentation. While ensuring the consistency of entities, TAT provides as diverse results as possible. We design two algorithms, exclusive and iterative algorithm, for web table augmentation that return Top-k results based on different requirements from users. The experiments show that TAT could return Top-k consistent results without loss of precision or coverage.
Ning Wang 0024
WISA3
2016 Summarizing Personal Dataspace Based on User Interests
abstract
A personal dataspace management system (PDSMS) is a platform to manage personal data with various data types. Facing huge volume of heterogeneous personal data and complex relationships between them, it is better for users to start with a simplified, easy-to-read schema and then explore in depth only the relevant schema elements during formulating queries. Existing approaches of database schema summarization neglect user interests, which is very important in a personal dataspace. We propose a framework for building a concise resource summary based on user interests automatically in PDSMS. Our method builds the initial summary by partitioning schema graph according to its linkage information and selecting representative elements based on a novel measure on schema element typicality. Then, user interested degree for schema node is introduced to measure user interests and the initial summary is refined according to user interests. Finally, we evaluate the quality of our resource summary through a comprehensive set of experiments, and results indicate that summaries generated by our system are more effective on reducing user efforts required in formulating queries.
Ning Wang 0024
Int. J. Softw. Eng. Knowl. Eng.1
2015 CrowdSR: A Crowd Enabled System for Semantic Recovering of Web Tables
Huaxi Liu, Ning Wang 0024, Xiangran Ren
WAIM2
2014 An Iterative Approach to Managing Uncertain Mappings in Dataspace Support Platforms
abstract
A DataSpace Support Platform (DSSP) is a self-sustained and self-managed system which needs to support uncertainty among its mediated schemas and its schema mappings. Some approaches for managing such uncertainty by assigning probabilities and reliability degrees to schema mappings have been proposed. Unfortunately, the number of mappings self-generated by a DSSP is usually too large and among those possible mappings, some might be totally correct and others partially correct. Therefore, providing probabilities or reliability degrees to the mappings is necessary but not sufficient to resolve uncertainty among them. This paper proposes a stepper-based approach called pos-mapping to managing reliable mappings using possibility theory. Instead of choosing a threshold for managing the reliable mappings, pos-mapping approach orders and divides the set of reliable mappings into subsets of possibility distributions and assigns to each of these subsets a recursive possibility degree function. The recursiveness of the possibility degree function leads to an incremental management of the possibility distributions. Experimental results show that our system is more efficient than the existing systems and the accuracy of the results increases with the number of reliable schemas in the DSSP.
Nathalie Cindy Kuicheu, Ning Wang 0024, Gile Narcisse Fanzou Tchuissang, De Xu, Guojun Dai, François Siewe
Int. J. Softw. Eng. Knowl. Eng.2
2013 Restoring: A Greedy Heuristic Approach Based on Neighborhood for Correlation Clustering
Ning Wang 0024
ADMA (1)1
2013 XML normalization based on entity segments
Xudong Lin 0002, Ning Wang 0024, Xiaoning Zeng, Yanyan Sun
Inf. Sci.2
2010 A novel XML keyword query approach using entity subtree
Xudong Lin 0002, Ning Wang 0024, De Xu, Xiaoning Zeng
J. Syst. Softw.2