EDBT 2026 Demo / reviewers in the wild / expert
Ruoyu Wang 0004
dblp:127/9829-4
· DBLP profile ↗
11ranked-venue papers
10as first author
7since 2021 · last 2026
0000-0002-8278-3269ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient data structures for fast and low-cost first-order logic rule miningabstractLogic rule mining discovers association patterns in the form of logic rules from structured data. Logic rules are widely applied in information systems to assist decisions in an interpretable way. However, too many computational resources are required in state-of-the-art systems, as most of these systems optimize rule mining algorithms from the perspectives of algorithms and architecture, while data efficiency has been overlooked. Although some start-of-the-art systems implement customized data structures to improve mining speed, the space overhead of the data structures is unaffordable when processing large-scale knowledge bases. Therefore, in this article, we propose data structures to improve data efficiency and accelerate logic rule mining. Our techniques implicitly represent the Cartesian product of variable substitutions in logic rules and build compact indices for a logic entailment cache. Furthermore, we create a pool and a lookup table for the cache so that cache components will not be repeatedly created. The evaluation results show that over 95% of memory can be reduced by our techniques, and mining procedures have been accelerated by about 20x on average. Most importantly, mining on large-scale knowledge bases is practical on normal hardware where only one thread and 20GB of memory are sufficient even for large-scale knowledge bases. Ruoyu Wang 0004, Raymond Wong, Daniel Sun 0004 |
Inf. Syst. | 1 |
| 2025 | SIB: Sorted-Integers-Based Index for Compact and Fast Caching in Top-Down Logic Rule Mining Targeting KB CompressionabstractBackground Mining logic rules from structured knowledge bases is the basis of knowledge engineering. Due to the NP‐hardness of the rule mining problem, logic rules cannot be efficiently induced from knowledge bases, especially large‐scale ones. Idea In this article, we propose a compact and efficient index structure for the maintenance of the intermediate data during top‐down rule mining, such that the memory consumption can be reduced and mining efficiency can be improved. Developing Points The index is based on a mapping from constant symbols to integers and the sorting of the mapped integers. Index update has been dissembled into four basic operations. Moreover, the index itself acts as the cache during top‐down mining. Value Most contributions in existing works employ algorithmic and architectural optimizations to improve efficiency. Data‐oriented optimizations have also been explored to some extent, but the data efficiency is relatively low, and the memory consumption is thus becoming a new challenge for state‐of‐the‐art systems. We tackle this challenge in this article, and our technique has been proven more efficient than state‐of‐the‐art systems. We evaluate our method on six datasets which contain up to 160 K records and are frequently used as benchmarks in tasks related to knowledge engineering. The experimental results show that the proposed technique speeds up the rule mining procedure by on average and reduces memory consumption by up to 70%. The space overhead of the data structure is about twice that of the indexed records, which is more than 80% lower than that of the state‐of‐the‐art technique. Ruoyu Wang 0004, Raymond K. Wong 0001, Daniel Sun 0004, Rajiv Ranjan 0001 |
Softw. Pract. Exp. | 1 |
| 2024 | Estimation-based optimizations for the semantic compression of RDF knowledge basesabstractStructured knowledge bases are critical for the interpretability of AI techniques. RDF KBs, which are the dominant representation of structured knowledge, are expanding extremely fast to increase their knowledge coverage, enhancing the capability of knowledge reasoning while bringing heavy burdens to downstream applications. Recent studies employ semantic compression to detect and remove knowledge redundancies via semantic models and use the induced model for further applications, such as knowledge completion and error detection. However, semantic models that are sufficiently expressive for semantic compression cannot be efficiently induced, especially for large-scale KBs, due to the hardness of logic induction. In this article, we present estimation-based optimizations for the semantic compression of RDF KBs from the perspectives of input and intermediate data involved in the induction of first-order logic rules. The negative sampling technique selects a representative subset of all negative tuples with respect to the closed-world assumption, reducing the cost of evaluating the quality of a logic rule used for knowledge inference. The number of logic inference operations used during a compression procedure is reduced by a statistical estimation technique that prunes logic rules of low quality. The evaluation results show that the two techniques are feasible for the purpose of semantic compression and accelerate the compression algorithm by up to 47x compared to the state-of-the-art system. • Negative sampling reduces the cost of a single logic inference. • Negative sampling speeds up semantic compression of RDF KBs by more than 2x. • Entailment estimation reduces the number of logic inferences in logic rule mining. • Entailment estimation speeds up semantic compression of RDF KBs by up to 47x. • Entailment estimation reduces up to 99% of memory cost of semantic compression. Ruoyu Wang 0004, Raymond K. Wong 0001, Daniel Sun 0004 |
Inf. Process. Manag. | 1 |
| 2023 | Horn rule discovery with batched caching and rule identifier for proficient compressor of knowledge dataabstractAbstract Knowledge data has been widely applied to artificial intelligence applications for interpretable and complex reasoning. Modern knowledge bases are constructed via automatic knowledge extraction from open‐accessible sources. Thus the sizes of KBs are continuously growing, heavily burdening the maintenance and application of the knowledge data. Besides the grammatical redundancies, semantically repeated information also frequently appears in knowledge bases but is still under‐explored. Existing semantic compressors fail to efficiently discover expressive patterns and thus perform unsatisfyingly on knowledge data. This article proposes SInC, a semantic inductive compressor, to efficiently induce first‐order Horn rules and semantically compress knowledge bases. SInC improves the scalability of top‐down rule mining by batching correlated records in the cache and further optimizes the pruning of duplication and specialization via an identifier structure of Horn rules. SInC was evaluated on real‐world and synthetic datasets and compared against the state‐of‐the‐art. The results show that the batched caching speed up the rule mining procedure by more than two orders while consuming fewer than three times memory space. The identifier technique speeds up the duplication and specialization pruning by orders of magnitude with less than 5‰ and 15% error rates, respectively. SInC outperforms the state‐of‐the‐art from the perspective of overall compression on both scalability and compression effect. Ruoyu Wang 0004, Daniel Sun 0004, Raymond K. Wong 0001, Rajiv Ranjan 0001 |
Softw. Pract. Exp. | 1 |
| 2023 | Symbolic Minimization on Relational DataabstractThe current wave of AI is heavily driven by data, especially for cognitive capabilities. Minimization of data semantics not only reveals core information but also becomes a guide in a wide range of domains. However, scalability is theoretically weak in pure semantic methodologies. In order to cooperate with large DBs, expressiveness is over-sacrificed in existing techniques. Thus, the quality of discovered patterns and redundancies are far from satisfactory. In this article, we formalize symbolic minimization on relational DBs and prove its NP-Completeness. A lossless technique is proposed by inducing generic first-order Horn rules that infer a subset of records from the others. More importantly, we further improve the scalability via effective caching and pruning without sacrificing the expressiveness of first-order Horn rules. A concrete system is implemented and comprehensively evaluated. Experiments show that our technique removes up to 70% contents and outperforms the state-of-the-art on minimization and scalability. The optimizations reduce up to 96% memory consumption and accelerate the performance by two orders. Our technique shows the practicality of pure semantic approaches in database mining. Ruoyu Wang 0004, Daniel Sun 0004, Raymond K. Wong 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | RDF Knowledge Base Summarization by Inducing First-Order Horn Rules
Ruoyu Wang 0004, Daniel Sun 0004, Raymond K. Wong 0001 |
ECML/PKDD (2) | 1 |
| 2022 | SInC: Semantic approach and enhancement for relational data compression
Ruoyu Wang 0004, Daniel Sun 0004, Raymond K. Wong 0001, Rajiv Ranjan 0001, Albert Y. Zomaya |
Knowl. Based Syst. | 1 |
| 2020 | Statistical Detection Of Collective Data FraudabstractStatistical divergence is widely applied in multimedia processing, basically due to regularity and interpretable features displayed in data. However, in a broader range of data realm, these advantages may no longer be feasible, and therefore a more general approach is required. In data detection, statistical divergence can be used as a similarity measurement based on collective features. In this paper, we present a collective detection technique based on statistical divergence. The technique extracts distribution similarities among data collections, and then uses the statistical divergence to detect collective anomalies. Evaluation shows that it is applicable in the real world. Ruoyu Wang 0004, Daniel Sun 0004, Guoqiang Li 0001, Raymond K. Wong 0001, Shiping Chen 0001, Jianquan Liu |
ICME | 1 |
| 2020 | Pipeline provenance for cloud-based big data analyticsabstractSummary Provenance is information about the origin and creation of data. In data science and engineering related with cloud environment, such information is useful and sometimes even critical. In data analytics, it is necessary for making data‐driven decisions to trace back history and reproduce final or intermediate results, even to tune models and adjust parameters in a real‐time fashion. Particularly, in cloud, users need to evaluate data and pipeline trustworthiness. In this paper, we propose a solution: LogProv, toward realizing these functionalities for big data provenance, which needs to renovate data pipelines or some of big data software infrastructure to generate structured logs for pipeline events, and then stores data and logs separately in cloud space. The data are explicitly linked to the logs, which implicitly record pipeline semantics. Semantic information can be retrieved from the logs easily since they are well defined and structured beforehand. We implemented and deployed LogProv in Nectar Cloud,* associated with Apache Pig, Hadoop ecosystem, and adopted Elasticsearch to provide query service. LogProv was evaluated and empirically case studied. The results show that LogProv is efficient since the performance overhead is no more than 10%; the query can be responded within 1 second; the trustworthiness is marked clearly; and there is no impact on the data processing logic of original pipelines. Ruoyu Wang 0004, Daniel Sun 0004, Guoqiang Li 0001, Raymond K. Wong 0001, Shiping Chen 0001 |
Softw. Pract. Exp. | 1 |
| 2017 | Efficient Density-Based Blocking for Record MatchingabstractRecord Matching in data engineering refers to searching for data records originating from the same entities across different data sources. In practice, the main challenge of record matching is that the amount of non-matches typically far exceeds the amount of matches. This is called imbalance problem, which notoriously affects efficiency and effectiveness of matching algorithms. To solve the imbalance problem, recently, density-based blocking algorithms have been studied and demonstrated an effective blocking performance. However, the efficiency of density-based blocking approaches is not good as their effectiveness. In this paper, we improve the efficiency of density-based blocking by exploiting the idea of pre-computing and pruning. Our approach optimizes the method of computing density to speed up the blocking process. Throughout experiments on real-world datasets, the proposed approach demonstrated a high performance on both blocking efficiency and blocking effectiveness. Chenxiao Dou, Ruoyu Wang 0004, Daniel Sun 0004, Muhammad Atif 0003 |
IDEAS | 2 |
| 2016 | LogProv: Logging events as provenance of big data analytics pipelines with trustworthinessabstractProvenance is information about the origin and creation of data. In data science and engineering, such information is useful and sometimes even critical. In spite of that, provenance for big data is under-explored due to the challenges from the `Vs' of big data. In data analytics, users need to query history, reproduce intermediate or final results, tune models, and adjust parameters in runtime for making data-driven decisions. In addition, users need to evaluate data and pipeline trustworthiness. Towards realising these functionalities for big data provenance, we propose a solution, called LogProv, which needs to renovate data pipelines or even some of big data software infrastructure to generate structured logs for pipeline events, and then stores data and logs separately. The data are explicitly linked to the logs, which implicitly record pipeline semantics. Semantic information can be retrieved from the logs easily since the logs are well defined and structured beforehand. We implemented LogProv in Apache Pig, and adopted ElasticSearch to provide query service. In this paper LogProv is evaluated in a Hadoop ecosystem hosted by a cloud and empirically case-studied. The results show that LogProv is efficient since the performance overhead is no more than 10%, the query can be responded within 1 second, the trustworthiness is marked clearly, and there is no impact on the data processing logic of original pipelines. Ruoyu Wang 0004, Daniel Sun 0004, Guoqiang Li 0001, Muhammad Atif 0003, Surya Nepal |
IEEE BigData | 1 |