Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jiaye Zheng

dblp:411/0708 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0002-4572-104XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 87% Data integration and cleaning · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
pattern mining
0.912025
Incremental Rule Discovery in Response to Parameter Updates · Proc. ACM Manag. Data 2025
Data mining › pattern mining
rule mining
0.912025
Incremental Rule Discovery in Response to Parameter Updates · Proc. ACM Manag. Data 2025
Data integration and cleaning › data quality
data quality rules
0.312025
Incremental Rule Discovery in Response to Parameter Updates · Proc. ACM Manag. Data 2025

Methods — techniques the papers use, named apart from their topics

parallelization · 0.9incremental algorithm · 0.9
YearPublicationVenuePosition
2025 Incremental Rule Discovery in Response to Parameter Updates
abstract
This paper studies incremental rule discovery. Given a dataset D, rule discovery is to mine the set of the rules on D such that their supports and confidences are above thresholds 𝜎 and 𝛅 , respectively. We formulate incremental problems in response to updates Δ𝜎 and/or Δ𝛅, to compute rules added and/or removed with respect to 𝜎 + Δ𝜎 and 𝛅 + Δ𝛅. The need for studying the problems is evident since practitioners often want to adjust their support and confidence thresholds during discovery. The objective is to minimize unnecessary recomputation during the adjustments, not to restart the costly discovery process from scratch. As a testbed, we consider entity enhancing rules, which subsume popular data quality rules as special cases. We develop three incremental algorithms, in response to Δ𝜎 , Δ𝜎 and both. We show that relative to a batch discovery algorithm, these algorithms are bounded, i.e., they incur the minimum cost among all incrementalizations of the batch one, and parallelly scalable, i.e., they guarantee to reduce runtime when given more processors. Using real-life data, we empirically verify that the incremental algorithms outperform the batch counterpart by up to 658× when Δ𝜎 and Δ𝜎 are either positive or negative.
Haoxian Chen 0001, Wenfei Fan, Jiaye Zheng
Proc. ACM Manag. Data3