Muyun Zhou

dblp:414/0534 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data integration and cleaning · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning › data preprocessing › data cleaning
constraint-based data cleaning
0.912025
Cleaning both Data Errors and Inaccurate Constraints on Numerical Sequential Data · Proc. VLDB Endow. 2025
Data integration and cleaning › integrity constraint discovery
denial constraint discovery
0.912025
$t$DCDiscover: Mining Threshold Denial Constraints from Time Series Data · ICDE 2025
Data integration and cleaning
dependency discovery
0.912025
$t$DCDiscover: Mining Threshold Denial Constraints from Time Series Data · ICDE 2025
Data integration and cleaning › data preprocessing › data cleaning
error detection and repair
0.312025
$t$DCDiscover: Mining Threshold Denial Constraints from Time Series Data · ICDE 2025
Data integration and cleaning › data preprocessing
time series data cleaning
0.312025
Cleaning both Data Errors and Inaccurate Constraints on Numerical Sequential Data · Proc. VLDB Endow. 2025

Methods — techniques the papers use, named apart from their topics

pruning strategies · 0.9evidence matrix · 0.9
YearPublicationVenuePosition
2026 Procore: Robust Core-Set Selection Via Pareto Multi-Dimensional Optimization From Noisy Data
Xiaoou Ding, Hongbin Hu, Songnan Jiang, Muyun Zhou, Chen Wang 0018, Jingru Yang, Hongzhi Wang 0001
ICDE4
2025 $t$DCDiscover: Mining Threshold Denial Constraints from Time Series Data
abstract
Denial constraints are vital in data quality management, but traditional mining algorithms struggle with time series data. To address this, we introduce a novel data quality rule, threshold Denial Constraints ($t$DCs), which enables predicate scaling in numerical contexts. We formalize the inference system for$t$DCs and demonstrate the monotonicity and abruptness of threshold predicates. To efficiently mine$t$DCs, we design the tDCDiscover algorithm, which leverages batch computation of differences and thresholds to significantly reduce the time required for acquiring homologous predicate evidence, achieving a 50% -66% decrease. Additionally, we introduce an evidence matrix to store evidence, lowering the complexity of evidence matching from$O(m)$to$O(1)$. We propose two pruning strategies: triviality pruning and prediction coverage pruning, to effectively decrease the search paths to one-fifth of their original number and eliminating at least 90% of unnecessary paths. We theoretically prove that tDCDiscover ensures minimal, valid, and complete results. Experimental results on eight real-world datasets demonstrate that, compared to the current state-of-the-art denial constraint mining techniques, tDCDiscover achieves more than double the efficiency when processing high-dimensional time series data. In downstream data cleaning tasks, tDCDiscover improves error detection precision by an average of 40% and repair accuracy by 18%, further offering advantages in time series data quality management.
Xiaoou Ding, Muyun Zhou, Yida Liu, Zekai Qian, Chen Wang 0018, Hongzhi Wang 0001, Jianmin Wang 0001
ICDE2
2025 Cleaning both Data Errors and Inaccurate Constraints on Numerical Sequential Data
Xiaoou Ding, Muyun Zhou, Yida Liu, Chen Wang 0018, Hongzhi Wang 0001, Jianmin Wang 0001
Proc. VLDB Endow.2