Finding Non-Redundant Simpson's Paradox in Multidimensional Data
Abstract
Simpson's paradox has broad impact across many scientific domains. Existing detection methods overlook a key issue: many detected paradoxes may be redundant, arising from equivalent data subsets, identical subpopulation partitions, or correlated outcome variables, thereby obscure insights and increase computational cost. In this paper, we present a framework for finding non-redundant Simpson's paradoxes by formalizing three sources of redundancy—sibling child, separator, and statistic equivalence—and showing that pairwise redundancy forms an equivalence relation. We further propose a concise representation that groups redundant paradoxes and develop efficient algorithms combining depth-first population materialization with redundancy-aware discovery. Experiments on real and synthetic datasets show that redundancy is prevalent (over 40% in some cases), while our methods scale to millions of records, achieve up to 6.72 times speedup over brute-force approaches and identify robust paradoxes, enabling efficient discovery, compact summarization, and clear interpretation in multidimensional data.
Assigned reviewers
No reviewers assigned yet.
Candidates from the panel ranked by taxonomy affinity
| # | Reviewer | Match | Load | Why |
|---|