EDBT 2026 Demo / reviewers in the wild / expert
Pengxin Bian
dblp:409/8106
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2025
0009-0007-6463-1863ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › pattern mining
frequent pattern mining |
0.9 | 1 | 2025 | Resilient Pattern Mining · ICDM 2025 |
Data mining
pattern mining |
0.9 | 1 | 2025 | Resilient Pattern Mining · ICDM 2025 |
Algorithms and data structures › sequence algorithms
string algorithms |
0.9 | 1 | 2025 | Resilient Pattern Mining · ICDM 2025 |
Bioinformatics and computational biology › genomics
genomic data analysis |
0.3 | 1 | 2025 | Resilient Pattern Mining · ICDM 2025 |
Methods — techniques the papers use, named apart from their topics
enhanced suffix array · 2.6dynamic programming · 2.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Resilient Pattern MiningabstractFrequent pattern mining is a flagship problem in data mining. In its most basic form, it asks for the set of substrings of a given string$S$of length$n$that occur at least$\tau$times in$S$, for some integer$\tau\epsilon[1,n]$. We introduce a resilient version of this classic problem, which we term the$(\tau,\ k)$-Resilient Pattern Mining (rpm) problem. Given a string$S$of length$n$and two integers$\tau, k\in[1, n\vert$, RPM asks for the set of substrings of$S$that occur at least$\tau$times in$S$, even when the letters at any$k$positions of$S$are substituted by other letters. Unlike frequent substrings, resilient ones account for the fact that changes to string$S$are often expensive to handle or are unknown. We make the following contributions. First, we present RPM-DP, a simple exact$\mathrm{O}(n^{{3}}k\log n)$-time and$\mathrm{O}(n^{2})$-space algorithm for RPM that is based on an existing dynamic programming algorithm. Second, we propose RPM-ESA, an exact$\mathrm{O}(n\log n)$-time and$\mathrm{O}(n)$-space algorithm for RPM, which employs advanced data structures and combinatorial insights. Third, we conduct experiments on real large-scale datasets from different domains demonstrating that: (I) The notion of resilient substrings is useful in analyzing genomic data and fundamentally different from that of frequent substrings, as frequent substrings are often not resilient and thus do not remain frequent for long in versioned datasets; (II) RPM-ESA is several orders of magnitude faster and more space-efficient than RPM-DP; and (III) Clustering based on resilient substrings is effective. Pengxin Bian, Panagiotis Charalampopoulos, Lorraine A. K. Ayad, Manal Mohamed 0001, Solon P. Pissis, Grigorios Loukides |
ICDM | 1 |