VLDB 2026 Research / reviewers in the wild / expert
Richard F. Helm
dblp:27/5481
· DBLP profile ↗
5ranked-venue papers
0as first author
0since 2021 · last 2010
0000-0001-5317-0925ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4Artificial intelligence and machine learning · 3Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Data mining · 100% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
pattern mining |
0.2 | 3 | 2008 | Algorithms for Storytelling · IEEE Trans. Knowl. Data Eng. 2008 Algorithms for storytelling · KDD 2006 Turning CARTwheels: an alternating algorithm for mining redescriptions · KDD 2004 |
Data mining › pattern mining › local pattern mining
redescription mining |
0.2 | 3 | 2008 | Algorithms for Storytelling · IEEE Trans. Knowl. Data Eng. 2008 Algorithms for storytelling · KDD 2006 Turning CARTwheels: an alternating algorithm for mining redescriptions · KDD 2004 |
Data mining
clustering |
0.1 | 1 | 2010 | Unifying dependent clustering and disparate clustering for non-homogeneous data · KDD 2010 |
Data mining › clustering
constrained clustering |
0.1 | 1 | 2010 | Unifying dependent clustering and disparate clustering for non-homogeneous data · KDD 2010 |
Mathematical optimization
constrained optimization |
0.1 | 1 | 2010 | Unifying dependent clustering and disparate clustering for non-homogeneous data · KDD 2010 |
Data mining › predictive modeling
classification |
0.0 | 1 | 2004 | Turning CARTwheels: an alternating algorithm for mining redescriptions · KDD 2004 |
Data mining › predictive modeling › classification
decision tree learning |
0.0 | 1 | 2004 | Turning CARTwheels: an alternating algorithm for mining redescriptions · KDD 2004 |
Bioinformatics and computational biology › functional genomics › functional enrichment analysis
gene set enrichment analysis |
0.0 | 2 | 2008 | Algorithms for Storytelling · IEEE Trans. Knowl. Data Eng. 2008 Algorithms for storytelling · KDD 2006 |
Methods — techniques the papers use, named apart from their topics
a* search · 0.3optimization framework · 0.2CARTwheels · 0.2CART wheels redescription mining · 0.1alternating optimization · 0.1CART · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2010 | Unifying dependent clustering and disparate clustering for non-homogeneous dataabstractModern data mining settings involve a combination of attribute-valued descriptors over entities as well as specified relationships between these entities. We present an approach to cluster such non-homogeneous datasets by using the relationships to impose either dependent clustering or disparate clustering constraints. Unlike prior work that views constraints as boolean criteria, we present a formulation that allows constraints to be satisfied or violated in a smooth manner. This enables us to achieve dependent clustering and disparate clustering using the same optimization framework by merely maximizing versus minimizing the objective function. We present results on both synthetic data as well as several real-world datasets. Mahmud Shahriar Hossain, Satish Tadepalli, Layne T. Watson, Ian Davidson, Richard F. Helm, Naren Ramakrishnan |
KDD | 5 |
| 2008 | Simultaneously Segmenting Multiple Gene Expression Time Courses by Analyzing Cluster Dynamics
Satish Tadepalli, Naren Ramakrishnan, Layne T. Watson, Bud Mishra, Richard F. Helm |
APBC | 5 |
| 2008 | Algorithms for StorytellingabstractWe formulate a new data mining problem called storytelling as a generalization of redescription mining. In traditional redescription mining, we are given a set of objects and a collection of subsets defined over these objects. The goal is to view the set system as a vocabulary and identify two expressions in this vocabulary that induce the same set of objects. Storytelling, on the other hand, aims to explicitly relate object sets that are disjoint (and hence, maximally dissimilar) by finding a chain of (approximate) redescriptions between the sets. This problem finds applications in bioinformatics, for instance, where the biologist is trying to relate a set of genes expressed in one experiment to another set, implicated in a different pathway. We outline an efficient storytelling implementation that embeds the CARTwheels redescription mining algorithm in an A* search procedure, using the former to supply next move operators on search branches to the latter. This approach is practical and effective for mining large datasets and, at the same time, exploits the structure of partitions imposed by the given vocabulary. Three application case studies are presented: a study of word overlaps in large English dictionaries, exploring connections between genesets in a bioinformatics dataset, and relating publications in the PubMed index of abstracts. Deept Kumar, Naren Ramakrishnan, Richard F. Helm, Malcolm Potts |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2006 | Algorithms for storytellingabstractWe formulate a new data mining problem called it storytelling as a generalization of redescription mining. In traditional redescription mining, we are given a set of objects and a collection of subsets defined over these objects. The goal is to view the set system as a vocabulary and identify two expressions in this vocabulary that induce the same set of objects. Storytelling, on the other hand, aims to explicitly relate object sets that are disjoint (and hence, maximally dissimilar) by finding a chain of (approximate) redescriptions between the sets. This problem finds applications in bioinformatics, for instance, where the biologist is trying to relate a set of genes expressed in one experiment to another set, implicated in a different pathway. We outline an efficient storytelling implementation that embeds the CART wheels redescription mining algorithm in an A* search procedure, using the former to supply next move operators on search branches to the latter. This approach is practical and effective for mining large datasets and, at the same time, exploits the structure of partitions imposed by the given vocabulary. Three application case studies are presented: a study of word overlaps in large English dictionaries, exploring connections between genesets in a bioinformatics dataset, and relating publications in the PubMed index of abstracts. Deept Kumar, Naren Ramakrishnan, Richard F. Helm, Malcolm Potts |
KDD | 3 |
| 2004 | Turning CARTwheels: an alternating algorithm for mining redescriptionsabstractWe present an unusual algorithm involving classification trees---CARTwheels---where two trees are grown in opposite directions so that they are joined at their leaves. This approach finds application in a new data mining task we formulate, called redescription mining. A redescription is a shift-of-vocabulary, or a different way of communicating information about a given subset of data; the goal of redescription mining is to find subsets of data that afford multiple descriptions. We highlight the importance of this problem in domains such as bioinformatics, which exhibit an underlying richness and diversity of data descriptors (e.g., genes can be studied in a variety of ways). CARTwheels exploits the duality between class partitions and path partitions in an induced classification tree to model and mine redescriptions. It helps integrate multiple forms of characterizing datasets, situates the knowledge gained from one dataset in the context of others, and harnesses high-level abstractions for uncovering cryptic and subtle features of data. Algorithm design decisions, implementation details, and experimental results are presented. Naren Ramakrishnan, Deept Kumar, Bud Mishra, Malcolm Potts, Richard F. Helm |
KDD | 5 |