David Savage

dblp:03/5007 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
0since 2021 · last 2017
0000-0002-1610-6673ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › pattern mining › subgroup discovery
contrast set mining
0.312017
Distributed Mining of Contrast Patterns · IEEE Trans. Parallel Distributed Syst. 2017
Data mining › big data analytics › large-scale data mining
distributed data mining
0.312017
Distributed Mining of Contrast Patterns · IEEE Trans. Parallel Distributed Syst. 2017
Data mining
pattern mining
0.312017
Distributed Mining of Contrast Patterns · IEEE Trans. Parallel Distributed Syst. 2017
Data mining › predictive modeling
classification
0.112017
Distributed Mining of Contrast Patterns · IEEE Trans. Parallel Distributed Syst. 2017

Methods — techniques the papers use, named apart from their topics

search-space partitioning · 0.3map-reduce framework · 0.3
YearPublicationVenuePosition
2017 Distributed Mining of Contrast Patterns
abstract
In this paper we propose a novel algorithm for mining contrast patterns using a distributed, map-reduce like framework. Contrast patterns describe differences between contrasted data sets and have previously been used for building highly accurate classifiers. However, mining for contrast patterns is a computationally expensive task and existing algorithms are designed to run in a sequential manner on a single machine. Consequently, existing approaches are unable to handle dense, high volume and high dimensional databases. Our algorithm addresses this problem by partitioning the search-space for contrast patterns into small, independent units. These units can be mined in parallel, providing a scalable solution for mining large data sets. Using three different real-world data sets we test an implementation of our algorithm on a Spark cluster. Results of these tests indicate that our algorithm achieves a high-degree of parallelism and scalability.
David Savage, Xiuzhen Zhang 0001, Pauline Lin, Xinghuo Yu 0001, Qingmai Wang
IEEE Trans. Parallel Distributed Syst.1
2015 Detection of opinion spam based on anomalous rating deviation
David Savage, Xiuzhen Zhang 0001, Xinghuo Yu 0001, Pauline Lin, Qingmai Wang
Expert Syst. Appl.1