EDBT 2026 Demo / reviewers in the wild / expert
Pauline Lin
dblp:07/2457 · also Pauline Chou, Pauline Lienhua Chou
· DBLP profile ↗
9ranked-venue papers
3as first author
1since 2021 · last 2022
0000-0001-6407-2028ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 1 first-authorSystems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data mining · 94% Query processing and optimization · 6% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
pattern mining |
0.4 | 2 | 2017 | Distributed Mining of Contrast Patterns · IEEE Trans. Parallel Distributed Syst. 2017 Efficient Computation of Iceberg Cubes by Bounding Aggregate Functions · IEEE Trans. Knowl. Data Eng. 2007 |
Data mining › pattern mining › subgroup discovery
contrast set mining |
0.3 | 1 | 2017 | Distributed Mining of Contrast Patterns · IEEE Trans. Parallel Distributed Syst. 2017 |
Data mining › big data analytics › large-scale data mining
distributed data mining |
0.3 | 1 | 2017 | Distributed Mining of Contrast Patterns · IEEE Trans. Parallel Distributed Syst. 2017 |
Data mining › predictive modeling
classification |
0.1 | 1 | 2017 | Distributed Mining of Contrast Patterns · IEEE Trans. Parallel Distributed Syst. 2017 |
Query processing and optimization
aggregate query processing |
0.1 | 1 | 2007 | Efficient Computation of Iceberg Cubes by Bounding Aggregate Functions · IEEE Trans. Knowl. Data Eng. 2007 |
Data mining › multidimensional data analysis
iceberg cube computation |
0.1 | 1 | 2007 | Efficient Computation of Iceberg Cubes by Bounding Aggregate Functions · IEEE Trans. Knowl. Data Eng. 2007 |
Methods — techniques the papers use, named apart from their topics
search-space partitioning · 0.3map-reduce framework · 0.3top-down cubing · 0.1bound pruning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Measurement of clustering effectiveness for document collectionsabstractAbstract Clustering of the contents of a document corpus is used to create sub-corpora with the intention that they are expected to consist of documents that are related to each other. However, while clustering is used in a variety of ways in document applications such as information retrieval, and a range of methods have been applied to the task, there has been relatively little exploration of how well it works in practice. Indeed, given the high dimensionality of the data it is possible that clustering may not always produce meaningful outcomes. In this paper we use a well-known clustering method to explore a variety of techniques, existing and novel, to measure clustering effectiveness. Results with our new, extrinsic techniques based on relevance judgements or retrieved documents demonstrate that retrieval-based information can be used to assess the quality of clustering, and also show that clustering can succeed to some extent at gathering together similar material. Further, they show that intrinsic clustering techniques that have been shown to be informative in other domains do not work for information retrieval. Whether clustering is sufficiently effective to have a significant impact on practical retrieval is unclear, but as the results show our measurement techniques can effectively distinguish between clustering methods. Justin Zobel, Pauline Lin |
Inf. Retr. J. | 3 |
| 2017 | Distributed Mining of Contrast PatternsabstractIn this paper we propose a novel algorithm for mining contrast patterns using a distributed, map-reduce like framework. Contrast patterns describe differences between contrasted data sets and have previously been used for building highly accurate classifiers. However, mining for contrast patterns is a computationally expensive task and existing algorithms are designed to run in a sequential manner on a single machine. Consequently, existing approaches are unable to handle dense, high volume and high dimensional databases. Our algorithm addresses this problem by partitioning the search-space for contrast patterns into small, independent units. These units can be mined in parallel, providing a scalable solution for mining large data sets. Using three different real-world data sets we test an implementation of our algorithm on a Spark cluster. Results of these tests indicate that our algorithm achieves a high-degree of parallelism and scalability. David Savage, Xiuzhen Zhang 0001, Pauline Lin, Xinghuo Yu 0001, Qingmai Wang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Detection of opinion spam based on anomalous rating deviation
David Savage, Xiuzhen Zhang 0001, Xinghuo Yu 0001, Pauline Lin, Qingmai Wang |
Expert Syst. Appl. | 4 |
| 2007 | Efficient Computation of Iceberg Cubes by Bounding Aggregate FunctionsabstractThe iceberg cubing problem is to compute the multidimensional group-by partitions that satisfy given aggregation constraints. Pruning unproductive computation for iceberg cubing when nonantimonotone constraints are present is a great challenge because the aggregate functions do not increase or decrease monotonically along the subset relationship between partitions. In this paper, we propose a novel bound prune cubing (BP-Cubing) approach for iceberg cubing with nonantimonotone aggregation constraints. Given a cube over n dimensions, an aggregate for any group-by partition can be computed from aggregates for the most specific n--dimensional partitions (MSPs). The largest and smallest aggregate values computed this way become the bounds for all partitions in the cube. We provide efficient methods to compute tight bounds for base aggregate functions and, more interestingly, arithmetic expressions thereof, from bounds of aggregates over the MSPs. Our methods produce tighter bounds than those obtained by previous approaches. We present iceberg cubing algorithms that combine bounding with efficient aggregation strategies. Our experiments on real-world and artificial benchmark data sets demonstrate that BP-Cubing algorithms achieve more effective pruning and are several times faster than state-of-the-art iceberg cubing algorithms and that BP-Cubing achieves the best performance with the top-down cubing approach. Xiuzhen Zhang 0001, Pauline Lin, Guozhu Dong |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2006 | Computing Iceberg Quotient Cubes with Bounding
Xiuzhen Zhang 0001, Pauline Lin, Kotagiri Ramamohanarao |
DaWaK | 2 |
| 2006 | Multiway Pruning for Efficient Iceberg Cubing
Xiuzhen Zhang 0001, Pauline Lin |
DEXA | 2 |
| 2005 | Multiway Iceberg Cubing on Trees
Pauline Lin, Xiuzhen Zhang 0001 |
WISE | 1 |
| 2004 | Computing Complex Iceberg Cubes by Multiway Aggregation and Bounding
Pauline Lin, Xiuzhen Zhang 0001 |
DaWaK | 1 |
| 2003 | Efficiently Computing Iceberg Cubes with Complex Constraints through Bounding
Pauline Lin, Xiuzhen Zhang 0001 |
PAKDD | 1 |