VLDB 2026 Research / reviewers in the wild / expert
Gustavo Henrique Orair
dblp:49/147
· DBLP profile ↗
2ranked-venue papers
1as first author
0since 2021 · last 2010
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
anomaly detection |
0.1 | 1 | 2010 | Distance-Based Outlier Detection: Consolidation and Renewed Bearing · Proc. VLDB Endow. 2010 |
Data mining › anomaly detection › outlier detection
distance-based outlier detection |
0.1 | 1 | 2010 | Distance-Based Outlier Detection: Consolidation and Renewed Bearing · Proc. VLDB Endow. 2010 |
Methods — techniques the papers use, named apart from their topics
factorial design experiment · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2010 | Distance-Based Outlier Detection: Consolidation and Renewed BearingabstractDetecting outliers in data is an important problem with interesting applications in a myriad of domains ranging from data cleaning to financial fraud detection and from network intrusion detection to clinical diagnosis of diseases. Over the last decade of research, distance-based outlier detection algorithms have emerged as a viable, scalable, parameter-free alternative to the more traditional statistical approaches. In this paper we assess several distance-based outlier detection approaches and evaluate them. We begin by surveying and examining the design landscape of extant approaches, while identifying key design decisions of such approaches. We then implement an outlier detection framework and conduct a factorial design experiment to understand the pros and cons of various optimizations proposed by us as well as those proposed in the literature, both independently and in conjunction with one another, on a diverse set of real-life datasets. To the best of our knowledge this is the first such study in the literature. The outcome of this study is a family of state of the art distance-based outlier detection algorithms. Our detailed empirical study supports the following observations. The combination of optimization strategies enables significant efficiency gains. Our factorial design study highlights the important fact that no single optimization or combination of optimizations (factors) always dominates on all types of data. Our study also allows us to characterize when a certain combination of optimizations is likely to prevail and helps provide interesting and useful insights for moving forward in this domain. Gustavo Henrique Orair, Carlos H. C. Teixeira, Wagner Meira Jr., Srinivasan Parthasarathy 0001 |
Proc. VLDB Endow. | 1 |
| 2006 | ParTriCluster: A Scalable Parallel Algorithm for Gene Expression AnalysisabstractAnalyzing gene expression patterns is becoming a highly relevant task in the bio informatics area. This analysis makes it possible to determine the behavior patterns of genes under various conditions, a fundamental information for treating diseases, among other applications. An advance in this area is the tricluster algorithm, which is the first algorithm capable of determining 3D clusters, that is, it determines clusters of sets of genes that behave similarly in a set of samples and set of time stamps. However, while biological experiments collect an increasing amount of data to be analyzed and correlated, the triclustering problem is NP-complete, and its parallelization seems to be an essential step towards obtaining feasible solutions. In this paper we propose and evaluate the implementation of a parallel version of the tricluster algorithm using the filter-labeled-stream paradigm supported by the Anthill parallel programming environment. The results show that our parallelization scales linearly with the data size. Further, the parallelization strategy is applicable to any depth-first searches Renata Braga Araújo, Guilherme Henrique Trielli Ferreira, Gustavo Henrique Orair, Wagner Meira Jr., Renato Ferreira 0001, Dorgival O. Guedes |
SBAC-PAD | 3 |