Adrian Nicoara

dblp:247/8054 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 67% Parallel and multicore computing · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › distributed data processing
data shuffling
0.412019
Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in Scope · Proc. VLDB Endow. 2019
Distributed systems › distributed database
distributed query processing
0.412019
Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in Scope · Proc. VLDB Endow. 2019
Parallel and multicore computing
parallel query processing
0.412019
Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in Scope · Proc. VLDB Endow. 2019

Methods — techniques the papers use, named apart from their topics

recursive partitioning · 0.4divide-and-conquer · 0.4
YearPublicationVenuePosition
2019 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in Scope
abstract
In distributed query processing, data shuffle is one of the most costly operations. We examined scaling limitations to data shuffle that current systems and the research literature do not solve. As the number of input and output partitions increases, naïve shuffling will result in high fan-out and fan-in. There are practical limits to fan-out, as a consequence of limits on memory buffers, network ports and I/O handles. There are practical limits to fan-in because it multiplies the communication errors due to faults in commodity clusters impeding progress. Existing solutions that limit fan-out and fan-in do so at the cost of scaling quadratically in the number of nodes in the data flow graph. This dominates the costs of shuffling large datasets. We propose a novel algorithm called Hyper Dimension Shuffle that we have introduced in production in SCOPE, Microsoft's internal big data analytics system. Hyper Dimension Shuffle is inspired by the divide and conquer concept, and utilizes a recursive partitioner with intermediate aggregations. It yields quasilinear complexity of the shuffling graph with tight guarantees on fan-out and fan-in. We demonstrate how it avoids the shuffling graph blow-up of previous algorithms to shuffle at petabyte-scale efficiently on both synthetic benchmarks and real applications.
Shi Qiao 0001, Adrian Nicoara, Marc T. Friedman, Hiren Patel, Jaliya Ekanayake
Proc. VLDB Endow.2