Junghee Ryu

dblp:206/7075 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 50% GPUs and heterogeneous computing · 50%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU and heterogeneous computing
0.312018
Randomized Algorithms Accelerated over CPU-GPU for Ultra-High Dimensional Similarity Search · SIGMOD Conference 2018
Storage systems
locality-sensitive hashing
0.312018
Randomized Algorithms Accelerated over CPU-GPU for Ultra-High Dimensional Similarity Search · SIGMOD Conference 2018
Algorithms and data structures › similarity search
high-dimensional similarity search
0.112018
Randomized Algorithms Accelerated over CPU-GPU for Ultra-High Dimensional Similarity Search · SIGMOD Conference 2018
Algorithms and data structures
similarity search
0.112018
Randomized Algorithms Accelerated over CPU-GPU for Ultra-High Dimensional Similarity Search · SIGMOD Conference 2018

Methods — techniques the papers use, named apart from their topics

reservoir sampling · 0.7minwise hashing · 0.7count-based estimation · 0.7
YearPublicationVenuePosition
2018 Randomized Algorithms Accelerated over CPU-GPU for Ultra-High Dimensional Similarity Search
abstract
We present FLASH (F ast L SH A lgorithm for S imilarity search accelerated with H PC), a similarity search system for ultra-high dimensional datasets on a single machine, that does not require similarity computations and is tailored for high-performance computing platforms. By leveraging a LSH style randomized indexing procedure and combining it with several principled techniques, such as reservoir sampling, recent advances in one-pass minwise hashing, and count based estimations, we reduce the computational and parallelization costs of similarity search, while retaining sound theoretical guarantees.
Yiqiu Wang, Anshumali Shrivastava, Junghee Ryu
SIGMOD Conference4