Ryan Wong 0002

dblp:198/1917-2 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2021
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 56% Distributed systems · 44%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
distributed data processing
0.512021
A Serverless Framework for Distributed Bulk Metadata Extraction · HPDC 2021
Cloud and datacenter computing
serverless computing
0.512021
A Serverless Framework for Distributed Bulk Metadata Extraction · HPDC 2021
Cloud and datacenter computing
cloud federation
0.112021
A Serverless Framework for Distributed Bulk Metadata Extraction · HPDC 2021

Methods — techniques the papers use, named apart from their topics

container-based execution · 0.5
YearPublicationVenuePosition
2021 A Serverless Framework for Distributed Bulk Metadata Extraction
abstract
We introduce Xtract, an automated and scalable system for bulk metadata extraction from large, distributed research data repositories. Xtract orchestrates the application of metadata extractors to groups of files, determining which extractors to apply to each file and, for each extractor and file, where to execute. A hybrid computing model, built on the funcX federated FaaS platform, enables Xtract to balance tradeoffs between extraction time and data transfer costs by dispatching each extraction task to the most appropriate location. Experiments on a range of clouds and supercomputers show that Xtract can efficiently process multi-million-file repositories by orchestrating the concurrent execution of container-based extractors on thousands of nodes. We highlight the flexibility of Xtract by applying it to a large, semi-curated scientific data repository and to an uncurated scientific Google Drive repository. We show that by remotely orchestrating metadata extraction across decentralized storage and compute nodes, Xtract can process large repositories in 50% of the time it takes just to transfer the same data to a machine within the same computing facility. We also show that when transferring data is necessary (e.g., no local compute is available), Xtract can scale to process files as fast as they are received, even over a multi-GB/s network.
Tyler J. Skluzacek, Ryan Wong 0002, Zhuozhao Li, Ryan Chard, Kyle Chard, Ian T. Foster
HPDC2