VLDB 2026 Research / reviewers in the wild / expert
Ryan Wong 0002
dblp:198/1917-2
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2021
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 56% Distributed systems · 44% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
distributed data processing |
0.5 | 1 | 2021 | A Serverless Framework for Distributed Bulk Metadata Extraction · HPDC 2021 |
Cloud and datacenter computing
serverless computing |
0.5 | 1 | 2021 | A Serverless Framework for Distributed Bulk Metadata Extraction · HPDC 2021 |
Cloud and datacenter computing
cloud federation |
0.1 | 1 | 2021 | A Serverless Framework for Distributed Bulk Metadata Extraction · HPDC 2021 |
Methods — techniques the papers use, named apart from their topics
container-based execution · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | A Serverless Framework for Distributed Bulk Metadata ExtractionabstractWe introduce Xtract, an automated and scalable system for bulk metadata extraction from large, distributed research data repositories. Xtract orchestrates the application of metadata extractors to groups of files, determining which extractors to apply to each file and, for each extractor and file, where to execute. A hybrid computing model, built on the funcX federated FaaS platform, enables Xtract to balance tradeoffs between extraction time and data transfer costs by dispatching each extraction task to the most appropriate location. Experiments on a range of clouds and supercomputers show that Xtract can efficiently process multi-million-file repositories by orchestrating the concurrent execution of container-based extractors on thousands of nodes. We highlight the flexibility of Xtract by applying it to a large, semi-curated scientific data repository and to an uncurated scientific Google Drive repository. We show that by remotely orchestrating metadata extraction across decentralized storage and compute nodes, Xtract can process large repositories in 50% of the time it takes just to transfer the same data to a machine within the same computing facility. We also show that when transferring data is necessary (e.g., no local compute is available), Xtract can scale to process files as fast as they are received, even over a multi-GB/s network. Tyler J. Skluzacek, Ryan Wong 0002, Zhuozhao Li, Ryan Chard, Kyle Chard, Ian T. Foster |
HPDC | 2 |