VLDB 2026 Research / reviewers in the wild / expert
Jack D. Marquez
dblp:243/0975
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-2673-3507ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ParaDyMS: Parallel Dynamic Motif Counting at ScaleabstractMotifs in graphs (or networks) are subgraphs induced by a small set of vertices such as triangles and cliques. The frequency of motifs is used to compare and align networks across various domains, such as biology, epidemiology, and social sciences. Recent advances have made it feasible to solve the computational challenge of counting motifs in networks with over a billion edges. However, these algorithms apply only for static networks where the structure remains unchanged. In reality, networks dynamically evolve, and understanding how motifs change with the dynamic nature of these networks remains an unsolved challenge despite its potential to provide essential insights into the system. We present ParaDyMS (Parallel Dynamic Motif Counting at Scale), the first parallel algorithm for updating motif counts in fully dynamic networks using batched updates. Our algorithm updates the frequencies of motifs only in the modified parts of the network instead of recomputing them from scratch. We provide proof of the algorithm's correctness and complexity and empirically compare its execution time with another state-of-the-art static algorithm on shared memory and GPUs using realworld networks. Our results show that our algorithm is highly scalable and can significantly reduce the time to compute motifs by more than 90% in the best case and, on average, by 69%. Nigel Tan, Jack D. Marquez, Michela Taufer, Sanjukta Bhowmick |
CCGrid | 3 |
| 2025 | Visualizing Temperature Hotspots and Their Impact Using the NEX-GDDP-CMIP6 DatasetabstractClimate events like tornadoes, hurricanes, and droughts are major twenty-first-century challenges. Mitigation efforts lag due to costly monitoring tools and complex data access. We present a dashboard with long-term temperature and weather trends, using nine global metrics from the NEX-GDDP-CMIP6 dataset, projecting to 2100. Our analysis flags countries facing major wet-bulb temperature hikes, linking these to expected economic impacts. Initial findings suggest that, in a worst-case scenario, some Northern Hemisphere nations may experience over a 10°C wet-bulb temperature rise between 2020 and 2090, potentially causing GDP losses exceeding 30%. Walter J. Ashworth, Jack D. Marquez, Michela Taufer |
eScience | 2 |
| 2025 | GEOtiled-SG: A Scalable Framework for High-Resolution Terrain Parameter ComputationabstractGEOtiled enables efficient computation of high-resolution terrain parameters from digital elevation models (DEMs) by decomposing large regions into smaller, parallelizable tiles. Originally developed as GEOtiled-G with support for three parameters (Slope, Aspect, Hillshade) via GDAL, we present GEOtiled-SG, an enhanced version that integrates the SAGA GIS library to compute over 15 parameters, expanding its utility in Earth science. To offset SAGA’s computational overhead, GEOtiled-SG introduces three optimizations: concurrent DEM cropping, buffer-aware mosaicking, and unified concurrency across workflow stages. Evaluations show that GEOtiled-SG maintains GEOtiled-G’s performance on the original parameters and offers consistent speedups across the expanded set. The framework is open source on GitHub, with data hosted on Dataverse, supporting reproducible, scalable terrain analysis. Gabriel Laboy, Ian Lumsden, Paula Olaya, Jack D. Marquez, Kin Wai Ng, Rodrigo Vargas, Michela Taufer |
eScience | 4 |
| 2025 | Advancing the GEOtiled Framework Through Scalable Terrain Parameter ComputationabstractThe GEOtiled framework facilitates the scalable and efficient computation of high-resolution terrain parameters using Digital Elevation Models (DEMs) across the Continental United States (CONUS). These parameters are essential in Earth Science applications. This paper presents significant advances to optimizing GEOtiled's performance. In GEOtiled 2.0, we introduce three key optimizations to speed up GEOtiled's overall runtime: concurrent cropping, efficient mosaic operation, and unified parallel processing. These optimizations improve data distribution and handling, reducing computation times while maintaining accuracy. We evaluate performance using DEM data at 30-meter resolution over a region covering the state of Tennessee. Our results demonstrate that the optimized framework provides substantial performance improvements for generating terrain parameters. GEOtiled 2.0 software, openly available on GitHub, advances reproducible, data-driven scientific discovery in Earth Science. Gabriel Laboy, Paula Olaya, Jack D. Marquez, Michael Sutherlin, Rodrigo Vargas, Michela Taufer |
HPDC | 3 |
| 2023 | Online Boosted Gaussian Learners for In-Situ Detection and Characterization of Protein Folding States in Molecular Dynamics SimulationsabstractMolecular Dynamics (MD) simulations are a crucial tool for understanding how proteins fold. In its easiest form, MD simulations can be scaled through data parallelism, this means that multiple folding trajectories can be spawned and executed in parallel, facilitating a more efficient exploration of the protein folding space. However, due to data dependencies, the analysis of MD simulations remains largely as a centralized process. In this work, we propose a data parallel, lightweight technique to learn the characteristics of protein folding states in MD simulations. Contrary to other methods, ours can differentiate relevant states in a single protein folding trajectory without requiring centralized global knowledge of the protein dynamics. As its processing and memory overheads are negligible (in the order of milliseconds per window of frames, and kilo bytes respectively) this technique can be coupled with the simulation for in-situ analysis. Harshita Sahni, Hector Carrillo-Cabada, Ekaterina D. Kots, Silvina Caíno-Lores, Jack D. Marquez, Ewa Deelman, Michel A. Cuendet, Harel Weinstein, Michela Taufer, Trilce Estrada |
e-Science | 5 |
| 2023 | Scalable Incremental Checkpointing using GPU-Accelerated De-DuplicationabstractWriting large amounts of data concurrently to stable storage is a typical I/O pattern of many HPC workflows. This pattern introduces high I/O overheads and results in increased storage space utilization especially for workflows that need to capture the evolution of data structures with high frequency as checkpoints. In this context, many applications, such as graph pattern matching, perform sparse updates to large data structures between checkpoints. For these applications, incremental checkpointing techniques that save only the differences from one checkpoint to another can dramatically reduce the checkpoint sizes, I/O bottlenecks, and storage space utilization. However, such techniques are not without challenges: it is non-trivial to transparently determine what data has changed since a previous checkpoint and assemble the differences in a compact fashion that does not result in excessive metadata. State-of-art data reduction techniques (e.g., compression and de-duplication) have significant limitations when applied to modern HPC applications that leverage GPUs: slow at detecting the differences, generate a large amount of metadata to keep track of the differences, and ignore crucial spatiotemporal checkpoint data redundancy. This paper addresses these challenges by proposing a Merkle tree-based incremental checkpointing method to exploit GPUs’ high memory bandwidth and massive parallelism. Experimental results at scale show a significant reduction of the I/O overhead and space utilization of checkpointing compared with state-of-the-art incremental checkpointing and compression techniques. Nigel Tan, Jakob Lüttgau, Jack D. Marquez, Keita Teranishi, Nicolas M. Morales, Sanjukta Bhowmick, Franck Cappello, Michela Taufer, Bogdan Nicolae |
ICPP | 3 |