VLDB 2026 Research / reviewers in the wild / expert
Thomas D. Uram
dblp:54/7282
· DBLP profile ↗
13ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Synthesis of Ground Truth for Neuronal SegmentationabstractSegmenting neuropil data for connectomics is limited by the high cost and effort of manually labeling Electron Microscopy (EM) volumes. To address this, we propose a two-stage framework to generate large amounts of realistic image data with ground-truth segmentations. First, we use Procedural Content Generation (PCG) with a parametric evolution strategy to synthesize diverse 3D neuron morphologies, which are combined into dense, noise-free segmentation masks. Second, these masks guide a conditional denoising diffusion model (DDPM) to generate realistic EM image patches, trained adversarially with a PatchGAN discriminator. This approach enables scalable production of high-quality training data, aiming to improve model accuracy and reduce manual annotation efforts. Jyotsna Rajaraman, Thomas D. Uram, Kevin M. Boergens, Michael E. Papka |
eScience | 2 |
| 2021 | Enabling discovery data science through cross-facility workflowsabstractExperimental and observational instruments for scientific research (such as light sources, genome sequencers, accelerators, telescopes and electron microscopes) increasingly require High Performance Computing (HPC) scale capabilities for data analysis and workflow processing. Next-generation instruments are being deployed with higher resolutions and faster data capture rates, creating a big data crunch that cannot be handled by modest institutional computing resources. Often these big data analysis pipelines also require near real-time computing and have higher resilience requirements than the simulation and modeling workloads more traditionally seen at HPC centers. While some facilities have enabled workflows to run at a single HPC facility, there is a growing need to integrate capabilities across HPC facilities to enable cross-facility workflows, either to provide resilience to an experiment, increase analysis throughput capabilities, or to better match a workflow to a particular architecture. In this paper we describe the barriers to executing complex data analysis workflows across HPC facilities and propose an architectural design pattern for enabling scientific discovery using cross-facility workflows that includes orchestration services, application programming interfaces (APIs), data access and co-scheduling. Katie Antypas, Deborah Bard, Johannes P. Blaschke, Shane Canon, Bjoern Enders, Mallikarjun Shankar, Suhas Somnath, Dale Stansberry, Thomas D. Uram, Sean R. Wilkinson |
IEEE BigData | 9 |
| 2021 | Extreme Scale Survey Simulation with Python WorkflowsabstractThe Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) will soon carry out an unprecedented wide, fast, and deep survey of the sky in multiple optical bands. The data from LSST will open up a new discovery space in astronomy and cosmology, simultaneously providing clues toward addressing burning issues of the day, such as the origin of dark energy and and the nature of dark matter, while at the same time yielding data that will, in turn, pose fresh new questions. To prepare for the imminent arrival of this remarkable data set, it is crucial that the associated scientific communities be able to develop the software needed to analyze it. Computational power now available allows us to generate synthetic data sets that can be used as a realistic training ground for such an effort. This effort raises its own challenges—the need to generate very large simulations of the night sky, scaling up simulation campaigns to large numbers of compute nodes across multiple computing centers with different architectures, and optimizing the complex workload around memory requirements and widely varying wall clock times. We describe here a large-scale workflow that melds together Python code to steer the workflow, Parsl to manage the large-scale distributed execution of workflow components, and containers to carry out the image simulation campaign across multiple sites. Taking advantage of these tools, we developed an extreme-scale computational framework and used it to simulate five years of observations for 300 square degrees of sky area. We describe our experiences and lessons learned in developing this workflow capability, and highlight how the scalability and portability of our approach enabled us to efficiently execute it on up to 4000 compute nodes on two supercomputers. A. S. Villarreal, Yadu N. Babuji, Thomas D. Uram, Daniel S. Katz, Kyle Chard, Katrin Heitmann |
e-Science | 3 |
| 2021 | AxonEM Dataset: 3D Axon Instance Segmentation of Brain Cortical Regions
Donglai Wei 0001, Kisuk Lee, J. Alexander Bae, Zequan Liu, Márcia dos Santos, Zudi Lin, Thomas D. Uram, Xueying Wang 0002, Ignacio Arganda-Carreras, Brian Matejek, Narayanan Kasthuri, Jeff Lichtman, Hanspeter Pfister |
MICCAI (1) | 10 |
| 2018 | DeepHyper: Asynchronous Hyperparameter Search for Deep Neural NetworksabstractHyperparameters employed by deep learning (DL) methods play a substantial role in the performance and reliability of these methods in practice. Unfortunately, finding performance optimizing hyperparameter settings is a notoriously difficult task. Hyperparameter search methods typically have limited production-strength implementations or do not target scalability within a highly parallel machine, portability across different machines, experimental comparison between different methods, and tighter integration with workflow systems. In this paper, we present DeepHyper, a Python package that provides a common interface for the implementation and study of scalable hyperparameter search methods. It adopts the Balsam workflow system to hide the complexities of running large numbers of hyperparameter configurations in parallel on high-performance computing (HPC) systems. We implement and study asynchronous model-based search methods that consist of sampling a small number of input hyperparameter configurations and progressively fitting surrogate models over the input-output space until exhausting a user-defined budget of evaluations. We evaluate the efficacy of these methods relative to approaches such as random search, genetic algorithms, Bayesian optimization, and hyperband on DL benchmarks on CPU-and GPU-based HPC systems. Prasanna Balaprakash, Michael Salim, Thomas D. Uram, Venkat Vishwanath, Stefan M. Wild |
HiPC | 3 |
| 2018 | Scalable pCT Image Reconstruction Delivered as a Cloud ServiceabstractWe describe a cloud-based medical image reconstruction service designed to meet a real-time and daily demand to reconstruct thousands of images from proton cancer treatment facilities worldwide. Rapid reconstruction of a three-dimensional Proton Computed Tomography (pCT) image can require the transfer of 100 GB of data and use of approximately 120 GPU-enabled compute nodes. The nature of proton therapy means that demand for such a service is sporadic and comes from potentially hundreds of clients worldwide. We thus explore the use of a commercial cloud as a scalable and cost-efficient platform for pCT reconstruction. To address the high performance requirements of this application we leverage Amazon Web Services' GPU-enabled cluster resources that are provisioned with high performance networks between nodes. To support episodic demand, we develop an on-demand multi-user provisioning service that can dynamically provision and resize clusters based on image reconstruction requirements, priorities, and wait times. We compare the performance of our pCT reconstruction service running on commercial cloud resources with that of the same application on dedicated local high performance computing resources. We show that we can achieve scalable and on-demand reconstruction of large scale pCT images for simultaneous multi-client requests, processing images in less than 10 minutes for less than $10 per image. Ryan Chard, Ravi K. Madduri, Nicholas T. Karonis, Kyle Chard, Kirk L. Duffin, Caesar E. Ordoñez, Thomas D. Uram, Justin Fleischauer, Ian T. Foster, Michael E. Papka, John Winans |
IEEE Trans. Cloud Comput. | 7 |
| 2013 | Distributed and hardware accelerated computing for clinical medical imaging using proton computed tomography (pCT)
Nicholas T. Karonis, Kirk L. Duffin, Caesar E. Ordoñez, Béla Erdélyi, Thomas D. Uram, Eric C. Olson, George Coutrakon, Michael E. Papka |
J. Parallel Distributed Comput. | 5 |
| 2011 | GROPHECY: GPU performance projection from CPU code skeletonsabstractWe propose GROPHECY, a GPU performance projection framework that can estimate the performance benefit of GPU acceleration without actual GPU programming or hardware. Users need only to skeletonize pieces of CPU code that are targets for GPU acceleration. Code skeletons are automatically transformed in various ways to mimic tuned GPU codes with characteristics resembling real implementations. The synthesized characteristics are used by an existing analytical model to project GPU performance. The cost and benefit of GPU development can then be estimated according to the transformed code skeleton that yields the best projected performance. With GROPHECY, users can leap toward GPU acceleration only when the cost-benefit makes sense. The framework is validated using kernel benchmarks and data-parallel codes in legacy scientific applications. The measured performance of manually tuned codes deviates from the projected performance by 17% in geometric mean. Jiayuan Meng, Vitali A. Morozov, Kalyan Kumaran, Venkatram Vishwanath, Thomas D. Uram |
SC | 5 |
| 2010 | A Web 2.0-Based Scientific Application FrameworkabstractA significant obstacle to building usable, web-based interfaces for computational science in a Grid environment is how to deploy scientific applications on computational resources and expose these applications as web services. To streamline the development of these interfaces, we propose a new application framework that can deliver user-defined scientific workflows as both web services and OpenSocial gadgets. Through this application framework, scientists can focus on defining computational workflows using domain-specific applications and can use the software tools in the framework to quickly generate gadgets for running the applications and visualizing the output from workflow executions. By assembling these domain-specific gadgets and some common gadgets predefined in the framework for workflow management, scientists can easily set up a customized computational workspace to meet their requirements. Wenjun Wu 0001, Thomas D. Uram, Michael Wilde, Mark Hereld, Michael E. Papka |
ICWS | 2 |
| 2009 | A hybrid multicast connectivity solution for multi-party collaborative environments
Namgon Kim, Jongwon Kim 0001, Thomas D. Uram |
Multim. Tools Appl. | 3 |
| 2008 | The Problem Solving Environments of TeraGrid, Science Gateways, and the Intersection of the TwoabstractProblem solving environments (PSEs) are increasingly important for scientific discovery. Today's most challenging problems often require multi-disciplinary teams, the ability to analyze very large amounts of data, and the need to rely on infrastructure built by others rather than reinventing solutions for each science team. The TeraGrid Science Gateways program recognizes these challenges and works with science teams to harness high-end resources that significantly extend a PSE's functionality. Jim Basney, Stuart Martin, John-Paul Navarro, Marlon E. Pierce, Tom Scavo, Leif Strand, Thomas D. Uram, Nancy Wilkins-Diehr, Wenjun Wu 0001, Choon-Han Youn |
eScience | 7 |
| 2005 | An Infrastructure of Network Services for Seamless Integration in Advanced Collaborative Computing EnvironmentsabstractAdvanced collaborative computing environments are one of the most important tools for integrating high-performance computers and computations and for interacting with colleagues around the world. However, heterogeneous characteristics such as network transfer rates, computational abilities, and hierarchical systems make the seamless integration of distributed resources a challenge. In this paper, we argue that advanced collaborative computing environments need an infrastructure of network services to support distributed and quality guaranteed multimedia applications. Accordingly, we propose the design of network services for high-performance collaborative computing. We present a collaborative environment network service infrastructure (CENSI) to embed network services into various systems intelligently and elastically. We also discuss three management modules: a three-party matching module (resources, requests, and network services), a module for performance monitoring and evaluation of group communications, and a module for distribution topology analysis Ivan R. Judson, Thomas D. Uram, S. Lefvert, Terry Disz, Michael E. Papka, Rick L. Stevens |
CLUSTER | 3 |
| 2004 | Capability matching of data streams with network servicesabstractDistributed computing middleware needs to support a wide range of resources, such as diverse software components, various hardware devices, and heterogeneous operating systems and architectures. Current technologies are unable to implement a maintenance-free platform to be compatible with such different computing environments. This situation is presenting an increasing challenge as Grid computing becomes more widespread. The infrastructure of network services (CENSA and CENSI) has been proposed to address this challenge. A seamless Grid computing environment, supported by network services, is composed of various streams such as data, video, audio, and text. We define a mathematical model of capability matching for three-party agreements: requests from users, resources, and network services. Based on the mathematical model, we provide a general approach for capability matching. We also present a new language schema for capability description. As an example, we embed the general matchmaker in the architecture of the access Grid. Several tests of accuracy and performance are discussed. Ivan R. Judson, Thomas D. Uram, Terry Disz, Michael E. Papka, Rick L. Stevens |
CCGRID | 3 |