Joel H. Saltz

dblp:s/JoelHSaltz · DBLP profile ↗
← Back
238ranked-venue papers
10as first author
31since 2021 · last 2026
0000-0002-3451-2165ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 114 · 8 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 83 · 2 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 14 since 2021Artificial intelligence and machine learning · 24 · 9 since 2021Databases, data management, data science and information retrieval · 14 · 2 since 2021Software engineering, systems software and programming languages · 6Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Measuring and predicting where and when pathologists focus their visual attention while grading whole slide images of cancer
Souradeep Chakraborty, Ruoyu Xue, Rajarsi Gupta 0001, Oksana Yaskiv, Constantin Friedman, Natallia Sheuka, Dana Perez, Paul Friedman, Won-Tak Choi, Waqas Mahmud, Beatrice S. Knudsen, Gregory J. Zelinsky, Joel H. Saltz, Dimitris Samaras
Medical Image Anal.13
2026 Position Paper: Artificial Intelligence in Medical Image Analysis: Advances, Clinical Translation, and Emerging Frontiers
abstract
Over the past five years, artificial intelligence (AI) has introduced new models and methods for addressing the challenges associated with the broader adoption of AI models and systems in medicine. This paper reviews recent advances in AI for medical image and video analysis, outlines emerging paradigms, highlights pathways for successful clinical translation, and provides recommendations for future work. Hybrid Convolutional Neural Network (CNN) Transformer architectures now deliver state-of-the-art results in segmentation, classification, reconstruction, synthesis, and registration. Foundation and generative AI models enable the use of transfer learning to smaller datasets with limited ground truth. Federated learning supports privacy-preserving collaboration across institutions. Explainable and trustworthy AI approaches have become essential to foster clinician trust, ensure regulatory compliance, and facilitate ethical deployment. Together, these developments pave the way for integrating AI into radiology, pathology, and wider healthcare workflows.
Andreas Panayides, Hao Chen 0011, Nenad Filipovic, Tijana Geroski, Junlin Hou, Karim Lekadir, Kostas Marias, George K. Matsopoulos, Giorgos Papanastasiou, Pinaki Sarder, Georgia D. Tourassi, Sotirios A. Tsaftaris, Huazhu Fu, Efthyvoulos C. Kyriacou, Christos P. Loizou, Michalis E. Zervakis, Joel H. Saltz, Farah Shamout, Ken C. L. Wong, Jianhua Yao 0001, Amir A. Amini, Dimitrios I. Fotiadis, Constantinos S. Pattichis, Marios S. Pattichis
IEEE J. Biomed. Health Informatics17
2026 Label-Efficient Deep Color Deconvolution of Brightfield Multiplex IHC Images
abstract
Brightfield Multiplex Immunohistochemistry (mIHC) provides simultaneous labeling of multiple protein biomarkers in the same tissue section. It enables the exploration of spatial relationships between the inflammatory microenvironment and tumor cells, and to uncover how tumor cell morphology relates to cancer biomarker expression. Color deconvolution is required to analyze and quantify the different cell phenotype populations present as indicated by the biomarkers. However, this becomes a challenging task as the number of multiplexed stains increase. In this work, we present self-supervised and semi-supervised approaches to mIHC color deconvolution. Our proposed methods are based on deep convolutional autoencoders and learn using innovative reconstruction losses inspired by physics. We show how we can integrate weak annotations and the abundant unlabeled data available to train a model to reliably unmix the multiplexed stains and generate stain segmentation maps. We demonstrate the effectiveness of our proposed methods through experiments on mIHC dataset of 7-plexed IHC images.
Shahira Abousamra, Danielle Fassler, Rajarsi Gupta 0001, Tahsin M. Kurç, Luisa F. Escobar-Hoyos, Dimitris Samaras, Kenneth Shroyer, Joel H. Saltz, Chao Chen 0012
IEEE Trans. Medical Imaging8
2025 ZoomLDM: Latent Diffusion Model for Multi-scale Image Generation
abstract
Diffusion models have revolutionized image generation, yet several challenges restrict their application to large-image domains, such as digital pathology and satellite imagery. Given that it is infeasible to directly train a model on ’whole’ images from domains with potential gigapixel sizes, diffusion-based generative methods have focused on synthesizing small, fixed-size patches extracted from these images. However, generating small patches has limited applicability since patch-based models fail to capture the global structures and wider context of large images, which can be crucial for synthesizing (semantically) accurate samples. To overcome this limitation, we present ZoomLDM, a diffusion model tailored for generating images across multiple scales. Central to our approach is a novel magnification-aware conditioning mechanism that utilizes self-supervised learning (SSL) embeddings and allows the diffusion model to synthesize images at different ’zoom’ levels, i.e., fixed-size patches extracted from large images at varying scales. ZoomLDM synthesizes coherent histopathology images that remain contextually accurate and detailed at different zoom levels, achieving state-of-the-art image generation quality across all scales and excelling in the data-scarce setting of generating thumbnails of entire large images. The multi-scale nature of ZoomLDM unlocks additional capabilities in large image generation, enabling computationally tractable and globally coherent image synthesis up to 4096 × 4096 pixels and 4 × super-resolution. Additionally, multi-scale features extracted from ZoomLDM are highly effective in multiple instance learning experiments.1
Srikar Yellapragada, Alexandros Graikos, Kostas Triaridis, Prateek Prasanna, Rajarsi Gupta 0001, Joel H. Saltz, Dimitris Samaras
CVPR6
2025 GECKO: Gigapixel Vision-Concept Contrastive Pretraining in Histopathology
abstract
Pretraining a Multiple Instance Learning (MIL) aggregator enables the derivation of Whole Slide Image (WSI)-level embeddings from patch-level representations without supervision. While recent multimodal MIL pretraining approaches leveraging auxiliary modalities have demonstrated performance gains over unimodal WSI pretraining, the acquisition of these additional modalities necessitates extensive clinical profiling. This requirement increases costs and limits scalability in existing WSI datasets lacking such paired modalities. To address this, we propose Gigapixel Vision-Concept Knowledge Contrastive pretraining (GECKO), which aligns WSIs with a Concept Prior derived from the available WSIs. First, we derive an inherently interpretable concept prior by computing the similarity between each WSI patch and textual descriptions of predefined pathology concepts. GECKO then employs a dual-branch MIL network: one branch aggregates patch embeddings into a WSI-level deep embedding, while the other aggregates the concept prior into a corresponding WSI-level concept embedding. Both aggregated embeddings are aligned using a contrastive objective, thereby pretraining the entire dual-branch MIL model. Moreover, when auxiliary modalities such as transcriptomics data are available, GECKO seamlessly integrates them. Across five diverse tasks, GECKO consistently outperforms prior unimodal and multimodal pretraining approaches while also delivering clinically meaningful interpretability that bridges the gap between computational models and pathology expertise. Code is made available at https://github.com/bmi-imaginelab/GECKO
Saarthak Kapse, Pushpak Pati, Srikar Yellapragada, Srijan Das, Rajarsi Gupta 0001, Joel H. Saltz, Dimitris Samaras, Prateek Prasanna
ICCV6
2025 Pathology Image Compression with Pre-trained Autoencoders
Srikar Yellapragada, Alexandros Graikos, Kostas Triaridis, Zilinghan Li, Tarak Nath Nandi, Ravi K. Madduri, Prateek Prasanna, Joel H. Saltz, Dimitris Samaras
MICCAI (2)8
2025 National COVID Cohort Collaborative data enhancements: a path for expanding common data models
abstract
OBJECTIVE: To support long COVID research in National COVID Cohort Collaborative (N3C), the N3C Phenotype and Data Acquisition team created data designs to aid contributing sites in enhancing their data. Enhancements include long COVID specialty clinic indicator; Admission, Discharge, and Transfer transactions; patient-level social determinants of health; and in-hospital use of oxygen supplementation. MATERIALS AND METHODS: For each enhancement, we defined the scope and wrote guidance on how to prepare and populate the data in a standardized way. RESULTS: As of June 2024, 29 sites have added at least one data enhancement to their N3C pipeline. DISCUSSION: The use of common data models is critical to the success of N3C; however, these data models cannot account for all needs. Project-driven data enhancement is required. This should be done in a standardized way in alignment with common data model specifications. Our approach offers a useful pathway for enhancing data to improve fit for purpose. CONCLUSION: In this initiative, we rapidly produced project-specific data modeling guidance and documentation in support of long COVID research while maintaining a commitment to terminology standards and harmonized data.
Kellie M. Walters, Marshall Clark, Sofia Dard, Stephanie S. Hong, Elizabeth Kelly, Kristin Kostka, Adam M. Lee, Robert T. Miller, Michele Morris, Matvey Palchuk, Emily R. Pfaff, Adam B. Wilcox, Alexis Graves, Alfred Anzalone, Amin Manna, Amit Saha, Amy Olex, Andrea Zhou, Andrew E. Williams, Andrew Southerland, Andrew T. Girvin, Anita Walden, Anjali A Sharathkumar, Benjamin R. C. Amor, Benjamin Bates, Brian Hendricks, Caleb Alexander, Carolyn T. Bramante, Cavin Ward-Caviness, Charisse R. Madlock-Brown, Christine Suver, Christopher G. Chute, Christopher Dillon, Chunlei Wu, Clare Schmitt, Cliff Takemoto, Dan Housman, Davera Gabriel, David Eichmann, Diego Mazzotti, Don Brown, Eilis A. Boudreau, Elaine L. Hill, Elizabeth Zampino, Emily Carlson Marti, Evan French, Farrukh M. Koraishy, Federico Mariona, Fred W. Prior, George Sokos, Greg Martin, Harold P. Lehmann, Heidi Spratt, Hemalkumar Mehta, Hythem Sidky, J. W. Awori Hayanga, Jami Pincavitch, Jaylyn Clark, Jeremy Richard Harper, Jessica Islam, Jin Ge, Joel Gagnier, Joel H. Saltz, Johanna Loomba, John Buse, Jomol P. Mathew, Joni L. Rutter, Julie A. McMurry, Justin Guinney, Justin Starren, Karen Crowley, Katie Rebecca Bradwell, Ken Wilkins, Kenneth R. Gersing, Kenrick Dwain Cato, Kimberly Murray, Lavance Northington, Lee Allan Pyles, Leonie Misquitta, Lesley Cottrell, Lili M. Portilla, Mariam Deacy, Mark M. Bissell, Mary Emmett, Mary Morrison Saltz, Melissa A. Haendel, Meredith C. B. Adams, Meredith Temple-O'Connor, Michael G. Kurilla, Nabeel Qureshi, Nasia Safdar, Nicole Garbarini, Noha Sharafeldin, Ofer Sadan, Patricia A. Francis, Penny Wung Burgoon, Peter N. Robinson, Philip R. O. Payne, Rafael Fuentes, Randeep Jawa, Rebecca Erwin-Cohen, Rena Patel, Richard A. Moffitt, Richard L. Zhu, Rishi Kamaleswaran, Robert Hurley, Saiju Pyarajan, Samuel G. Michael, Samuel Bozzette, Sandeep Mallipattu, Satyanarayana Vedula, Scott Chapman, Shawn T. O'Neil, Soko Setoguchi, Tellen D. Bennett, Tiffany Callahan, Umit Topaloglu, Usman Sheikh, Valery Gordon, Vignesh Subbian, Warren A. Kibbe, Wenndy Hernandez, Will Beasley, Will Cooper, William Hillegass, Xiaohan Tanner Zhang
J. Am. Medical Informatics Assoc.65
2024 Pan-Cancer Tumor Infiltrating Lymphocyte Detection based on Federated Learning
abstract
Advances in deep learning (DL) have shown great promise in revolutionizing healthcare, notwithstanding their success hinging on the availability of centralized large and diverse data. Such centralization is challenging because of numerous concerns relating to privacy, data-ownership, intellectual property, and compliance with varying regulatory policies. Federated learning (FL), offers a new decentralized paradigm to train DL models in healthcare. In this study, we evaluate the effect of FL in developing DL models for the analysis of digitized tissue sections, specifically whole slide images (WSIs). A classification application was considered as the example use case, to quantify the distribution of Tumor Infiltrating Lymphocytes (TILs), which are a critical biomarker in cancer research, providing valuable insights into patient outcomes. We trained a VGG classification model using 50 × 50 micron patches extracted from the WSIs with their associated TIL/nonTIL label. We simulated a FL environment, where different cancer types are included across each collaborating node. Our results show that the model trained with the federated training approach achieves similar performance, both quantitatively and qualitatively, to that of a model trained with all the training data pooled at a centralized location. Our study shows that FL has tremendous potential for enabling the development of more robust and accurate models for histopathology image analysis without having to collect large and diverse training data at a single location. Particularly for TILs, our FL approach yields a single DL model trained across numerous anatomical sites and able to robustly generalize to unseen cancer types.
Ujjwal Baid, Sarthak Pati, Tahsin M. Kurç, Rajarsi Gupta 0001, Erich Bremer, Shahira Abousamra, Siddhesh P. Thakur, Joel H. Saltz, Spyridon Bakas
IEEE Big Data8
2024 Learned Representation-Guided Diffusion Models for Large-Image Generation
abstract
To synthesize high-fidelity samples, diffusion models typically require auxiliary data to guide the generation process. However, it is impractical to procure the painstaking patch-level annotation effort required in specialized domains like histopathology and satellite imagery; it is often performed by domain experts and involves hundreds of millions of patches. Modern-day self-supervised learning (SSL) representations encode rich semantic and visual information. In this paper, we posit that such representations are expressive enough to act as proxies to fine-grained human labels. We introduce a novel approach that trains diffusion models conditioned on embeddings from SSL. Our diffusion models successfully project these features back to high-quality histopathology and remote sensing images. In addition, we construct larger images by assembling spatially consistent patches inferred from SSL embeddings, preserving long-range dependencies. Augmenting real data by generating variations of real images improves downstream classifier accuracy for patch-level and larger, image-scale classification tasks. Our models are effective even on datasets not encountered during training, demonstrating their robustness and generalizability. Generating images from learned embeddings is agnostic to the source of the embeddings. The SSL embeddings used to generate a large image can either be extracted from a reference image, or sampled from an auxiliary model conditioned on any related modality (e.g. class labels, text, genomic data). As proof of concept, we introduce the text-to-large image synthesis paradigm where we successfully synthesize large pathology and satellite images out of text descriptions.
Alexandros Graikos, Srikar Yellapragada, Minh-Quan Le, Saarthak Kapse, Prateek Prasanna, Joel H. Saltz, Dimitris Samaras
CVPR6
2024 SI-MIL: Taming Deep MIL for Self-Interpretability in Gigapixel Histopathology
abstract
Introducing interpretability and reasoning into Multiple Instance Learning (MIL) methods for Whole Slide Image (WSI) analysis is challenging, given the complexity of gigapixel slides. Traditionally, MIL interpretability is limited to identifying salient regions deemed pertinent for downstream tasks, offering little insight to the end-user (pathologist) regarding the rationale behind these selections. To address this, we propose Self-Interpretable MIL (SI-MIL), a method intrinsically designed for interpretability from the very outset. SI-MIL employs a deep MIL framework to guide an interpretable branch grounded on handcrafted pathological features, facilitating linear predictions. Beyond identifying salient regions, SI-MIL uniquely provides feature-level interpretations rooted in pathological insights for WSIs. Notably, SI-MIL, with its linear prediction constraints, challenges the prevalent myth of an inevitable trade-off between model interpretability and performance, demonstrating competitive results compared to state-of-the-art methods on WSI-level prediction tasks across three cancer types. In addition, we thoroughly benchmark the local-and global-interpretability of SI-MIL in terms of statistical analysis, a domain expert study, and desiderata of interpretability, namely, user-friendliness and faithfulness.
Saarthak Kapse, Pushpak Pati, Srijan Das, Chao Chen 0012, Maria Vakalopoulou, Joel H. Saltz, Dimitris Samaras, Rajarsi Gupta 0001, Prateek Prasanna
CVPR7
2024 ∞-Brush: Controllable Large Image Synthesis with Diffusion Models in Infinite Dimensions
Minh-Quan Le, Alexandros Graikos, Srikar Yellapragada, Rajarsi Gupta 0001, Joel H. Saltz, Dimitris Samaras
ECCV (32)5
2024 Decoding the Visual Attention of Pathologists to Reveal Their Level of Expertise
Souradeep Chakraborty, Rajarsi Gupta 0001, Oksana Yaskiv, Constantin Friedman, Natallia Sheuka, Dana Perez, Paul Friedman, Gregory J. Zelinsky, Joel H. Saltz, Dimitris Samaras
MICCAI (3)9
2024 Semi-supervised Contrastive VAE for Disentanglement of Digital Pathology Images
Mahmudul Hasan 0006, Xiaoling Hu 0002, Shahira Abousamra, Prateek Prasanna, Joel H. Saltz, Chao Chen 0012
MICCAI (4)5
2024 PathLDM: Text conditioned Latent Diffusion Model for Histopathology
abstract
To achieve high-quality results, diffusion models must be trained on large datasets. This can be notably prohibitive for models in specialized domains, such as computational pathology. Conditioning on labeled data is known to help in data-efficient model training. Therefore, histopathology reports, which are rich in valuable clinical information, are an ideal choice as guidance for a histopathology generative model. In this paper, we introduce PathLDM, the first text-conditioned Latent Diffusion Model tailored for generating high-quality histopathology images. Leveraging the rich contextual information provided by pathology text reports, our approach fuses image and textual data to enhance the generation process. By utilizing GPT's capabilities to distill and summarize complex text reports, we establish an effective conditioning mechanism. Through strategic conditioning and necessary architectural enhancements, we achieved a SoTA FID score of 7.64 for text-to-image generation on the TCGA-BRCA dataset, significantly outperforming the closest text-conditioned competitor with FID 30.1.
Srikar Yellapragada, Alexandros Graikos, Prateek Prasanna, Tahsin M. Kurç, Joel H. Saltz, Dimitris Samaras
WACV5
2024 Attention De-sparsification Matters: Inducing diversity in digital pathology representation learning
Saarthak Kapse, Srijan Das, Rajarsi Gupta 0001, Joel H. Saltz, Dimitris Samaras, Prateek Prasanna
Medical Image Anal.5
2024 High-Performance Spatial Data Analytics: Systematic R&D for Scale-Out and Scale-Up Solutions from the Past to Now
abstract
We released open-source software Hadoop-GIS in 2011, and presented and published the work in VLDB 2013. This work initiated the development of a new spatial data analytical ecosystem characterized by its large-scale capacity in both computing and data storage, high scalability, compatibility with low-cost commodity processors in clusters and open-source software. After more than a decade of research and development, this ecosystem has matured and is now serving many applications across various fields. In this paper, we provide the background on why we started this project and give an overview of the original Hadoop-GIS software architecture, along with its unique technical contributions and legacy. We present the evolution of the ecosystem and its current state-of-the-art, which has been influenced by the Hadoop-GIS project. We also describe the ongoing efforts to further enhance this ecosystem with hardware accelerations to meet the increasing demands for low latency and high throughput in various spatial data analysis tasks. Finally, we will summarize the insights gained and lessons learned over more than a decade in pursuing high-performance spatial data analytics.
Fusheng Wang 0001, Rubao Lee, Dejun Teng, Xiaodong Zhang 0001, Joel H. Saltz
Proc. VLDB Endow.5
2023 Topology-Guided Multi-Class Cell Context Generation for Digital Pathology
abstract
In digital pathology, the spatial context of cells is important for cell classification, cancer diagnosis and prognosis. To model such complex cell context, however, is challenging. Cells form different mixtures, lineages, clusters and holes. To model such structural patterns in a learnable fashion, we introduce several mathematical tools from spatial statistics and topological data analysis. We incorporate such structural descriptors into a deep generative model as both conditional inputs and a differentiable loss. This way, we are able to generate high quality multi-class cell layouts for the first time. We show that the topology-rich cell layouts can be used for data augmentation and improve the performance of downstream tasks such as cell classification.
Shahira Abousamra, Rajarsi Gupta 0001, Tahsin M. Kurç, Dimitris Samaras, Joel H. Saltz, Chao Chen 0012
CVPR5
2023 Prompt-MIL: Boosting Multi-instance Learning Schemes via Task-Specific Prompt Tuning
Saarthak Kapse, Ke Ma 0005, Prateek Prasanna, Joel H. Saltz, Maria Vakalopoulou, Dimitris Samaras
MICCAI (8)5
2023 Effective and efficient active learning for deep learning-based tissue image analysis
abstract
MOTIVATION: Deep learning attained excellent results in digital pathology recently. A challenge with its use is that high quality, representative training datasets are required to build robust models. Data annotation in the domain is labor intensive and demands substantial time commitment from expert pathologists. Active learning (AL) is a strategy to minimize annotation. The goal is to select samples from the pool of unlabeled data for annotation that improves model accuracy. However, AL is a very compute demanding approach. The benefits for model learning may vary according to the strategy used, and it may be hard for a domain specialist to fine tune the solution without an integrated interface. RESULTS: We developed a framework that includes a friendly user interface along with run-time optimizations to reduce annotation and execution time in AL in digital pathology. Our solution implements several AL strategies along with our diversity-aware data acquisition (DADA) acquisition function, which enforces data diversity to improve the prediction performance of a model. In this work, we employed a model simplification strategy [Network Auto-Reduction (NAR)] that significantly improves AL execution time when coupled with DADA. NAR produces less compute demanding models, which replace the target models during the AL process to reduce processing demands. An evaluation with a tumor-infiltrating lymphocytes classification application shows that: (i) DADA attains superior performance compared to state-of-the-art AL strategies for different convolutional neural networks (CNNs), (ii) NAR improves the AL execution time by up to 4.3×, and (iii) target models trained with patches/data selected by the NAR reduced versions achieve similar or superior classification quality to using target CNNs for data selection. AVAILABILITY AND IMPLEMENTATION: Source code: https://github.com/alsmeirelles/DADA.
André L. S. Meirelles, Tahsin M. Kurç, Jun Kong 0002, Renato Ferreira 0001, Joel H. Saltz, George Teodoro
Bioinform.5
2023 An open natural language processing (NLP) framework for EHR-based clinical research: a case demonstration using the National COVID Cohort Collaborative (N3C)
abstract
Despite recent methodology advancements in clinical natural language processing (NLP), the adoption of clinical NLP models within the translational research community remains hindered by process heterogeneity and human factor variations. Concurrently, these factors also dramatically increase the difficulty in developing NLP models in multi-site settings, which is necessary for algorithm robustness and generalizability. Here, we reported on our experience developing an NLP solution for Coronavirus Disease 2019 (COVID-19) signs and symptom extraction in an open NLP framework from a subset of sites participating in the National COVID Cohort (N3C). We then empirically highlight the benefits of multi-site data for both symbolic and statistical methods, as well as highlight the need for federated annotation and evaluation to resolve several pitfalls encountered in the course of these efforts.
Sijia Liu 0002, Andrew Wen, Liwei Wang 0010, Sunyang Fu, Robert T. Miller, Andrew E. Williams, Daniel R. Harris, Ramakanth Kavuluru, Noor Abu-El-Rub, Dalton Schutte, Rui Zhang 0028, Masoud Rouhizadeh, John D. Osborne, Yongqun He, Umit Topaloglu, Stephanie S. Hong, Joel H. Saltz, Thomas Schaffter, Emily R. Pfaff, Christopher G. Chute, Tim Duong, Melissa A. Haendel, Rafael Fuentes, Peter Szolovits, Hua Xu 0001
J. Am. Medical Informatics Assoc.19
2022 Learning Topological Interactions for Multi-Class Medical Image Segmentation
Saumya Gupta, Xiaoling Hu 0002, James Kaan, Michael Jin, Mutshipay Mpoy, Katherine Chung, Mary M. Saltz, Tahsin M. Kurç, Joel H. Saltz, Apostolos Tassiopoulos, Prateek Prasanna, Chao Chen 0012
ECCV (29)10
2022 Gigapixel Whole-Slide Images Classification Using Locally Supervised Learning
Ke Ma 0005, Rajarsi Gupta 0001, Joel H. Saltz, Maria Vakalopoulou, Dimitris Samaras
MICCAI (2)5
2022 Harmonizing units and values of quantitative data elements in a very large nationally pooled electronic health record (EHR) dataset
abstract
OBJECTIVE: The goals of this study were to harmonize data from electronic health records (EHRs) into common units, and impute units that were missing. MATERIALS AND METHODS: The National COVID Cohort Collaborative (N3C) table of laboratory measurement data-over 3.1 billion patient records and over 19 000 unique measurement concepts in the Observational Medical Outcomes Partnership (OMOP) common-data-model format from 55 data partners. We grouped ontologically similar OMOP concepts together for 52 variables relevant to COVID-19 research, and developed a unit-harmonization pipeline comprised of (1) selecting a canonical unit for each measurement variable, (2) arriving at a formula for conversion, (3) obtaining clinical review of each formula, (4) applying the formula to convert data values in each unit into the target canonical unit, and (5) removing any harmonized value that fell outside of accepted value ranges for the variable. For data with missing units for all the results within a lab test for a data partner, we compared values with pooled values of all data partners, using the Kolmogorov-Smirnov test. RESULTS: Of the concepts without missing values, we harmonized 88.1% of the values, and imputed units for 78.2% of records where units were absent (41% of contributors' records lacked units). DISCUSSION: The harmonization and inference methods developed herein can serve as a resource for initiatives aiming to extract insight from heterogeneous EHR collections. Unique properties of centralized data are harnessed to enable unit inference. CONCLUSION: The pipeline we developed for the pooled N3C data enables use of measurements that would otherwise be unavailable for analysis.
Katie R. Bradwell, Jacob T. Wooldridge, Benjamin R. C. Amor, Tellen D. Bennett, Adit Anand, Carolyn Bremer, Yun Jae Yoo, Zhenglong Qian, Steven G. Johnson, Emily R. Pfaff, Andrew T. Girvin, Amin Manna, Emily Niehaus, Stephanie S. Hong, Xiaohan Tanner Zhang, Richard L. Zhu, Mark Bissell, Nabeel Qureshi, Joel H. Saltz, Melissa A. Haendel, Christopher G. Chute, Harold P. Lehmann, Richard A. Moffitt
J. Am. Medical Informatics Assoc.19
2022 Demonstrating an approach for evaluating synthetic geospatial and temporal epidemiologic data utility: results from analyzing >1.8 million SARS-CoV-2 tests in the United States National COVID Cohort Collaborative (N3C)
abstract
OBJECTIVE: This study sought to evaluate whether synthetic data derived from a national coronavirus disease 2019 (COVID-19) dataset could be used for geospatial and temporal epidemic analyses. MATERIALS AND METHODS: Using an original dataset (n = 1 854 968 severe acute respiratory syndrome coronavirus 2 tests) and its synthetic derivative, we compared key indicators of COVID-19 community spread through analysis of aggregate and zip code-level epidemic curves, patient characteristics and outcomes, distribution of tests by zip code, and indicator counts stratified by month and zip code. Similarity between the data was statistically and qualitatively evaluated. RESULTS: In general, synthetic data closely matched original data for epidemic curves, patient characteristics, and outcomes. Synthetic data suppressed labels of zip codes with few total tests (mean = 2.9 ± 2.4; max = 16 tests; 66% reduction of unique zip codes). Epidemic curves and monthly indicator counts were similar between synthetic and original data in a random sample of the most tested (top 1%; n = 171) and for all unsuppressed zip codes (n = 5819), respectively. In small sample sizes, synthetic data utility was notably decreased. DISCUSSION: Analyses on the population-level and of densely tested zip codes (which contained most of the data) were similar between original and synthetically derived datasets. Analyses of sparsely tested populations were less similar and had more data suppression. CONCLUSION: In general, synthetic data were successfully used to analyze geospatial and temporal trends. Analyses using small sample sizes or populations were limited, in part due to purposeful data label suppression-an attribute disclosure countermeasure. Users should consider data fitness for use in these cases.
Jason A. Thomas, Randi E. Foraker, Noa Zamstein, Jon D. Morrow, Philip R. O. Payne, Adam B. Wilcox, Melissa A. Haendel, Christopher G. Chute, Kenneth R. Gersing, Anita Walden, Tellen D. Bennett, David Eichmann, Justin Guinney, Warren A. Kibbe, Emily R. Pfaff, Peter N. Robinson, Joel H. Saltz, Heidi Spratt, Justin Starren, Christine Suver, Chunlei Wu, Davera Gabriel, Stephanie S. Hong, Kristin Kostka, Harold P. Lehmann, Richard A. Moffitt, Michele Morris, Matvey Palchuk, Xiaohan Tanner Zhang, Richard L. Zhu, Benjamin R. C. Amor, Mark M. Bissell, Marshall Clark, Andrew T. Girvin, Adam M. Lee, Robert T. Miller, Kellie M. Walters, Yooree Chae, Connor Cook, Alexandra Dest, Racquel R. Dietz, Thomas Dillon, Patricia A. Francis, Rafael Fuentes, Alexis Graves, Andrew J. Neumann, Shawn T. O'Neil, Usman Sheikh, Andréa M. Volz, Elizabeth Zampino, Christopher P. Austin, Samuel Bozzette, Mariam Deacy, Nicole Garbarini, Michael G. Kurilla, Samuel G. Michael, Joni L. Rutter, Meredith Temple-O'Connor, Katie Rebecca Bradwell, Amin Manna, Nabeel Qureshi, Mary Morrison Saltz, Julie A. McMurry, Carolyn T. Bramante, Jeremy Richard Harper, Wenndy Hernandez, Farrukh M. Koraishy, Federico Mariona, Saidulu Mattapally, Amit Saha, Satyanarayana Vedula, Yujuan Fu, Nisha Mathews, Ofer Mendelevitch
J. Am. Medical Informatics Assoc.18
2022 Efficient microscopy image analysis on CPU-GPU systems with cost-aware irregular data partitioning
Willian de Oliveira Barreiros Junior, Alba Cristina Magalhaes Alves de Melo, Jun Kong 0002, Renato Ferreira 0001, Tahsin M. Kurç, Joel H. Saltz, George Teodoro
J. Parallel Distributed Comput.6
2021 Informatics to Power Post-COVID Care: A Framework for Patient Care and Secondary Data Use
Sritha Rajupet, Rachel Wong, Donna Moller, Lisa Maldonado, Tricia Weiss, Tahsin M. Kurç, Janos G. Hajagos, Hasit Shah, Mary M. Saltz, Joel H. Saltz, Veena Lingam
AMIA10
2021 Generating Longitudinal Synthetic EHR Data with Recurrent Autoencoders and Generative Adversarial Networks
Siao Sun, Fusheng Wang 0001, Sina Rashidian, Tahsin M. Kurç, Kayley Abell-Hart, Janos G. Hajagos, Wei Zhu 0008, Mary M. Saltz, Joel H. Saltz
AMIA9
2021 Multi-Class Cell Detection Using Spatial Context Representation
abstract
In digital pathology, both detection and classification of cells are important for automatic diagnostic and prognostic tasks. Classifying cells into subtypes, such as tumor cells, lymphocytes or stromal cells is particularly challenging. Existing methods focus on morphological appearance of individual cells, whereas in practice pathologists often infer cell classes through their spatial context. In this paper, we propose a novel method for both detection and classification that explicitly incorporates spatial contextual information. We use the spatial statistical function to describe local density in both a multi-class and a multi-scale manner. Through representation learning and deep clustering techniques, we learn advanced cell representation with both appearance and spatial context. On various benchmarks, our method achieves better performance than state-of-the-arts, especially on the classification task. We also create a new dataset for multi-class cell detection and classification in breast cancer and we make both our code and data publicly available.
Shahira Abousamra, David Belinsky, John S. Van Arnam, Felicia Allard, Eric Yee, Rajarsi Gupta 0001, Tahsin M. Kurç, Dimitris Samaras, Joel H. Saltz, Chao Chen 0012
ICCV9
2021 Identifying risk of opioid use disorder for patients taking opioid medications with deep learning
abstract
OBJECTIVE: The United States is experiencing an opioid epidemic. In recent years, there were more than 10 million opioid misusers aged 12 years or older annually. Identifying patients at high risk of opioid use disorder (OUD) can help to make early clinical interventions to reduce the risk of OUD. Our goal is to develop and evaluate models to predict OUD for patients on opioid medications using electronic health records and deep learning methods. The resulting models help us to better understand OUD, providing new insights on the opioid epidemic. Further, these models provide a foundation for clinical tools to predict OUD before it occurs, permitting early interventions. METHODS: Electronic health records of patients who have been prescribed with medications containing active opioid ingredients were extracted from Cerner's Health Facts database for encounters between January 1, 2008, and December 31, 2017. Long short-term memory models were applied to predict OUD risk based on five recent prior encounters before the target encounter and compared with logistic regression, random forest, decision tree, and dense neural network. Prediction performance was assessed using F1 score, precision, recall, and area under the receiver-operating characteristic curve. RESULTS: The long short-term memory (LSTM) model provided promising prediction results which outperformed other methods, with an F1 score of 0.8023 (about 0.016 higher than dense neural network (DNN)) and an area under the receiver-operating characteristic curve (AUROC) of 0.9369 (about 0.145 higher than DNN). CONCLUSIONS: LSTM-based sequential deep learning models can accurately predict OUD using a patient's history of electronic health records, with minimal prior domain knowledge. This tool has the potential to improve clinical decision support for early intervention and prevention to combat the opioid epidemic.
Jianyuan Deng, Sina Rashidian, Kayley Abell-Hart, Richard N. Rosenthal, Mary M. Saltz, Joel H. Saltz, Fusheng Wang 0001
J. Am. Medical Informatics Assoc.8
2021 The National COVID Cohort Collaborative (N3C): Rationale, design, infrastructure, and deployment
abstract
OBJECTIVE: Coronavirus disease 2019 (COVID-19) poses societal challenges that require expeditious data and knowledge sharing. Though organizational clinical data are abundant, these are largely inaccessible to outside researchers. Statistical, machine learning, and causal analyses are most successful with large-scale data beyond what is available in any given organization. Here, we introduce the National COVID Cohort Collaborative (N3C), an open science community focused on analyzing patient-level data from many centers. MATERIALS AND METHODS: The Clinical and Translational Science Award Program and scientific community created N3C to overcome technical, regulatory, policy, and governance barriers to sharing and harmonizing individual-level clinical data. We developed solutions to extract, aggregate, and harmonize data across organizations and data models, and created a secure data enclave to enable efficient, transparent, and reproducible collaborative analytics. RESULTS: Organized in inclusive workstreams, we created legal agreements and governance for organizations and researchers; data extraction scripts to identify and ingest positive, negative, and possible COVID-19 cases; a data quality assurance and harmonization pipeline to create a single harmonized dataset; population of the secure data enclave with data, machine learning, and statistical analytics tools; dissemination mechanisms; and a synthetic data pilot to democratize data access. CONCLUSIONS: The N3C has demonstrated that a multisite collaborative learning health network can overcome barriers to rapidly build a scalable infrastructure incorporating multiorganizational clinical data for COVID-19 analytics. We expect this effort to save lives by enabling rapid collaboration among clinicians, researchers, and data scientists to identify treatments and specialized care and thereby reduce the immediate and long-term impacts of COVID-19.
Melissa A. Haendel, Christopher G. Chute, Tellen D. Bennett, David Eichmann, Justin Guinney, Warren A. Kibbe, Philip R. O. Payne, Emily R. Pfaff, Peter N. Robinson, Joel H. Saltz, Heidi Spratt, Christine Suver, John Wilbanks, Adam B. Wilcox, Andrew E. Williams, Chunlei Wu, Clair Blacketer, Robert L. Bradford, James J. Cimino, Marshall Clark, Evan W. Colmenares, Patricia A. Francis, Davera Gabriel, Alexis Graves, Raju Hemadri, Stephanie S. Hong, George Hripcsak, Dazhi Jiao, Jeffrey G. Klann, Kristin Kostka, Adam M. Lee, Harold P. Lehmann, Lora Lingrey, Robert T. Miller, Michele Morris, Shawn N. Murphy, Karthik Natarajan, Matvey Palchuk, Usman Sheikh, Harold R. Solbrig, Shyam Visweswaran, Anita Walden, Kellie M. Walters, Griffin M. Weber, Xiaohan Tanner Zhang, Richard L. Zhu, Benjamin R. C. Amor, Andrew T. Girvin, Amin Manna, Nabeel Qureshi, Michael G. Kurilla, Samuel G. Michael, Lili M. Portilla, Joni L. Rutter, Christopher P. Austin, Kenneth R. Gersing
J. Am. Medical Informatics Assoc.10
2021 Predicting opioid overdose risk of patients with opioid prescriptions using electronic health records based on temporal deep learning
Jianyuan Deng, Sina Rashidian, Richard N. Rosenthal, Mary M. Saltz, Joel H. Saltz, Fusheng Wang 0001
J. Biomed. Informatics7
2020 SMOOTH-GAN: Towards Sharp and Smooth Synthetic EHR Data Generation
Sina Rashidian, Fusheng Wang 0001, Richard A. Moffitt, Anurag Dutt, Vishwam Pandya, Janos G. Hajagos, Mary M. Saltz, Joel H. Saltz
AIME10
2020 Optimizing parameter sensitivity analysis of large-scale microscopy image analysis workflows with multilevel computation reuse
abstract
Parameter sensitivity analysis (SA) is an effective tool to gain knowledge about complex analysis applications and assess the variability in their analysis results. However, it is an expensive process as it requires the execution of the target application multiple times with a large number of different input parameter values. In this work, we propose optimizations to reduce the overall computation cost of SA in the context of analysis applications that segment high-resolution slide tissue images, ie, images with resolutions of 100k × 100k pixels. Two cost-cutting techniques are combined to efficiently execute SA: use of distributed hybrid systems for parallel execution and computation reuse at multiple levels of an analysis pipeline to reduce the amount of computation. These techniques were evaluated using a cancer image analysis workflow on a hybrid cluster with 256 nodes, each with an Intel Phi and a dual socket CPU. Our parallel execution method attained an efficiency of over 90% on 256 nodes. The hybrid execution on the CPU and Intel Phi improved the performance by 2×. Multilevel computation reuse led to performance gains of over 2.9×.
Willian de Oliveira Barreiros Junior, Jeremias Moreira, Tahsin M. Kurç, Jun Kong 0002, Alba Cristina Magalhaes Alves de Melo, Joel H. Saltz, George Teodoro
Concurr. Comput. Pract. Exp.6
2020 AI in Medical Imaging Informatics: Current Challenges and Future Directions
abstract
This paper reviews state-of-the-art research solutions across the spectrum of medical imaging informatics, discusses clinical translation, and provides future directions for advancing clinical practice. More specifically, it summarizes advances in medical imaging acquisition technologies for different modalities, highlighting the necessity for efficient medical data management strategies in the context of AI in big healthcare data analytics. It then provides a synopsis of contemporary and emerging algorithmic methods for disease classification and organ/ tissue segmentation, focusing on AI and deep learning architectures that have already become the de facto approach. The clinical benefits of in-silico modelling advances linked with evolving 3D reconstruction and visualization applications are further documented. Concluding, integrative analytics approaches driven by associate research branches highlighted in this study promise to revolutionize imaging informatics as known today across the healthcare continuum for both radiology and digital pathology applications. The latter, is projected to enable informed, more accurate diagnosis, timely prognosis, and effective treatment planning, underpinning precision medicine.
Andreas Panayides, Amir A. Amini, Nenad Filipovic, Ashish Sharma 0001, Sotirios A. Tsaftaris, Alistair A. Young, David J. Foran, Nhan Do, Spyretta Golemati, Tahsin M. Kurç, Kun Huang 0001, Konstantina S. Nikita, Benjamin Veasey, Michalis E. Zervakis, Joel H. Saltz, Constantinos S. Pattichis
IEEE J. Biomed. Health Informatics15
2019 Machine Learning Based Opioid Overdose Prediction Using Electronic Health Records
Sina Rashidian, Yu Wang 0137, Janos G. Hajagos, Richard N. Rosenthal, Jun Kong 0002, Mary M. Saltz, Joel H. Saltz, Fusheng Wang 0001
AMIA9
2019 Exascale Deep Learning to Accelerate Cancer Research
abstract
Deep learning, through the use of neural networks, has demonstrated remarkable ability to automate many routine tasks when presented with sufficient data for training. The neural network architecture (e.g. number of layers, types of layers, connections between layers, etc.) plays a critical role in determining what, if anything, the neural network is able to learn from the training data. The trend for neural network architectures, especially those trained on ImageNet, has been to grow ever deeper and more complex. The result has been ever increasing accuracy on benchmark datasets with the cost of increased computational demands. In this paper we demonstrate that neural network architectures can be automatically generated, tailored for a specific application, with dual objectives: accuracy of prediction and speed of prediction. Using MENNDL- an HPC-enabled software stack for neural architecture search-we generate a neural network with comparable accuracy to state-of-the-art networks on a cancer pathology dataset that is also 16× faster at inference. The speedup in inference is necessary because of the volume and velocity of cancer pathology data; specifically, the previous state-of-the-art networks are too slow for individual researchers without access to HPC systems to keep pace with the rate of data generation. Our new model enables researchers with modest computational resources to analyze newly generated data faster than it is collected.
Robert M. Patton, Shahira Abousamra, Dimitris Samaras, Joel H. Saltz, J. Travis Johnston, Steven R. Young, Catherine D. Schuman, Thomas E. Potok, Derek C. Rose, Seung-Hwan Lim, Junghoon Chae, Le Hou
IEEE BigData4
2019 Robust Histopathology Image Analysis: To Label or to Synthesize?
abstract
Detection, segmentation and classification of nuclei are fundamental analysis operations in digital pathology. Existing state-of-the-art approaches demand extensive amount of supervised training data from pathologists and may still perform poorly in images from unseen tissue types. We propose an unsupervised approach for histopathology image segmentation that synthesizes heterogeneous sets of training image patches, of every tissue type. Although our synthetic patches are not always of high quality, we harness the motley crew of generated samples through a generally applicable importance sampling method. This proposed approach, for the first time, re-weighs the training loss over synthetic data so that the ideal (unbiased) generalization loss over the true data distribution is minimized. This enables us to use a random polygon generator to synthesize approximate cellular structures (i.e., nuclear masks) for which no real examples are given in many tissue types, and hence, GAN-based methods are not suited. In addition, we propose a hybrid synthesis pipeline that utilizes textures in real histopathology patches and GAN models, to tackle heterogeneity in tissue textures. Compared with existing state-of-the-art supervised models, our approach generalizes significantly better on cancer types without training data. Even in cancer types with training data, our approach achieves the same performance without supervision cost. We release code and segmentation results on over 5000 Whole Slide Images (WSI) in The Cancer Genome Atlas (TCGA) repository, a dataset that would be orders of magnitude larger than what is available today.
Le Hou, Ayush Agarwal, Dimitris Samaras, Tahsin M. Kurç, Rajarsi Gupta 0001, Joel H. Saltz
CVPR6
2019 Label super-resolution networks
Kolya Malkin, Caleb Robinson, Le Hou, Rachel Soobitsky, Jacob Czawlytko, Dimitris Samaras, Joel H. Saltz, Lucas Joppa, Nebojsa Jojic
ICLR (Poster)7
2019 Pancreatic Cancer Detection in Whole Slide Images Using Noisy Label Annotations
Han Le, Dimitris Samaras, Tahsin M. Kurç, Rajarsi Gupta 0001, Kenneth Shroyer, Joel H. Saltz
MICCAI (1)6
2019 Sparse autoencoder for unsupervised nucleus detection and representation in histopathology images
Le Hou, Vu Nguyen 0004, Ariel B. Kanevsky, Dimitris Samaras, Tahsin M. Kurç, Rajarsi Gupta 0001, Yi Gao 0002, Wenjin Chen, David J. Foran, Joel H. Saltz
Pattern Recognit.11
2018 Systems Demonstration: Towards a Services Oriented Platform for Combined Radiology-Pathology Image Analysis and Interpretation
Tahsin M. Kurç, Joel H. Saltz, Fred W. Prior, Ashish Sharma 0001
AMIA2
2018 Social Media Based Analysis of Opioid Epidemic Using Reddit
Sheetal Pandrekar Mangesh, Xin Chen 0022, Gaurav Gopalkrishna, Avi Srivastava, Mary M. Saltz, Joel H. Saltz, Fusheng Wang 0001
AMIA6
2018 Cooperative and out-of-core execution of the irregular wavefront propagation pattern on hybrid machines with Intel® Xeon Phi™
abstract
The Irregular Wavefront Propagation Pattern (IWPP) is a core computing structure in several image analysis operations. Efficient implementation of IWPP on the Intel Xeon Phi is difficult because of the irregular data access and computation characteristics. The traditional IWPP algorithm relies on atomic instructions, which are not available in the SIMD set of the Intel Phi. To overcome this limitation, we have proposed a new IWPP algorithm that can take advantage of non-atomic SIMD instructions supported on the Intel Xeon Phi. We have also developed and evaluated methods to use CPU and Intel Phi cooperatively for parallel execution of the IWPP algorithms. Our new cooperative IWPP version is also able to handle large out-of-core images that would not fit into the memory of the accelerator. The new IWPP algorithm is used to implement the Morphological Reconstruction and Fill Holes operations, which are operations commonly found in image analysis applications. The vectorization implemented with the new IWPP has attained improvements of up to about 5× on top of the original IWPP and significant gains as compared to state-of-the-art the CPU and GPU versions. The new version running on an Intel Phi is 6.21× and 3.14× faster than running on a 16-core CPU and on a GPU, respectively. Finally, the cooperative execution using two Intel Phi devices and a multi-core CPU has reached performance gains of 2.14× as compared to the execution using a single Intel Xeon Phi.
Jeremias M. Gomes, Alba Cristina Magalhaes Alves de Melo, Jun Kong 0002, Tahsin M. Kurç, Joel H. Saltz, George Teodoro
Concurr. Comput. Pract. Exp.5
2017 ConvNets with Smooth Adaptive Activation Functions for Regression
abstract
Within Neural Networks (NN), the parameters of Adaptive Activation Functions (AAF) control the shapes of activation functions. These parameters are trained along with other parameters in the NN. AAFs have improved performance of Convolutional Neural Networks (CNN) in multiple classification tasks. In this paper, we propose and apply AAFs on CNNs for regression tasks. We argue that applying AAFs in the regression (second-to-last) layer of a NN can significantly decrease the bias of the regression NN. However, using existing AAFs may lead to overfitting. To address this problem, we propose a Smooth Adaptive Activation Function (SAAF) with a piecewise polynomial form which can approximate any continuous function to arbitrary degree of error, while having a bounded Lipschitz constant for given bounded model parameters. As a result, NNs with SAAF can avoid overfitting by simply regularizing model parameters. We empirically evaluated CNNs with SAAFs and achieved state-of-the-art results on age and pose estimation datasets.
Le Hou, Dimitris Samaras, Tahsin M. Kurç, Yi Gao 0002, Joel H. Saltz
AISTATS5
2017 Large-scale Analysis of Opioid Poisoning Related Hospital Visits in New York State
Xin Chen 0022, Yu Wang 0137, Xiaxia Yu, Elinor Schoenfeld, Mary M. Saltz, Joel H. Saltz, Fusheng Wang 0001
AMIA6
2017 20 Years of Digital Pathology - An Overview of the Road Travelled and What is on the Horizon
Joel H. Saltz, Ashish Sharma 0001, Alexis B. Carter, Liron Pantanowitz, Tahsin M. Kurç
AMIA1
2017 Parallel and Efficient Sensitivity Analysis of Microscopy Image Segmentation Workflows in Hybrid Systems
abstract
We investigate efficient sensitivity analysis (SA) of algorithms that segment and classify image features in a large dataset of high-resolution images. Algorithm SA is the process of evaluating variations of methods and parameter values to quantify differences in the output. A SA can be very compute demanding because it requires re-processing the input dataset several times with different parameters to assess variations in output. In this work, we introduce strategies to efficiently speed up SA via runtime optimizations targeting distributed hybrid systems and reuse of computations from runs with different parameters. We evaluate our approach using a cancer image analysis workflow on a hybrid cluster with 256 nodes, each with an Intel Phi and a dual socket CPU. The SA attained a parallel efficiency of over 90% on 256 nodes. The cooperative execution using the CPUs and the Phi available in each node with smart task assignment strategies resulted in an additional speedup of about 2×. Finally, multi-level computation reuse lead to an additional speedup of up to 2.46× on the parallel version. The level of performance attained with the proposed optimizations will allow the use of SA in large-scale studies.
Willian de Oliveira Barreiros Junior, George Teodoro, Tahsin M. Kurç, Jun Kong 0002, Alba Cristina Magalhaes Alves de Melo, Joel H. Saltz
CLUSTER6
2017 SparkGIS: Resource Aware Efficient In-Memory Spatial Query Processing
abstract
Much effort has been devoted to support high performance spatial queries on large volumes of spatial data in distributed spatial computing systems, especially in the MapReduce paradigm. Recent works have focused on extending spatial MapReduce frameworks to leverage high performance in-memory distributed processing capabilities of systems such as Spark. However, the performance advantage comes with the requirement of having enough memory and comprehensive configuration. Failing to fulfill this falls back to disk IO, defeating the purpose of such systems or in worst case gets out of memory and fails the job. The problem is aggravated further for spatial processing since the underlying in-memory systems are oblivious of spatial data features and characteristics. In this paper we present SparkGIS - an in-memory oriented spatial data querying system for high throughput and low latency spatial query handling by adapting Apache Spark's distributed processing capabilities. It supports basic spatial queries including containment, spatial join and k-nearest neighbor and allows extending these to complex query pipelines. SparkGIS mitigates skew in distributed processing by supporting several dynamic partitioning algorithms suitable for a rich set of contemporary application scenarios. Multilevel global and local, pre-generated and on-demand in-memory indexes, allow SparkGIS to prune input data and apply compute intensive operations on a subset of relevant spatial objects only. Finally, SparkGIS employs dynamic query rewriting to gracefully manage large spatial query workflows that exceed available distributed resources. Our comparative evaluation has shown that the performance of SparkGIS is on par with contemporary Spark based platforms for relatively smaller queries and outperforms them for larger data and memory intensive workflows by dynamic query rewriting and efficient spatial data management.
Furqan Baig, Hoang Vo, Tahsin M. Kurç, Joel H. Saltz, Fusheng Wang 0001
SIGSPATIAL/GIS4
2017 Classification of Pancreatic Cysts in Computed Tomography Images Using a Random Forest and Convolutional Neural Network Ensemble
Konstantin Dmitriev, Arie E. Kaufman, Ammar A. Javed, Ralph H. Hruban, Elliot K. Fishman, Anne Marie Lennon, Joel H. Saltz
MICCAI (3)7
2017 Center-Focusing Multi-task CNN with Injected Features for Classification of Glioma Nuclear Images
abstract
Classifying the various shapes and attributes of a glioma cell nucleus is crucial for diagnosis and understanding of the disease. We investigate the automated classification of the nuclear shapes and visual attributes of glioma cells, using Convolutional Neural Networks (CNNs) on pathology images of automatically segmented nuclei. We propose three methods that improve the performance of a previously-developed semi-supervised CNN. First, we propose a method that allows the CNN to focus on the most important part of an image-the image's center containing the nucleus. Second, we inject (concatenate) pre-extracted VGG features into an intermediate layer of our Semi-Supervised CNN so that during training, the CNN can learn a set of additional features. Third, we separate the losses of the two groups of target classes (nuclear shapes and attributes) into a single-label loss and a multi-label loss in order to incorporate prior knowledge of inter-label exclusiveness. On a dataset of 2078 images, the combination of the proposed methods reduces the error rate of attribute and shape classification by 21.54% and 15.07% respectively compared to the existing state-of-the-art method on the same dataset.
Veda Murthy, Le Hou, Dimitris Samaras, Tahsin M. Kurç, Joel H. Saltz
WACV5
2016 Safe "cloudification" of large images through picker APIs
Erich Bremer, Tahsin M. Kurç, Yi Gao 0002, Joel H. Saltz, Jonas S. Almeida
AMIA4
2016 POE: A Pathology Extraction Tool for Finding Attribute-Value Pairs in Glioma Pathology Reports
Veronica E. Lynn, Niranjan Balasubramanian, Tahsin M. Kurç, Joel H. Saltz, Rebecca S. Jacobson
AMIA4
2016 S2CR3UM: A Solution to the In Silico Relevance, Reliability & Reproducibility Conundrum
Sarah B. Putney, Janos G. Hajagos, Joel H. Saltz, Jonas S. Almeida, Mary M. Saltz
AMIA4
2016 Patch-Based Convolutional Neural Network for Whole Slide Tissue Image Classification
abstract
Convolutional Neural Networks (CNN) are state-of-the-art models for many image classification tasks. However, to recognize cancer subtypes automatically, training a CNN on gigapixel resolution Whole Slide Tissue Images (WSI) is currently computationally impossible. The differentiation of cancer subtypes is based on cellular-level visual features observed on image patch scale. Therefore, we argue that in this situation, training a patch-level classifier on image patches will perform better than or similar to an image-level classifier. The challenge becomes how to intelligently combine patch-level classification results and model the fact that not all patches will be discriminative. We propose to train a decision fusion model to aggregate patch-level predictions given by patch-level CNNs, which to the best of our knowledge has not been shown before. Furthermore, we formulate a novel Expectation-Maximization (EM) based method that automatically locates discriminative patches robustly by utilizing the spatial relationships of patches. We apply our method to the classification of glioma and non-small-cell lung carcinoma cases into subtypes. The classification accuracy of our method is similar to the inter-observer agreement between pathologists. Although it is impossible to train CNNs on WSIs, we experimentally demonstrate using a comparable non-cancer dataset of smaller images that a patch-based CNN can outperform an image-based CNN.
Le Hou, Dimitris Samaras, Tahsin M. Kurç, Yi Gao 0002, James Davis 0001, Joel H. Saltz
CVPR6
2015 OpenHealth Platform for Interactive Contextualization of Population Health Open Data
Jonas S. Almeida, Janos G. Hajagos, Ivan Crnosija, Tahsin M. Kurç, Mary M. Saltz, Joel H. Saltz
AMIA6
2015 Integrative Informatics and Predictive Modeling Support for Population Health
Mary M. Saltz, Joel H. Saltz, Janos G. Hajagos, Charles Boicey, Jim Murry, Ivan Crnosija, Tahsin M. Kurç, Erich Bremer, Jonas S. Almeida
AMIA2
2015 Efficient Irregular Wavefront Propagation Algorithms on Intel(R) Xeon Phi(TM)
abstract
We investigate the execution of the Irregular Wave front Propagation Pattern (IWPP), a fundamental computing structure used in several image analysis operations, on the Intel® Xeon PhiTM co-processor. An efficient implementation of IWPP on the Xeon Phi is a challenging problem because of IWPP's irregularity and the use of atomic instructions in the original IWPP algorithm to resolve race conditions. On the Xeon Phi, the use of SIMD and vectorization instructions is critical to attain high performance. However, SIMD atomic instructions are not supported. Therefore, we propose a new IWPP algorithm that can take advantage of the supported SIMD instruction set. We also evaluate an alternate storage container (priority queue) to track active elements in the wave front in an effort to improve the parallel algorithm efficiency. The new IWPP algorithm is evaluated with Morphological Reconstruction and Imfill operations as use cases. Our results show performance improvements of up to 5.63× on top of the original IWPP due to vectorization. Moreover, the new IWPP achieves speedups of 45.7× and 1.62×, respectively, as compared to efficient CPU and GPU implementations.
Jeremias M. Gomes, George Teodoro, Alba Cristina Magalhaes Alves de Melo, Jun Kong 0002, Tahsin M. Kurç, Joel H. Saltz
SBAC-PAD6
2015 Scalable analysis of Big pathology image data cohorts using efficient methods and high-performance computing strategies
abstract
BACKGROUND: We describe a suite of tools and methods that form a core set of capabilities for researchers and clinical investigators to evaluate multiple analytical pipelines and quantify sensitivity and variability of the results while conducting large-scale studies in investigative pathology and oncology. The overarching objective of the current investigation is to address the challenges of large data sizes and high computational demands. RESULTS: The proposed tools and methods take advantage of state-of-the-art parallel machines and efficient content-based image searching strategies. The content based image retrieval (CBIR) algorithms can quickly detect and retrieve image patches similar to a query patch using a hierarchical analysis approach. The analysis component based on high performance computing can carry out consensus clustering on 500,000 data points using a large shared memory system. CONCLUSIONS: Our work demonstrates efficient CBIR algorithms and high performance computing can be leveraged for efficient analysis of large microscopy images to meet the challenges of clinically salient applications in pathology. These technologies enable researchers and clinical investigators to make more effective use of the rich informational content contained within digitized microscopy specimens.
Tahsin M. Kurç, Xin Qi 0007, Daihou Wang, Fusheng Wang 0001, George Teodoro, Lee A. D. Cooper, Michael Nalisnik, Lin Yang 0002, Joel H. Saltz, David J. Foran
BMC Bioinform.9
2014 Comparative Performance Analysis of Intel (R) Xeon Phi (TM), GPU, and CPU: A Case Study from Microscopy Image Analysis
abstract
We study and characterize the performance of operations in an important class of applications on GPUs and Many Integrated Core (MIC) architectures. Our work is motivated by applications that analyze low-dimensional spatial datasets captured by high resolution sensors, such as image datasets obtained from whole slide tissue specimens using microscopy scanners. Common operations in these applications involve the detection and extraction of objects (object segmentation), the computation of features of each extracted object (feature computation), and characterization of objects based on these features (object classification). In this work, we have identify the data access and computation patterns of operations in the object segmentation and feature computation categories. We systematically implement and evaluate the performance of these operations on modern CPUs, GPUs, and MIC systems for a microscopy image analysis application. Our results show that the performance on a MIC of operations that perform regular data access is comparable or sometimes better than that on a GPU. On the other hand, GPUs are significantly more efficient than MICs for operations that access data irregularly. This is a result of the low performance of MICs when it comes to random data access. We also have examined the coordinated use of MICs and CPUs. Our experiments show that using a performance aware task strategy for scheduling application operations improves performance about 1.29× over a first-come-first-served strategy. This allows applications to obtain high performance efficiency on CPU-MIC systems - the example application attained an efficiency of 84% on 192 nodes (3072 CPU cores and 192 MICs).
George Teodoro, Tahsin M. Kurç, Jun Kong 0002, Lee A. D. Cooper, Joel H. Saltz
IPDPS5
2014 Efficient Execution of Microscopy Image Analysis on CPU, GPU, and MIC Equipped Cluster Systems
abstract
High performance computing is experiencing a major paradigm shift with the introduction of accelerators, such as graphics processing units (GPUs) and Intel Xeon Phi (MIC). These processors have made available a tremendous computing power at low cost, and are transforming machines into hybrid systems equipped with CPUs and accelerators. Although these systems can deliver a very high peak performance, making full use of its resources in real-world applications is a complex problem. Most current applications deployed to these machines are still being executed in a single processor, leaving other devices underutilized. In this paper we explore a scenario in which applications are composed of hierarchical data flow tasks which are allocated to nodes of a distributed memory machine in coarse-grain, but each of them may be composed of several finer-grain tasks which can be allocated to different devices within the node. We propose and implement novel performance aware scheduling techniques that can be used to allocate tasks to devices. We evaluate our techniques using a pathology image analysis application used to investigate brain cancer morphology, and our experimental evaluation shows that the proposed scheduling strategies significantly outperforms other efficient scheduling techniques, such as Heterogeneous Earliest Finish Time - HEFT, in cooperative executions using CPUs, GPUs, and MICs. We also experimentally show that our strategies are less sensitive to inaccuracy in the scheduling input data and that the performance gains are maintained as the application scales.
Guilherme Andrade, Renato Ferreira 0001, George Teodoro, Leonardo Rocha 0001, Joel H. Saltz, Tahsin M. Kurç
SBAC-PAD5
2014 GlycoPattern: a web platform for glycan array mining
abstract
UNLABELLED: GlycoPattern is Web-based bioinformatics resource to support the analysis of glycan array data for the Consortium for Functional Glycomics. This resource includes algorithms and tools to discover structural motifs, a heatmap visualization to compare multiple experiments, hierarchical clustering of Glycan Binding Proteins with respect to their binding motifs and a structural search feature on the experimental data. AVAILABILITY AND IMPLEMENTATION: GlycoPattern is freely available on the Web at http://glycopattern.emory.edu with all major browsers supported.
Sanjay Agravat, Joel H. Saltz, Richard D. Cummings, David F. Smith
Bioinform.2
2014 Parallel content-based sub-image retrieval using hierarchical searching
abstract
MOTIVATION: The capacity to systematically search through large image collections and ensembles and detect regions exhibiting similar morphological characteristics is central to pathology diagnosis. Unfortunately, the primary methods used to search digitized, whole-slide histopathology specimens are slow and prone to inter- and intra-observer variability. The central objective of this research was to design, develop, and evaluate a content-based image retrieval system to assist doctors for quick and reliable content-based comparative search of similar prostate image patches. METHOD: Given a representative image patch (sub-image), the algorithm will return a ranked ensemble of image patches throughout the entire whole-slide histology section which exhibits the most similar morphologic characteristics. This is accomplished by first performing hierarchical searching based on a newly developed hierarchical annular histogram (HAH). The set of candidates is then further refined in the second stage of processing by computing a color histogram from eight equally divided segments within each square annular bin defined in the original HAH. A demand-driven master-worker parallelization approach is employed to speed up the searching procedure. Using this strategy, the query patch is broadcasted to all worker processes. Each worker process is dynamically assigned an image by the master process to search for and return a ranked list of similar patches in the image. RESULTS: The algorithm was tested using digitized hematoxylin and eosin (H&E) stained prostate cancer specimens. We have achieved an excellent image retrieval performance. The recall rate within the first 40 rank retrieved image patches is ∼90%. AVAILABILITY AND IMPLEMENTATION: Both the testing data and source code can be downloaded from http://pleiad.umdnj.edu/CBII/Bioinformatics/.
Lin Yang 0002, Xin Qi 0007, Fuyong Xing, Tahsin M. Kurç, Joel H. Saltz, David J. Foran
Bioinform.5
2014 Region templates: Data representation and management for high-throughput image analysis
abstract
We introduce a region template abstraction and framework for the efficient storage, management and processing of common data types in analysis of large datasets of high resolution images on clusters of hybrid computing nodes. The region template abstraction provides a generic container template for common data structures, such as points, arrays, regions, and object sets, within a spatial and temporal bounding box. It allows for different data management strategies and I/O implementations, while providing a homogeneous, unified interface to applications for data storage and retrieval. A region template application is represented as a hierarchical dataflow in which each computing stage may be represented as another dataflow of finer-grain tasks. The execution of the application is coordinated by a runtime system that implements optimizations for hybrid machines, including performance-aware scheduling for maximizing the utilization of computing devices and techniques to reduce the impact of data transfers between CPUs and GPUs. An experimental evaluation on a state-of-the-art hybrid cluster using a microscopy imaging application shows that the abstraction adds negligible overhead (about 3%) and achieves good scalability and high data transfer rates. Optimizations in a high speed disk based storage implementation of the abstraction to support asynchronous data transfers and computation result in an application performance gain of about 1.13×. Finally, a processing rate of 11,730 4K×4K tiles per minute was achieved for the microscopy imaging application on a cluster with 100 nodes (300 GPUs and 1,200 CPU cores). This computation rate enables studies with very large datasets.
George Teodoro, Tony Pan, Tahsin M. Kurç, Jun Kong 0002, Lee A. D. Cooper, Scott Klasky, Joel H. Saltz
Parallel Comput.7
2014 Approximate similarity search for online multimedia services on distributed CPU-GPU platforms
George Teodoro, Eduardo Valle, Nathan Mariano, Ricardo da Silva Torres, Wagner Meira Jr., Joel H. Saltz
VLDB J.6
2013 Temporal Abstraction-based Clinical Phenotyping with Eureka!
Andrew R. Post, Tahsin M. Kurç, Richie Willard, Himanshu Rathod, Michel Mansour, Akshatha Kalsanka Pai, William M. Torian, Sanjay Agravat, Suzanne Sturm, Joel H. Saltz
AMIA10
2013 Clinical Phenotyping with the Analytic Information Warehouse
Andrew R. Post, Tahsin M. Kurç, Richie Willard, Himanshu Rathod, Michel Mansour, Akshatha Kalsanka Pai, William M. Torian, Sanjay Agravat, Suzanne Sturm, Joel H. Saltz
AMIA10
2013 High-performance computational analysis of glioblastoma pathology images with database support identifies molecular and survival correlates
abstract
In this paper, we present a novel framework for microscopic image analysis of nuclei, data management, and high performance computation to support translational research involving nuclear morphometry features, molecular data, and clinical outcomes. Our image analysis pipeline consists of nuclei segmentation and feature computation facilitated by high performance computing with coordinated execution in multi-core CPUs and Graphical Processor Units (GPUs). All data derived from image analysis are managed in a spatial relational database supporting highly efficient scientific queries. We applied our image analysis workflow to 159 glioblastomas (GBM) from The Cancer Genome Atlas dataset. With integrative studies, we found statistics of four specific nuclear features were significantly associated with patient survival. Additionally, we correlated nuclear features with molecular data and found interesting results that support pathologic domain knowledge. We found that Proneural subtype GBMs had the smallest mean of nuclear Eccentricity and the largest mean of nuclear Extent, and MinorAxisLength. We also found gene expressions of stem cell marker MYC and cell proliferation maker MKI67 were correlated with nuclear features. To complement and inform pathologists of relevant diagnostic features, we queried the most representative nuclear instances from each patient population based on genetic and transcriptional classes. Our results demonstrate that specific nuclear features carry prognostic significance and associations with transcriptional and genetic classes, highlighting the potential of high throughput pathology image analysis as a complementary approach to human-based review and translational research.
Jun Kong 0002, Fusheng Wang 0001, George Teodoro, Lee A. D. Cooper, Carlos Sanchez Moreno, Tahsin M. Kurç, Tony Pan, Joel H. Saltz, Daniel J. Brat
BIBM8
2013 Demonstration of Hadoop-GIS: a spatial data warehousing system over MapReduce
abstract
- a scalable and high performance spatial query system over MapReduce. Hadoop-GIS provides an efficient spatial query engine to process spatial queries, data and space based partitioning, and query pipelines that parallelize queries implicitly on MapReduce. Hadoop-GIS also provides an expressive, SQL-like spatial query language for workload specification. We will demonstrate how spatial queries are expressed in spatially extended SQL queries, and submitted through a command line/web interface for execution. Parallel to our system demonstration, we explain the system architecture and details on how queries are translated to MapReduce operators, optimized, and executed on Hadoop. In addition, we will showcase how the system can be used to support two representative real world use cases: large scale pathology analytical imaging, and geo-spatial data warehousing.
Ablimit Aji, Xiling Sun, Hoang Vo, Qiaoling Liu, Rubao Lee, Xiaodong Zhang 0001, Joel H. Saltz, Fusheng Wang 0001
SIGSPATIAL/GIS7
2013 High-throughput Analysis of Large Microscopy Image Datasets on CPU-GPU Cluster Platforms
abstract
Analysis of large pathology image datasets offers significant opportunities for the investigation of disease morphology, but the resource requirements of analysis pipelines limit the scale of such studies. Motivated by a brain cancer study, we propose and evaluate a parallel image analysis application pipeline for high throughput computation of large datasets of high resolution pathology tissue images on distributed CPU-GPU platforms. To achieve efficient execution on these hybrid systems, we have built runtime support that allows us to express the cancer image analysis application as a hierarchical data processing pipeline. The application is implemented as a coarse-grain pipeline of stages, where each stage may be further partitioned into another pipeline of fine-grain operations. The fine-grain operations are efficiently managed and scheduled for computation on CPUs and GPUs using performance aware scheduling techniques along with several optimizations, including architecture aware process placement, data locality conscious task assignment, data prefetching, and asynchronous data copy. These optimizations are employed to maximize the utilization of the aggregate computing power of CPUs and GPUs and minimize data copy overheads. Our experimental evaluation shows that the cooperative use of CPUs and GPUs achieves significant improvements on top of GPU-only versions (up to 1.6×) and that the execution of the application as a set of fine-grain operations provides more opportunities for runtime optimizations and attains better performance than coarser-grain, monolithic implementations used in other works. An implementation of the cancer image analysis pipeline using the runtime support was able to process an image dataset consisting of 36,848 4Kx4K-pixel image tiles (about 1.8TB uncompressed) in less than 4 minutes (150 tiles/second) on 100 nodes of a state-of-the-art hybrid cluster system.
George Teodoro, Tony Pan, Tahsin M. Kurç, Jun Kong 0002, Lee A. D. Cooper, Norbert Podhorszki, Scott Klasky, Joel H. Saltz
IPDPS8
2013 Research and applications: Cancer Digital Slide Archive: an informatics resource to support integrated in silico analysis of TCGA pathology data
abstract
BACKGROUND: The integration and visualization of multimodal datasets is a common challenge in biomedical informatics. Several recent studies of The Cancer Genome Atlas (TCGA) data have illustrated important relationships between morphology observed in whole-slide images, outcome, and genetic events. The pairing of genomics and rich clinical descriptions with whole-slide imaging provided by TCGA presents a unique opportunity to perform these correlative studies. However, better tools are needed to integrate the vast and disparate data types. OBJECTIVE: To build an integrated web-based platform supporting whole-slide pathology image visualization and data integration. MATERIALS AND METHODS: All images and genomic data were directly obtained from the TCGA and National Cancer Institute (NCI) websites. RESULTS: The Cancer Digital Slide Archive (CDSA) produced is accessible to the public (http://cancer.digitalslidearchive.net) and currently hosts more than 20,000 whole-slide images from 22 cancer types. DISCUSSION: The capabilities of CDSA are demonstrated using TCGA datasets to integrate pathology imaging with associated clinical, genomic and MRI measurements in glioblastomas and can be extended to other tumor types. CDSA also allows URL-based sharing of whole-slide images, and has preliminary support for directly sharing regions of interest and other annotations. Images can also be selected on the basis of other metadata, such as mutational profile, patient age, and other relevant characteristics. CONCLUSIONS: With the increasing availability of whole-slide scanners, analysis of digitized pathology images will become increasingly important in linking morphologic observations with genomic and clinical endpoints.
David A. Gutman, Jake Cobb, Dhananjaya Somanna, Yuna Park, Fusheng Wang 0001, Tahsin M. Kurç, Joel H. Saltz, Daniel J. Brat, Lee A. D. Cooper
J. Am. Medical Informatics Assoc.7
2013 The Analytic Information Warehouse (AIW): A platform for analytics using electronic health record data
Andrew R. Post, Tahsin M. Kurç, Sharath R. Cholleti, Xia Lin, William Bornstein, Dedra Cantrell, Sam Hohmann, Joel H. Saltz
J. Biomed. Informatics10
2013 Efficient irregular wavefront propagation algorithms on hybrid CPU-GPU machines
George Teodoro, Tony Pan, Tahsin M. Kurç, Jun Kong 0002, Lee A. D. Cooper, Joel H. Saltz
Parallel Comput.6
2013 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce
abstract
Support of high performance queries on large volumes of spatial data becomes increasingly important in many application domains, including geospatial problems in numerous fields, location based services, and emerging scientific applications that are increasingly data- and compute-intensive. The emergence of massive scale spatial data is due to the proliferation of cost effective and ubiquitous positioning technologies, development of high resolution imaging technologies, and contribution from a large number of community users. There are two major challenges for managing and querying massive spatial data to support spatial queries: the explosion of spatial data, and the high computational complexity of spatial queries. In this paper, we present Hadoop-GIS - a scalable and high performance spatial data warehousing system for running large scale spatial queries on Hadoop. Hadoop-GIS supports multiple types of spatial queries on MapReduce through spatial partitioning, customizable spatial query engine RESQUE, implicit parallel spatial query execution on MapReduce, and effective methods for amending query results through handling boundary objects. Hadoop-GIS utilizes global partition indexing and customizable on demand local spatial indexing to achieve efficient query processing. Hadoop-GIS is integrated into Hive to support declarative spatial queries with an integrated architecture. Our experiments have demonstrated the high efficiency of Hadoop-GIS on query response and high scalability to run on commodity clusters. Our comparative experiments have showed that performance of Hadoop-GIS is on par with parallel SDBMS and outperforms SDBMS for compute-intensive queries. Hadoop-GIS is available as a set of library for processing spatial queries, and as an integrated software package in Hive.
Ablimit Aji, Fusheng Wang 0001, Hoang Vo, Rubao Lee, Qiaoling Liu, Xiaodong Zhang 0001, Joel H. Saltz
Proc. VLDB Endow.7
2012 Leveraging Derived Data Elements in Data Analytic Models for Understanding and Predicting Hospital Readmissions
Sharath R. Cholleti, Andrew R. Post, Xia Lin, William Bornstein, Dedra Cantrell, Joel H. Saltz
AMIA7
2012 A Pipeline Supporting Efficient and Flexible Analytics Using Longitudinal and High-dimensional Extracts from Clinical Databases
Andrew R. Post, Sharath R. Cholleti, Himanshu Rathod, William Bornstein, Dedra Cantrell, Joel H. Saltz
AMIA7
2012 Systematic Modeling, Testing, and Monitoring of Information Integrity in Federated Ontology-driven Data Sources
Mijung Kim, Jake Cobb, Tahsin M. Kurç, Alessandro Orso, Mary Jean Harrold, Andrew R. Post, Shamkant B. Navathe, Joel H. Saltz
AMIA8
2012 High Performance Computing for Integrative Analysis of Large Pathology Image Datasets
Tahsin M. Kurç, Joel H. Saltz, George Teodoro, Tony Pan, Lee A. D. Cooper, Jun Kong 0002, David A. Gutman, Daniel J. Brat, Fusheng Wang 0001
AMIA2
2012 Enabling ontology based semantic queries in biomedical database systems
abstract
While current biomedical ontology repositories offer primitive query capabilities, it is difficult or cumbersome to support ontology based semantic queries directly in semantically annotated biomedical databases. The problem may be largely attributed to the mismatch between the models of the ontologies and the databases, and the mismatch between the query interfaces of the two systems. To fully realize semantic query capabilities based on ontologies, we develop a system DBOntoLink to provide unified semantic query interfaces by extending database query languages. With DBOntoLink, semantic queries can be directly and naturally specified as extended functions of the database query languages without any programming needed. DBOntoLink is adaptable to different ontologies through customizations and supports major biomedical ontologies hosted at the NCBO BioPortal. We demonstrate the use of DBOntoLink in a real world biomedical database with semantically annotated medical image annotations.
Shuai Zheng 0003, Fusheng Wang 0001, James J. Lu, Joel H. Saltz
CIKM4
2012 Towards building a high performance spatial query system for large scale medical imaging data
abstract
Support of high performance queries on large volumes of scientific spatial data is becoming increasingly important in many applications. This growth is driven by not only geospatial problems in numerous fields, but also emerging scientific applications that are increasingly data- and compute-intensive. For example, digital pathology imaging has become an emerging field during the past decade, where examination of high resolution images of human tissue specimens enables more effective diagnosis, prediction and treatment of diseases. Systematic analysis of large-scale pathology images generates tremendous amounts of spatially derived quantifications of micro-anatomic objects, such as nuclei, blood vessels, and tissue regions. Analytical pathology imaging provides high potential to support image based computer aided diagnosis. One major requirement for this is effective querying of such enormous amount of data with fast response, which is faced with two major challenges: the "big data" challenge and the high computation complexity. In this paper, we present our work towards building a high performance spatial query system for querying massive spatial data on MapReduce. Our framework takes an on demand index building approach for processing spatial queries and a partition-merge approach for building parallel spatial query pipelines, which fits nicely with the computing model of MapReduce. We demonstrate our framework on supporting multi-way spatial joins for algorithm evaluation and nearest neighbor queries for microanatomic objects. To reduce query response time, we propose cost based query optimization to mitigate the effect of data skew. Our experiments show that the framework can efficiently support complex analytical spatial queries on MapReduce.
Ablimit Aji, Fusheng Wang 0001, Joel H. Saltz
SIGSPATIAL/GIS3
2012 Accelerating Large Scale Image Analyses on Parallel, CPU-GPU Equipped Systems
abstract
The past decade has witnessed a major paradigm shift in high performance computing with the introduction of accelerators as general purpose processors. These computing devices make available very high parallel computing power at low cost and power consumption, transforming current high performance platforms into heterogeneous CPU-GPU equipped systems. Although the theoretical performance achieved by these hybrid systems is impressive, taking practical advantage of this computing power remains a very challenging problem. Most applications are still deployed to either GPU or CPU, leaving the other resource under- or un-utilized. In this paper, we propose, implement, and evaluate a performance aware scheduling technique along with optimizations to make efficient collaborative use of CPUs and GPUs on a parallel system. In the context of feature computations in large scale image analysis applications, our evaluations show that intelligently co-scheduling CPUs and GPUs can significantly improve performance over GPU-only or multi-core CPU-only approaches.
George Teodoro, Tahsin M. Kurç, Tony Pan, Lee A. D. Cooper, Jun Kong 0002, Patrick M. Widener, Joel H. Saltz
IPDPS7
2012 Efficient regression testing of ontology-driven systems
abstract
To manage and integrate information gathered from heterogeneous databases, an ontology is often used. Like all systems, ontology-driven systems evolve over time and must be regression tested to gain confidence in the behavior of the modified system. Because rerunning all existing tests can be extremely expensive, researchers have developed regression-test-selection (RTS) techniques that select a subset of the available tests that are affected by the changes, and use this subset to test the modified system. Existing RTS techniques have been shown to be effective, but they operate on the code and are unable to handle changes that involve ontologies. To address this limitation, we developed and present in this paper a novel RTS technique that targets ontology-driven systems. Our technique creates representations of the old and new ontologies, compares them to identify entities affected by the changes, and uses this information to select the subset of tests to rerun. We also describe in this paper OntoRetest, a tool that implements our technique and that we used to empirically evaluate our approach on two biomedical ontology-driven database systems. The results of our evaluation show that our technique is both efficient and effective in selecting tests to rerun and in reducing the overall time required to perform regression testing.
Mijung Kim, Jake Cobb, Mary Jean Harrold, Tahsin M. Kurç, Alessandro Orso, Joel H. Saltz, Andrew R. Post, Kunal Malhotra, Shamkant B. Navathe
ISSTA6
2012 Integrated morphologic analysis for the identification and characterization of disease subtypes
abstract
BACKGROUND AND OBJECTIVE: Morphologic variations of disease are often linked to underlying molecular events and patient outcome, suggesting that quantitative morphometric analysis may provide further insight into disease mechanisms. In this paper a methodology for the subclassification of disease is developed using image analysis techniques. Morphologic signatures that represent patient-specific tumor morphology are derived from the analysis of hundreds of millions of cells in digitized whole slide images. Clustering these signatures aggregates tumors into groups with cohesive morphologic characteristics. This methodology is demonstrated with an analysis of glioblastoma, using data from The Cancer Genome Atlas to identify a prognostically significant morphology-driven subclassification, in which clusters are correlated with transcriptional, genetic, and epigenetic events. MATERIALS AND METHODS: Methodology was applied to 162 glioblastomas from The Cancer Genome Atlas to identify morphology-driven clusters and their clinical and molecular correlates. Signatures of patient-specific tumor morphology were generated from analysis of 200 million cells in 462 whole slide images. Morphology-driven clusters were interrogated for associations with patient outcome, response to therapy, molecular classifications, and genetic alterations. An additional layer of deep, genome-wide analysis identified characteristic transcriptional, epigenetic, and copy number variation events. RESULTS AND DISCUSSION: Analysis of glioblastoma identified three prognostically significant patient clusters (median survival 15.3, 10.7, and 13.0 months, log rank p=1.4e-3). Clustering results were validated in a separate dataset. Clusters were characterized by molecular events in nuclear compartment signaling including developmental and cell cycle checkpoint pathways. This analysis demonstrates the potential of high-throughput morphometrics for the subclassification of disease, establishing an approach that complements genomics.
Lee A. D. Cooper, Jun Kong 0002, David A. Gutman, Fusheng Wang 0001, Christina Appin, Sharath R. Cholleti, Tony Pan, Ashish Sharma 0001, Lisa Scarpace, Tom Mikkelsen, Tahsin M. Kurç, Carlos Sanchez Moreno, Daniel J. Brat, Joel H. Saltz
J. Am. Medical Informatics Assoc.15
2012 Digital Pathology: Data-Intensive Frontier in Medical Imaging
abstract
Pathology is a medical subspecialty that practices the diagnosis of disease. Microscopic examination of tissue reveals information enabling the pathologist to render accurate diagnoses and to guide therapy. The basic process by which anatomic pathologists render diagnoses has remained relatively unchanged over the last century, yet advances in information technology now offer significant opportunities in image-based diagnostic and research applications. Pathology has lagged behind other healthcare practices such as radiology where digital adoption is widespread. As devices that generate whole slide images become more practical and affordable, practices will increasingly adopt this technology and eventually produce an explosion of data that will quickly eclipse the already vast quantities of radiology imaging data. These advances are accompanied by significant challenges for data management and storage, but they also introduce new opportunities to improve patient care by streamlining and standardizing diagnostic approaches and uncovering disease mechanisms. Computer-based image analysis is already available in commercial diagnostic systems, but further advances in image analysis algorithms are warranted in order to fully realize the benefits of digital pathology in medical discovery and patient care. In coming decades, pathology image analysis will extend beyond the streamlining of diagnostic workflows and minimizing interobserver variability and will begin to provide diagnostic assistance, identify therapeutic targets, and predict patient outcomes and therapeutic responses.
Lee A. D. Cooper, Alexis B. Carter, Alton B. Farris, Fusheng Wang 0001, Jun Kong 0002, David A. Gutman, Patrick M. Widener, Tony Pan, Sharath R. Cholleti, Ashish Sharma 0001, Tahsin M. Kurç, Daniel J. Brat, Joel H. Saltz
Proc. IEEE13
2012 Accelerating Pathology Image Data Cross-Comparison on CPU-GPU Hybrid Systems
abstract
As an important application of spatial databases in pathology imaging analysis, cross-comparing the spatial boundaries of a huge amount of segmented micro-anatomic objects demands extremely data- and compute-intensive operations, requiring high throughput at an affordable cost. However, the performance of spatial database systems has not been satisfactory since their implementations of spatial operations cannot fully utilize the power of modern parallel hardware. In this paper, we provide a customized software solution that exploits GPUs and multi-core CPUs to accelerate spatial cross-comparison in a cost-effective way. Our solution consists of an efficient GPU algorithm and a pipelined system framework with task migration support. Extensive experiments with real-world data sets demonstrate the effectiveness of our solution, which improves the performance of spatial cross-comparison by over 18 times compared with a parallelized spatial database approach.
Kaibo Wang, Yin Huai, Rubao Lee, Fusheng Wang 0001, Xiaodong Zhang 0001, Joel H. Saltz
Proc. VLDB Endow.6
2011 Computer-Based Image Analysis of Liver Steatosis with Large-Scale Microscopy Imagery and Correlation with Magnetic Resonance Imaging Lipid Analysis
abstract
Most pathology analyses and measurements are prevalently carried out by trained reviewers in both clinical and research settings. Therefore, the resulting outputs are inexorably biased by interpreters and degraded with poor reproducibility. In this paper, we propose a computerized image analysis paradigm enabling quantitative characterizations of steatosis areas in microscopy images of pediatric liver biopsies. With the same set of patients, we also acquired the lipid measurements from magnetic resonance imaging data analysis for correlation investigation. Our preliminary results suggest a high correlation between the steatosis areas quantized with microscopy images and the lipid percentages calculated from radiology imaging data. Additionally, we compared the per formance of the proposed analysis method with those of three certified pathologists and a popular commercial algorithm. The results suggest the superiority of our method to both human reviewers and the commercial method in terms of the steatosis lipid correlation strength. This demonstrates that the developed method is promising for generating quantitative and reliable analysis results to better support further liver disease study.
Jun Kong 0002, Michael J. Lee, Pelin Bagci, Diego R. Martín, N. Volkan Adsay, Joel H. Saltz, Alton B. Farris
BIBM7
2011 Detection of Conflicts and Inconsistencies in Taxonomy-Based Authorization Policies
abstract
The values of data elements stored in biomedical databases often draw from biomedical ontologies. Authorization rules can be defined on these ontologies to control access to sensitive and private data elements in such databases. Authorization rules may be specified by different authorities at different times for various purposes. Since such policy rules can conflict with each other, access to sensitive information may inadvertently be allowed. Another problem in biomedical data protection is inference attacks, in which a user who has legitimate access to some data elements is able to infer information related to other data elements. We propose and evaluate two strategies; one for detecting policy inconsistencies to avoid potential inference attacks and the other for detecting policy conflicts.
Apurva Mohan, Douglas M. Blough, Tahsin M. Kurç, Andrew R. Post, Joel H. Saltz
BIBM5
2011 ImageMiner: a software system for comparative analysis of tissue microarrays using content-based image retrieval, high-performance computing, and grid technology
abstract
OBJECTIVE AND DESIGN: The design and implementation of ImageMiner, a software platform for performing comparative analysis of expression patterns in imaged microscopy specimens such as tissue microarrays (TMAs), is described. ImageMiner is a federated system of services that provides a reliable set of analytical and data management capabilities for investigative research applications in pathology. It provides a library of image processing methods, including automated registration, segmentation, feature extraction, and classification, all of which have been tailored, in these studies, to support TMA analysis. The system is designed to leverage high-performance computing machines so that investigators can rapidly analyze large ensembles of imaged TMA specimens. To support deployment in collaborative, multi-institutional projects, ImageMiner features grid-enabled, service-based components so that multiple instances of ImageMiner can be accessed remotely and federated. RESULTS: The experimental evaluation shows that: (1) ImageMiner is able to support reliable detection and feature extraction of tumor regions within imaged tissues; (2) images and analysis results managed in ImageMiner can be searched for and retrieved on the basis of image-based features, classification information, and any correlated clinical data, including any metadata that have been generated to describe the specified tissue and TMA; and (3) the system is able to reduce computation time of analyses by exploiting computing clusters, which facilitates analysis of larger sets of tissue samples.
David J. Foran, Lin Yang 0002, Wenjin Chen, Lauri A. Goodell, Michael Reiss, Fusheng Wang 0001, Tahsin M. Kurç, Tony Pan, Ashish Sharma 0001, Joel H. Saltz
J. Am. Medical Informatics Assoc.11
2011 Optimizing latency and throughput of application workflows on clusters
Nagavijayalakshmi Vydyanathan, Ümit V. Çatalyürek, Tahsin M. Kurç, P. Sadayappan, Joel H. Saltz
Parallel Comput.5
2010 Texture based image recognition in microscopy images of diffuse gliomas with multi-class gentle boosting mechanism
abstract
The diagnosis of diffuse gliomas requires the careful inspection of large amounts of visual data. Identifying tissue regions that inform diagnosis is a cumbersome task for human reviewers and is a process prone to inter-reader variability. In this paper we present an automatic method for identifying critical diagnostic regions within whole-slide microscopy images of gliomas. We frame the problem of critical region identification as a texture-based content retrieval task in the sense that each image is represented by a set of texture features. Both linear and nonlinear dimensionality reduction techniques are utilized to explore the intrinsic dimensionality of the feature space where images are classified by classification and regression trees with performances improved by a newly extended multi-class gentle boosting (MCGB) mechanism. The proposed method is demonstrated on 1200 sample regions using a five-fold cross validation, achieving a 96.25% classification accuracy.
Jun Kong 0002, Lee A. D. Cooper, Ashish Sharma 0001, Tahsin M. Kurç, Daniel J. Brat, Joel H. Saltz
ICASSP6
2010 Exploring the performance of massively multithreaded architectures
abstract
Abstract We present a new scheme for evaluating the performance of multithreaded computers and demonstrate its application to the Cray MTA‐2 and XMT supercomputers. Our scheme is based on the concept of clock cycles per element,\documentclass{article}\footskip=0pc\pagestyle{empty}\begin{document}${\cal C}$\end{document} , plotted againstbothproblem sizeandthe number of processors. This scheme clearly shows if an implementation has achieved its asymptotic efficiency and is more general than (but includes) the commonly used speedup metric. It permits the discovery of any imperfections in both the software as well as the hardware, and is expected to permit a unified comparison of many different parallel architectures. Measurements on a number of well‐known parallel algorithms, ranging from matrix multiply to quicksort, are presented for the MTA‐2 and XMT and highlight some interesting differences between these machines. The performance of sequence alignment using dynamic programming is evaluated on the MTA‐2, XMT, IBM x3755 and SGI Altix 350 and provides a useful comparison of the capabilities of the Cray machines with more conventional shared memory architectures. Copyright © 2009 John Wiley & Sons, Ltd.
Shahid H. Bokhari, Joel H. Saltz
Concurr. Comput. Pract. Exp.2
2009 An integrated framework for performance-based optimization of scientific workflows
abstract
Data analysis processes in scientific applications can be expressed as coarse-grain workflows of complex data processing operations with data flow dependencies between them. Performance optimization of these workflows can be viewed as a search for a set of optimal values in a multi-dimensional parameter space. While some performance parameters such as grouping of workflow components and their mapping to machines do not a ect the accuracy of the output, others may dictate trading the output quality of individual components (and of the whole workflow) for performance. This paper describes an integrated framework which is capable of supporting performance optimizations along multiple dimensions of the parameter space. Using two real-world applications in the spatial data analysis domain, we present an experimental evaluation of the proposed framework.
Vijay S. Kumar, P. Sadayappan, Gaurang Mehta, Karan Vahi, Ewa Deelman, Varun Ratnakar, Jihie Kim, Yolanda Gil, Mary W. Hall, Tahsin M. Kurç, Joel H. Saltz
HPDC11
2009 Architectural implications for spatial object association algorithms
abstract
Spatial object association, also referred to as crossmatch of spatial datasets, is the problem of identifying and comparing objects in two or more datasets based on their positions in a common spatial coordinate system. In this work, we evaluate two crossmatch algorithms that are used for astronomical sky surveys, on the following database system architecture configurations: (1) Netezza Performance Serverreg, a parallel database system with active disk style processing capabilities, (2) MySQL Cluster, a high-throughput network database system, and (3) a hybrid configuration consisting of a collection of independent database system instances with data replication support. Our evaluation provides insights about how architectural characteristics of these systems affect the performance of the spatial crossmatch algorithms. We conducted our study using real use-case scenarios borrowed from a large-scale astronomy application known as the large synoptic survey telescope (LSST).
Vijay S. Kumar, Tahsin M. Kurç, Joel H. Saltz, Ghaleb Abdulla, Scott R. Kohn, Celeste Matarazzo
IPDPS3
2009 Tensor classification of N-point correlation function features for histology tissue segmentation
Kishore Mosaliganti, Firdaus Janoos, M. Okan Irfanoglu, Randall Ridgway, Raghu Machiraju, Kun Huang 0001, Joel H. Saltz, Gustavo Leone, Michael C. Ostrowski
Medical Image Anal.7
2009 Computer-aided evaluation of neuroblastoma on whole-slide histology images: Classifying grade of neuroblastic differentiation
Jun Kong 0002, Olcay Sertel, Hiroyuki Shimada, Kim L. Boyer, Joel H. Saltz, Metin Nafi Gürcan
Pattern Recognit.5
2009 Computer-aided prognosis of neuroblastoma on whole-slide images: Classification of stromal development
Olcay Sertel, Jun Kong 0002, Hiroyuki Shimada, Ümit V. Çatalyürek, Joel H. Saltz, Metin Nafi Gürcan
Pattern Recognit.5
2009 An Integrated Approach to Locality-Conscious Processor Allocation and Scheduling of Mixed-Parallel Applications
abstract
Complex parallel applications can often be modeled as directed acyclic graphs of coarse-grained application tasks with dependences. These applications exhibit both task and data parallelism, and combining these two (also called mixed parallelism) has been shown to be an effective model for their execution. In this paper, we present an algorithm to compute the appropriate mix of task and data parallelism required to minimize the parallel completion time (makespan) of these applications. In other words, our algorithm determines the set of tasks that should be run concurrently and the number of processors to be allocated to each task. The processor allocation and scheduling decisions are made in an integrated manner and are based on several factors such as the structure of the task graph, the runtime estimates and scalability characteristics of the tasks, and the intertask data communication volumes. A locality-conscious scheduling strategy is used to improve intertask data reuse. Evaluation through simulations and actual executions of task graphs derived from real applications and synthetic graphs shows that our algorithm consistently generates schedules with a lower makespan as compared to Critical Path Reduction (CPR) and Critical Path and Allocation (CPA), two previously proposed scheduling algorithms. Our algorithm also produces schedules that have a lower makespan than pure task- and data-parallel schedules. For task graphs with known optimal schedules or lower bounds on the makespan, our algorithm generates schedules that are closer to the optima than other scheduling approaches.
Nagavijayalakshmi Vydyanathan, Sriram Krishnamoorthy, Gerald Sabin, Ümit V. Çatalyürek, Tahsin M. Kurç, P. Sadayappan, Joel H. Saltz
IEEE Trans. Parallel Distributed Syst.7
2008 Multi-hop path splitting and multi-pathing optimizations for data transfers over shared wide-area networks using gridFTP
abstract
In this paper, we propose to employ two optimizations - multi-hop path splitting and multi-pathing - to improve the performance of data transfers over shared public networks. We present a path determination algorithm which integrates the aforesaid optimizations in order to improve the performance of single file transfers. Finally, we develop a file transfer scheduling algorithm based on this framework, and evaluate its effectiveness on a wide-area testbed.
Gaurav Khanna 0002, Ümit V. Çatalyürek, Tahsin M. Kurç, P. Sadayappan, Joel H. Saltz, Rajkumar Kettimuthu, Ian T. Foster
HPDC5
2008 Texture classification using nonlinear color quantization: Application to histopathological image analysis
abstract
In this paper, a novel color texture classification approach is introduced and applied to computer-assisted grading of follicular lymphoma from whole-slide tissue samples. The digitized tissue samples of follicular lymphoma were classified into histological grades under a statistical framework. The proposed method classifies the image either into low or high grades based on the amount of cytological components. To further discriminate the lower grades into low and mid grades, we proposed a novel color texture analysis approach. This approach modifies the gray level co-occurrence matrix method by using a non-linear color quantization with self-organizing feature maps (SOFMs). This is particularly useful for the analysis of H&E stained pathological images whose dynamic color range is considerably limited. Experimental results on real follicular lymphoma images demonstrate that the proposed approach outperforms the gray level based texture analysis.
Olcay Sertel, Jun Kong 0002, Gerard Lozanski, Arwa Shanaah, Ümit V. Çatalyürek, Joel H. Saltz, Metin Nafi Gürcan
ICASSP6
2008 A Duplication Based Algorithm for Optimizing Latency Under Throughput Constraints for Streaming Workflows
abstract
Scheduling, in many application domains, involves the optimization of multiple performance metrics. For example, application workflows with real-time constraints have strict throughput requirements and also desire a low latency or response time. In this paper, we present a novel algorithm for the scheduling of workflows that act on a stream of input data. Our algorithm focuses on the two performance metrics: latency and throughput, and minimizes the latency of workflows while satisfying strict throughput requirements. We leverage pipelined, task and data parallelism in a coordinated manner to meet these objectives and investigate the benefit of task duplication in alleviating communication overheads in the pipelined schedule for different workflow characteristics. The proposed algorithm is designed for a realistic k-port communication model, where each processor can simultaneously communicate with at most k distinct processors. Evaluation using synthetic and application benchmarks shows that our algorithm consistently produces lower-latency schedules and meets throughput requirements, even when previously proposed schemes fail.
Nagavijayalakshmi Vydyanathan, Ümit V. Çatalyürek, Tahsin M. Kurç, P. Sadayappan, Joel H. Saltz
ICPP5
2008 A dynamic scheduling approach for coordinated wide-area data transfers using GridFTP
abstract
Many scientific applications need to stage large volumes of files from one set of machines to another set of machines in a wide-area network. Efficient execution of such data transfers needs to take into account the heterogeneous nature of the environment and dynamic availability of shared resources. This paper proposes an algorithm that dynamically schedules a batch of data transfer requests with the goal of minimizing the overall transfer time. The proposed algorithm performs simultaneous transfer of chunks of files from multiple file replicas, if the replicas exist. Adaptive replica selection is employed to transfer different chunks of the same file by taking dynamically changing network bandwidths into account. We utilize GridFTP as the underlying mechanism for data transfers. The algorithm makes use of information from past GridFTP transfers to estimate network bandwidths and resource availability. The efficiency of the algorithm is evaluated on a wide-area testbed.
Gaurav Khanna 0002, Ümit V. Çatalyürek, Tahsin M. Kurç, Rajkumar Kettimuthu, P. Sadayappan, Joel H. Saltz
IPDPS6
2008 Designing and parameterizing a workflow for optimization: A case study in biomedical imaging
abstract
This paper describes our experience to date employing the systematic mapping and optimization of large- scale scientific application workflows to current and future parallel platforms. The overall goal of the project is to integrate a set of system layers - application program, compiler, run-time environment, knowledge representation, optimization framework, and workflow manager - and through a systematic strategy for workflow mapping, our approach will exploit the vast machine resources available in such parallel platforms to dramatically increase the productivity of application programmers. In this paper, we describe the representation of a biomedical imaging application as a workflow, our early experiences in integrating the set of tools brought together for this project, and implications for future applications.
Vijay S. Kumar, Mary W. Hall, Jihie Kim, Yolanda Gil, Tahsin M. Kurç, Ewa Deelman, Varun Ratnakar, Joel H. Saltz
IPDPS8
2008 Translational research design templates, Grid computing, and HPC
abstract
Design templates that involve discovery, analysis, and integration of information resources commonly occur in many scientific research projects. In this paper we present examples of design templates from the biomedical translational research domain and discuss the requirements imposed on Grid middleware infrastructures by them. Using caGrid, which is a Grid middleware system based on the model driven architecture (MDA) and the service oriented architecture (SOA) paradigms, as a starting point, we discuss architecture directions for MDA and SOA based systems like caGrid to support common design templates.
Joel H. Saltz, Scott Oster, Shannon Hastings, Stephen Langella, Renato Ferreira 0001, Justin Permar, Ashish Sharma 0001, David Ervin, Tony Pan, Ümit V. Çatalyürek, Tahsin M. Kurç
IPDPS1
2008 Using overlays for efficient data transfer over shared wide-area networks
abstract
Data-intensive applications frequently transfer large amounts of data over wide-area networks. The performance achieved in such settings can often be improved by routing data via intermediate nodes chosen to increase aggregate bandwidth. We explore the benefits of overlay network approaches by designing and implementing a service-oriented architecture that incorporates two key optimizations - multi-hop path splitting andmulti-pathing - within the GridFTP file transfer protocol. We develop a file transfer scheduling algorithm that incorporates the two optimizations in conjunction with the use of available file replicas. The algorithm makes use of information from past GridFTP transfers to estimate network bandwidths and resource availability. The effectiveness of these optimizations is evaluated using several application file transfer patterns: one-to-all broadcast, all-to-one gather, and data redistribution, on a wide-area testbed. The experimental results show that our architecture and algorithm achieve significant performance improvement.
Gaurav Khanna 0002, Ümit V. Çatalyürek, Tahsin M. Kurç, Rajkumar Kettimuthu, P. Sadayappan, Ian T. Foster, Joel H. Saltz
SC7
2008 Model Formulation: Sharing Data and Analytical Resources Securely in a Biomedical Research Grid Environment
abstract
OBJECTIVES: To develop a security infrastructure to support controlled and secure access to data and analytical resources in a biomedical research Grid environment, while facilitating resource sharing among collaborators. DESIGN: A Grid security infrastructure, called Grid Authentication and Authorization with Reliably Distributed Services (GAARDS), is developed as a key architecture component of the NCI-funded cancer Biomedical Informatics Grid (caBIG). The GAARDS is designed to support in a distributed environment 1) efficient provisioning and federation of user identities and credentials; 2) group-based access control support with which resource providers can enforce policies based on community accepted groups and local groups; and 3) management of a trust fabric so that policies can be enforced based on required levels of assurance. MEASUREMENTS: GAARDS is implemented as a suite of Grid services and administrative tools. It provides three core services: Dorian for management and federation of user identities, Grid Trust Service for maintaining and provisioning a federated trust fabric within the Grid environment, and Grid Grouper for enforcing authorization policies based on both local and Grid-level groups. RESULTS: The GAARDS infrastructure is available as a stand-alone system and as a component of the caGrid infrastructure. More information about GAARDS can be accessed at http://www.cagrid.org. CONCLUSIONS: GAARDS provides a comprehensive system to address the security challenges associated with environments in which resources may be located at different sites, requests to access the resources may cross institutional boundaries, and user credentials are created, managed, revoked dynamically in a de-centralized manner.
Stephen Langella, Shannon Hastings, Scott Oster, Tony Pan, Ashish Sharma 0001, Justin Permar, David Ervin, Berkant Barla Cambazoglu, Tahsin M. Kurç, Joel H. Saltz
J. Am. Medical Informatics Assoc.10
2008 Model Formulation: caGrid 1.0: An Enterprise Grid Infrastructure for Biomedical Research
abstract
OBJECTIVE: To develop software infrastructure that will provide support for discovery, characterization, integrated access, and management of diverse and disparate collections of information sources, analysis methods, and applications in biomedical research. DESIGN: An enterprise Grid software infrastructure, called caGrid version 1.0 (caGrid 1.0), has been developed as the core Grid architecture of the NCI-sponsored cancer Biomedical Informatics Grid (caBIG) program. It is designed to support a wide range of use cases in basic, translational, and clinical research, including 1) discovery, 2) integrated and large-scale data analysis, and 3) coordinated study. MEASUREMENTS: The caGrid is built as a Grid software infrastructure and leverages Grid computing technologies and the Web Services Resource Framework standards. It provides a set of core services, toolkits for the development and deployment of new community provided services, and application programming interfaces for building client applications. RESULTS: The caGrid 1.0 was released to the caBIG community in December 2006. It is built on open source components and caGrid source code is publicly and freely available under a liberal open source license. The core software, associated tools, and documentation can be downloaded from the following URL: https://cabig.nci.nih.gov/workspaces/Architecture/caGrid. CONCLUSIONS: While caGrid 1.0 is designed to address use cases in cancer research, the requirements associated with discovery, analysis and integration of large scale data, and coordinated studies are common in other biomedical fields. In this respect, caGrid 1.0 is the realization of a framework that can benefit the entire biomedical community.
Scott Oster, Stephen Langella, Shannon Hastings, David Ervin, Ravi K. Madduri, Joshua Phillips, Tahsin M. Kurç, Frank Siebenlist, Peter A. Covitz, Krishnakant Shanbhag, Ian T. Foster, Joel H. Saltz
J. Am. Medical Informatics Assoc.12
2008 An imaging workflow for characterizing phenotypical change in large histological mouse model datasets
Kishore Mosaliganti, Tony Pan, Randall Ridgway, Richard Sharp, Lee A. D. Cooper, Alexandra Gulacy, Ashish Sharma 0001, M. Okan Irfanoglu, Raghu Machiraju, Tahsin M. Kurç, Alain de Bruin, Pamela Wenzel, Gustavo Leone, Joel H. Saltz, Kun Huang 0001
J. Biomed. Informatics14
2008 Large-Scale Biomedical Image Analysis in Grid Environments
abstract
This paper presents the application of a component-based Grid middleware system for processing extremely large images obtained from digital microscopy devices. We have developed parallel, out-of-core techniques for different classes of data processing operations employed on images from confocal microscopy scanners. These techniques are combined into a data preprocessing and analysis pipeline using the component-based middleware system. The experimental results show that: 1) our implementation achieves good performance and can handle very large datasets on high-performance Grid nodes, consisting of computation and/or storage clusters and 2) it can take advantage of Grid nodes connected over high-bandwidth wide-area networks by combining task and data parallelism.
Vijay S. Kumar, Benjamin Rutt, Tahsin M. Kurç, Ümit V. Çatalyürek, Tony Pan, Sunny K. Chow, Stephan Lamont, Maryann E. Martone, Joel H. Saltz
IEEE Trans. Inf. Technol. Biomed.9
2008 Reconstruction of Cellular Biological Structures from Optical Microscopy Data
abstract
Developments in optical microscopy imaging have generated large high-resolution data sets that have spurred medical researchers to conduct investigations into mechanisms of disease, including cancer at cellular and subcellular levels. The work reported here demonstrates that a suitable methodology can be conceived that isolates modality-dependent effects from the larger segmentation task and that 3D reconstructions can be cognizant of shapes as evident in the available 2D planar images. In the current realization, a method based on active geodesic contours is first deployed to counter the ambiguity that exists in separating overlapping cells on the image plane. Later, another segmentation effort based on a variant of Voronoi tessellations improves the delineation of the cell boundaries using a Bayesian formulation. In the next stage, the cells are interpolated across the third dimension thereby mitigating the poor structural correlation that exists in that dimension. We deploy our methods on three separate data sets obtained from light, confocal, and phase-contrast microscopy and validate the results appropriately.
Kishore Mosaliganti, Lee A. D. Cooper, Richard Sharp, Raghu Machiraju, Gustavo Leone, Kun Huang 0001, Joel H. Saltz
IEEE Trans. Vis. Comput. Graph.7
2007 Computerized Pathological Image Analysis For Neuroblastoma Prognosis
Metin Nafi Gürcan, Jun Kong 0002, Olcay Sertel, Berkant Barla Cambazoglu, Joel H. Saltz, Ümit V. Çatalyürek
AMIA5
2007 The Cancer Biomedical Informatics Grid (caBIG™) Security Infrastructure
Stephen Langella, Scott Oster, Shannon Hastings, Frank Siebenlist, Joshua Phillips, David Ervin, Justin Permar, Tahsin M. Kurç, Joel H. Saltz
AMIA9
2007 caGrid 1.0: A Grid Enterprise Architecture for Cancer Research
Scott Oster, Stephen Langella, Shannon Hastings, David Ervin, Ravi K. Madduri, Tahsin M. Kurç, Frank Siebenlist, Peter A. Covitz, Krishnakant Shanbhag, Ian T. Foster, Joel H. Saltz
AMIA11
2007 The GPU on biomedical image processing for color and phenotype analysis
abstract
The computational power and memory bandwidth of graphics processing units (GPUs) have turned them into attractive platforms for general-purpose applications. In this paper, we exploit this power in the context of biomedical image processing by establishing a cooperative environment between the CPU and the GPU. We deal with phenotype and color analysis on a wide variety of microscopic images from studies of cartilage and bone tissue regeneration using stem cells and genetics involving cancer pathology. Both processors are used in parallel to map algorithms for computing color histograms, contour detection using the Canny filter and pattern recognition based on the Hough transform. Task, data and instruction parallelism are exploited in the GPU to accomplish performance gains between 4x and 100x more than the typical CPU code.
Antonio Ruiz 0001, Manuel Ujaldon, Jose Antonio Andrades, Jose Becerra, Kun Huang 0001, Tony Pan, Joel H. Saltz
BIBE7
2007 Pathological Image Analysis Using the GPU: Stroma Classification for Neuroblastoma
abstract
Neuroblastoma is one of the most malignant childhood cancers affecting infants mostly. The current prognosis is based on microscopic examination of slides by expert pathologists, a process that is error-prone, time consuming and may lead to inter- and intra-reader variations. Therefore, we are developing a Computer Aided Prognosis (CAP) system which provides computerized image analysis to assist pathologist in their prognosis. Since this system operates on relatively large- scale images and requires sophisticated algorithms, it takes a long time to process whole-slide images. In this paper, we propose a novel and efficient approach for the execution of a CAP system for neuroblastoma prognosis, using the graphics processing unit (GPU). By leveraging high memory bandwidth and strong floating point operation capabilities of the GPU, our goal is to achieve order of magnitude reduction in the overall execution time as compared to that on a CPU alone. The proposed approach was tested on a set of testing images with a promising accuracy of 99.4% and an execution performance gain factor up to 45 times compared to C++ code running on the CPU.
Antonio Ruiz 0001, Olcay Sertel, Manuel Ujaldon, Ümit V. Çatalyürek, Joel H. Saltz, Metin Nafi Gürcan
BIBM5
2007 An Efficient and Reliable Scientific Workflow System
abstract
This paper presents a fault tolerance framework for applications that process data using a distributed network of user-defined operations in a pipelined fashion. The framework saves intermediate results and messages exchanged among application components in a distributed data management system to facilitate quick recovery from failures. The experimental results show that the framework scales well and our approach introduces very little overhead to application execution.
Tulio Tavares, George Teodoro, Tahsin M. Kurç, Renato Ferreira 0001, Dorgival O. Guedes, Wagner Meira Jr., Ümit V. Çatalyürek, Shannon Hastings, Scott Oster, Stephen Langella, Joel H. Saltz
CCGRID11
2007 Performance vs. accuracy trade-offs for large-scale image analysis applications
abstract
In many data analysis applications, application-level parameters influence the execution time of the data analysis method or program. Some of these parameters also affect the accuracy of output of the analysis. In this work, we investigate execution strategies for adaptive data analysis applications where the user is willing to trade-off accuracy of output for performance gain and vice-versa. In order to meet the user defined quality of service requirements, the system must dynamically select values for the parameters during execution. We propose algorithms for adaptive processing of image tiles at different resolutions so that user defined requirements in terms of accuracy of the result and execution time constraints can be satisfied. We develop heuristics for estimation of accuracy vs performance characteristics of image tiles and for scheduling of the tiles for processing. We implement a demand-driven strategy for parallel execution of these heuristics on a parallel machine. We evaluate our approach for analysis of large images from digitized microscopy scanners.
Vijay S. Kumar, Tahsin M. Kurç, Jun Kong 0002, Ümit V. Çatalyürek, Metin Nafi Gürcan, Joel H. Saltz
CLUSTER6
2007 Scheduling File Transfers for Data-Intensive Jobs on Heterogeneous Clusters
Gaurav Khanna 0002, Ümit V. Çatalyürek, Tahsin M. Kurç, P. Sadayappan, Joel H. Saltz
Euro-Par5
2007 Toward Optimizing Latency Under Throughput Constraints for Application Workflows on Clusters
Nagavijayalakshmi Vydyanathan, Ümit V. Çatalyürek, Tahsin M. Kurç, P. Sadayappan, Joel H. Saltz
Euro-Par5
2007 Computer-Aided Grading of Neuroblastic Differentiation: Multi-Resolution and Multi-Classifier Approach
abstract
In this paper, the development of a computer-aided system for the classification of grade of neuroblastic differentiation is presented. This automated process is carried out within a multi-resolution framework that follows a coarse-to-fine strategy. Additionally, a novel segmentation approach using the Fisher-Rao criterion, embedded in the generic expectation-maximization algorithm, is employed. Multiple decisions from a classifier group are aggregated using a two-step classifier combiner that consists of a majority voting process and a weighted sum rule using priori classifier accuracies. The developed system, when tested on 14,616 image tiles, had the best overall accuracy of 96.89%. Furthermore, multi-resolution scheme combined with automated feature selection process resulted in 34% savings in computational costs on average when compared to a previously developed single-resolution system. Therefore, the performance of this system shows good promise for the computer-aided pathological assessment of the neuroblastic differentiation in clinical practice.
Jun Kong 0002, Olcay Sertel, Hiroyuki Shimada, Kim L. Boyer, Joel H. Saltz, Metin Nafi Gürcan
ICIP (5)5
2007 Intelligent Optimization of Parallel and Distributed Applications
abstract
This paper describes a new project that systematically addresses the enormous complexity of mapping applications to current and future parallel platforms. By integrating the system layers - domain-specific environment, application program, compiler, run-time environment, performance models and simulation, and workflow manager - and through a systematic strategy for application mapping, our approach exploit the vast machine resources available in such parallel platforms to dramatically increase the productivity of application programmers. This project brings together computer scientists in the areas represented by the system layers (i.e., language extensions, compilers, run-time systems, workflows) together with expertise in knowledge representation and machine learning. With expert domain scientists in molecular dynamics (MD) simulation, we are developing our approach in the context of a specific application class which already targets environments consisting of several hundreds of processors. In this way, we gain valuable insight into a generalizable strategy, while simultaneously producing performance benefits for existing and important applications.
Bhupesh Bansal, Ümit V. Çatalyürek, Jacqueline Chame, Chun Chen 0002, Ewa Deelman, Yolanda Gil, Mary W. Hall, Vijay S. Kumar, Tahsin M. Kurç, Kristina Lerman, Aiichiro Nakano, Yoon-Ju Lee Nelson, Joel H. Saltz, Ashish Sharma 0001, Priya Vashishta
IPDPS13
2007 Knowledge and Cache Conscious Algorithm Design and Systems Support for Data Mining Algorithms
abstract
The knowledge discovery process is interactive in nature and therefore minimizing query response time is imperative. The compute and memory intensive nature of data mining algorithms makes this task challenging. We propose to improve the performance of data mining algorithms by re-architecting algorithms and designing effective systems support. From the view point of re-architecting algorithms, knowledge-conscious and cache-conscious design strategies are presented. Knowledge-conscious algorithm designs try and re-use repeated computation between iterations and across executions of a data mining algorithm. Cache-conscious algorithm designs on the other hand reduce execution time by maximizing data locality and reuse. The design of systems support that allows a variety of data mining algorithms to leverage knowledge-caching and cache-conscious placement with minimal implementation efforts is also presented.
Amol Ghoting, Gregory Buehrer, Matthew Goyder, Shirish Tatikonda, Srinivasan Parthasarathy 0001, Tahsin M. Kurç, Joel H. Saltz
IPDPS8
2007 Toward terabyte pattern mining: an architecture-conscious solution
abstract
We present a strategy for mining frequent item sets from terabyte-scale data sets on cluster systems. The algorithm embraces the holistic notion of architecture-conscious datamining, taking into account the capabilities of the processor, the memory hierarchy and the available network interconnects. Optimizations have been designed for lowering communication costs using compressed data structures and a succinct encoding. Optimizations for improving cache, memory and I/O utilization using pruningand tiling techniques, and smart data placement strategies are also employed. We leverage the extended memory spaceand computational resources of a distributed message-passing clusterto design a scalable solution, where each node can extend its metastructures beyond main memory by leveraging 64-bit architecture support. Our solution strategy is presented in the context of FPGrowth, a well-studied and rather efficient frequent pattern mining algorithm. Results demonstrate that the proposed strategy result in near-linearscaleup on up to 48 nodes.
Gregory Buehrer, Srinivasan Parthasarathy 0001, Shirish Tatikonda, Tahsin M. Kurç, Joel H. Saltz
PPoPP5
2007 Parallel four-dimensional Haralick texture analysis for disk-resident image datasets
abstract
Abstract Texture analysis is one possible method of detecting features in biomedical images. During texture analysis, texture‐related information is found by examining local variations in image brightness. Four‐dimensional (4D) Haralick texture analysis is a method that extracts local variations along space and time dimensions and represents them as a collection of 14 statistical parameters. However, application of the 4D Haralick method on large time‐dependent image datasets is hindered by data retrieval, computation, and memory requirements. This paper describes a parallel implementation using a distributed component‐based framework of 4D Haralick texture analysis on PC clusters. The experimental performance results show that good performance can be achieved for this application via combined use of task‐ and data‐parallelism. In addition, we show that our 4D texture analysis implementation can be used to classify imaged tissues. Copyright © 2006 John Wiley & Sons, Ltd.
Brent Woods, Bradley D. Clymer, Johannes T. Heverhagen, Michael V. Knopp, Joel H. Saltz, Tahsin M. Kurç
Concurr. Comput. Pract. Exp.5
2007 Introduce: An Open Source Toolkit for Rapid Development of Strongly Typed Grid Services
abstract
Service-oriented architectures and applications have gained wide acceptance in the Grid computing community. A number of tools and middleware systems have been developed to support application development using Grid Services architectures. Most of these efforts, however, have focused on low-level support for management and execution of Grid services, management of Grid-enabled resources, and deployment and execution of applications that make use of Grid services. Simple-to-use service development tools, which would allow a Grid service developer to leverage Grid technologies without needing to know low-level details, are becoming increasingly important for wider application of the Grid. In this paper, we describe an open-source, extensible toolkit, called Introduce, that supports easy development and deployment of Web Services Resource Framework (WSRF) compliant services. Introduce is designed to reduce the service development and deployment effort by hiding low level details of the Globus Toolkit and to enable the implementation of strongly typed services. In strongly typed services, a service produces and consumes data types that are well-defined and published in the Grid. This enables data-level syntactic interoperability so that clients and services can access and consume data elements programmatically and correctly. We expect that enabling strongly typed Grid services while lowering the difficulty of entry to the Grid via toolkits like Introduce will have a major impact to the success of the Grid and its wider adoption as a viable technology of choice in the commercial sector as well as in academic, medical, and government research.
Shannon Hastings, Scott Oster, Stephen Langella, David Ervin, Tahsin M. Kurç, Joel H. Saltz
J. Grid Comput.6
2007 Active semantic caching to optimize multidimensional data analysis in parallel and distributed environments
Henrique Andrade, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
Parallel Comput.4
2007 Detection and Visualization of Surface-Pockets to Enable Phenotyping Studies
abstract
In this paper, we propose a technique for detecting pockets on a surface-of-interest. A sequence of propagating fronts converging to the target surface is used as the basis for inspection. We compute a correspondence function between the initial and the target surface. This leads to a natural definition of the local feature size measured as the evolution distance between mapped points. Surface pockets are then extracted as salient clusters embedded in the feature space. The level-set initialization also determines the scale-space of the extracted pockets. Results are presented on a case-study in which the focus is to chronicle the phenotyping differences in genetically modified mouse placenta. Our results are validated based on manually verified ground-truth.
Kishore Mosaliganti, Firdaus Janoos, Richard Sharp, Randall Ridgway, Raghu Machiraju, Kun Huang 0001, Pamela Wenzel, Alain de Bruin, Gustavo Leone, Joel H. Saltz
IEEE Trans. Medical Imaging10
2006 Prototype Integration of Protein Electrophoresis Laboratory Results In an Information Warehouse to Improve Workflow and Data Analysis
Scott A. Silvey, Michael G. Bissell, Joel H. Saltz, Jyoti Kamal
AMIA4
2006 Dorian: Grid Service Infrastructure for Identity Management and Federation
abstract
Identity management and federation is becoming an ever present problem in large multi-institutional environments. By their nature, Grids span multiple institutional administration boundaries and aim to provide support for the sharing of applications, data, and computational resources in a collaborative environment. One underlying problem is to enable participating institutions to manage the identities of their own members by leveraging existing institutional identity management systems, while at the same time facilitating the participation in larger Grids through the deployment of grid-wide user credentials. Those grid-wide identities are used for features such as single sign-on, secure communication, and are the basis for authorization decisions. In this paper we presented the design and implementation of Dorian, a grid service infrastructure component that enables the federation of users across the collaboration
Stephen Langella, Scott Oster, Shannon Hastings, Frank Siebenlist, Tahsin M. Kurç, Joel H. Saltz
CBMS6
2006 Locality Conscious Processor Allocation and Scheduling for Mixed Parallel Applications
abstract
Complex applications can often be viewed as a collection of coarse-grained data-parallel application components with precedence constraints. It has been shown that combining task and data parallelism (mixed parallelism) can be an effective execution paradigm for these applications. In this paper, we present an algorithm to compute the appropriate mix of task and data parallelism based on the scalability characteristics of the tasks as well as the intertask data communication costs, such that the parallel completion time (makespan) is minimized. The algorithm iteratively reduces the makespan by increasing the degree of data parallelism of tasks on the critical path that have good scalability and a low degree of potential task parallelism. Data communication costs along the critical path are minimized by exploiting parallel transfer mechanisms and use of a locality conscious backfill scheduler. Evaluation using benchmark task graphs derived from real applications as well as synthetic graphs shows that our algorithm consistently performs better than previous scheduling schemes
Nagavijayalakshmi Vydyanathan, Sriram Krishnamoorthy, Gerald Sabin, Ümit V. Çatalyürek, Tahsin M. Kurç, P. Sadayappan, Joel H. Saltz
CLUSTER7
2006 A New Deformable Model for Boundary Tracking in Cardiac MRI and Its Application to the Detection of Intra-Ventricular Dyssynchrony
abstract
We present a new deformable model technique following a snake-like approach and using a complex Fourier shape descriptors parameterization to efficiently formulate the forces that constrain contour deformation. The method was successfully applied to track the left ventricle’s (LV) endocardial and epicardial boundaries in sequences of shortaxis magnetic resonance images depicting complete cardiac cycles. The extracted shapes show the method’s robustness to weak contrast, noisy edge maps and to papillary muscle anatomy. Our second contribution is a statistical pattern recognition approach for the detection of asynchronous activation of the LV walls. We applied our deformable model method to provide spatio-temporal characterizations of complete cardiac cycles and then designed a linear classifier using the popular combination of Principal Component Analysis and Linear Discriminant Analysis. From a database comprising 33 patients, our approach provided a correct classification performance of 90.9% showing its potential in providing improved dyssynchrony characterization as an adjunct to current criteria for selecting patients for therapy, which provided a classification accuracy of just 62.5% on the same database.
Paulo F. U. Gotardo, Kim L. Boyer, Joel H. Saltz, Subha V. Raman
CVPR (1)3
2006 Task Scheduling and File Replication for Data-Intensive Jobs with Batch-shared I/O
abstract
This paper addresses the problem of efficient execution of a batch of data-intensive tasks with batch-shared I/O behavior, on coupled storage and compute clusters. Two scheduling schemes are proposed: 1) a 0-1 integer programming (IP) based approach, which couples task scheduling and data replication, and 2) a bi-level hypergraph partitioning based heuristic approach (BiPartition), which decouples task scheduling and data replication. The experimental results show that: 1) the IP scheme achieves the best batch execution time, but has significant scheduling overhead, thereby restricting its application to small scale workloads, and 2) the BiPartition scheme is a better fit for larger workloads and systems - it has very low scheduling overhead and no more than 5-10% degradation in solution quality, when compared with the IP based approach
Gaurav Khanna 0002, Nagavijayalakshmi Vydyanathan, Ümit V. Çatalyürek, Tahsin M. Kurç, Sriram Krishnamoorthy, P. Sadayappan, Joel H. Saltz
HPDC7
2006 On Creating Efficient Object-relational Views of Scientific Datasets
abstract
Scientific datasets are often large and distributed in flat files across several storage nodes. Scientists frequently want to analyze subsets of these datasets. A data source abstraction that provides an object-relational view of data while hiding the details of storage and transport mechanisms and dataset layouts is useful in this regard. In this abstraction, basic data sources (BDS) interpret flat files as a set of records and are the building blocks of the view mechanism. Derived data sources (DDS) may be built on top of BDSs and provide more complex objects that serve the scientists' needs. The simplest DDS is one that supports a join based view over BDSs. We investigate issues involving building such DDSs for scientific applications and consider distributed versions of the indexed join and the grace hash join algorithms. We construct cost models that capture their performance in a restricted space of dataset and system parameters and compare them analytically and experimentally
Sivaramakrishnan Narayanan, Tahsin M. Kurç, Ümit V. Çatalyürek, Joel H. Saltz
ICPP4
2006 An Integrated Approach for Processor Allocation and Scheduling of Mixed-Parallel Applications
abstract
Computationally complex applications can often be viewed as a collection of coarse-grained data-parallel tasks with precedence constraints. Researchers have shown that combining task and data parallelism (mixed parallelism) can be an effective approach for executing these applications, as compared to pure task or data parallelism. In this paper, we present an approach to determine the appropriate mix of task and data parallelism, i.e., the set of tasks that should be run concurrently and the number of processors to be allocated to each task. An iterative algorithm is proposed that couples processor allocation and scheduling of mixed-parallel applications on compute clusters so as to minimize the parallel completion time (makespan). Our algorithm iteratively reduces the makespan by increasing the degree of data parallelism of tasks on the critical path that have good scalability and a low degree of potential task parallelism. The approach employs a look-ahead technique to escape local minima and uses priority based backfill scheduling to efficiently schedule the parallel tasks onto processors. Evaluation using benchmark task graphs derived from real applications as well as synthetic graphs shows that our algorithm consistently performs better than CPR and CPA, two previously proposed scheduling schemes, as well as pure task and data parallelism
Nagavijayalakshmi Vydyanathan, Sriram Krishnamoorthy, Gerald Sabin, Ümit V. Çatalyürek, Tahsin M. Kurç, P. Sadayappan, Joel H. Saltz
ICPP7
2006 Using Space and Attribute Partitioned Partial Replicas for Data Subsetting and Aggregation Queries
abstract
Partial replication is one type of optimization to speed up execution of queries submitted to large datasets. In partial replication, a portion of the dataset is extracted, re-organized, and re-distributed across the storage system. In this paper we investigate methods for efficient execution of queries when replicas of a dataset exist; we assume the replicas have already been created and do not target the replica creation problem. We propose a cost model and algorithm for combined use of space partitioned and attribute partitioned replicas for executing data subsetting range queries. We extend the cost model and propose a greedy algorithm to address range queries with aggregation operations. The extended replica selection algorithm allows uneven partitioning of replicas across storage nodes. Different replicas can be partitioned across different subsets of storage nodes. We have implemented these techniques as part of an automatic data virtualization system and have evaluated the benefits of our techniques using this system. We demonstrate the efficacy of the algorithms on parallel machines using queries on datasets from oil reservoir simulation studies and satellite data processing applications
Li Weng, Ümit V. Çatalyürek, Tahsin M. Kurç, Gagan Agrawal, Joel H. Saltz
ICPP5
2006 I/O conscious algorithm design and systems support for data analysis on emerging architectures
abstract
Advances in data collection and storage technologies have given rise to large dynamic data stores. In order to effectively manage and mine such stores on modern and emerging architectures, one must consider both designing effective middleware support and re-architecting algorithms, to derive performance that commensurates with technological advances. In this article, we present a top-down view of how one can achieve this goal for next generation data analysis centers. Specifically, we present a case study on frequent pattern algorithms, and show how such algorithms can be re-structured to be cache, memory and I/O conscious. Furthermore, motivated by such algorithms, we present a services oriented middleware framework for the derivation of high performance on next generation architectures
Gregory Buehrer, Amol Ghoting, Shirish Tatikonda, Srinivasan Parthasarathy 0001, Tahsin M. Kurç, Joel H. Saltz
IPDPS7
2006 Design and analysis of a multi-dimensional data sampling service for large scale data analysis applications
abstract
Sampling is a widely used technique to increase efficiency in database and data mining applications operating on large dataset. In this paper, we present a scalable sampling implementation that supports efficient, multi-dimensional spatio-temporal sample generation on dynamic, large scale datasets stored on a storage cluster The proposed algorithm leverages Hilbert space-filling curves in order to provide an approximate linear order of multidimensional data while maintaining spatial locality. This new implementation is then bootstrapped on top of our previous implementation, which efficiently samples large datasets along a single dimension (e.g., time), thereby realizing a service for spatio-temporal sampling. We evaluate the performance of our approach comparing it to the popular R-tree based technique. The experimental results show that our approach achieves up to an order of magnitude higher efficiency and scalability
Tahsin M. Kurç, Joel H. Saltz, Srinivasan Parthasarathy 0001
IPDPS3
2006 A Data Locality Aware Online Scheduling Approach for I/O-Intensive Jobs with File Sharing
Gaurav Khanna 0002, Ümit V. Çatalyürek, Tahsin M. Kurç, P. Sadayappan, Joel H. Saltz
JSSPP5
2006 Preventing Signal Degradation During Elastic Matching of Noisy DCE-MR Eye Images
Kishore Mosaliganti, Guang Jia, Johannes T. Heverhagen, Raghu Machiraju, Joel H. Saltz, Michael V. Knopp
MICCAI (1)5
2006 A Run-time System for Efficient Execution of Scientific Workflows on Distributed Environments
abstract
Scientific workflow systems have been introduced in response to the demand of researchers from several domains of science who need to process and analyze increasingly larger datasets. The design of these systems is largely based on the observation that data analysis applications can be composed as pipelines or networks of computations on data. In this paper we present a run-time support system that is designed to facilitate this type of computation in distributed computing environments. Our system is optimized for data-intensive workflows, in which efficient management and retrieval of data, coordination of data processing and data movement, and check-pointing of intermediate results are critical and challenging issues. Experimental evaluation of our system shows that linear speedups can be achieved for sophisticated applications, which are implemented as a network of multiple data processing components
George Teodoro, Tulio Tavares, Renato Ferreira 0001, Tahsin M. Kurç, Wagner Meira Jr., Dorgival O. Guedes, Tony Pan, Joel H. Saltz
SBAC-PAD8
2006 Imaging and visual analysis - Large image correction and warping in a cluster environment
abstract
This paper is concerned with efficient execution of a pipeline of data processing operations on very large images obtained from confocal microscopy instruments. We describe parallel, out-of-core algorithms for each operation in this pipeline. One of the challenging steps in the pipeline is the warping operation using inverse mapping based methods. We propose and investigate a set of algorithms to handle the warping computations on storage clusters. Our experimental results show that the proposed approaches are scalable both in terms of number of processors and the size of images.
Vijay S. Kumar, Benjamin Rutt, Tahsin M. Kurç, Ümit V. Çatalyürek, Joel H. Saltz, Sunny K. Chow, Stephan Lamont, Maryann E. Martone
SC5
2006 caGrid: design and implementation of the core architecture of the cancer biomedical informatics grid
abstract
MOTIVATION: The complexity of cancer is prompting researchers to find new ways to synthesize information from diverse data sources and to carry out coordinated research efforts that span multiple institutions. There is a need for standard applications, common data models, and software infrastructure to enable more efficient access to and sharing of distributed computational resources in cancer research. To address this need the National Cancer Institute (NCI) has initiated a national-scale effort, called the cancer Biomedical Informatics Grid (caBIGtrade mark), to develop a federation of interoperable research information systems. RESULTS: At the heart of the caBIG approach to federated interoperability effort is a Grid middleware infrastructure, called caGrid. In this paper we describe the caGrid framework and its current implementation, caGrid version 0.5. caGrid is a model-driven and service-oriented architecture that synthesizes and extends a number of technologies to provide a standardized framework for the advertising, discovery, and invocation of data and analytical resources. We expect caGrid to greatly facilitate the launch and ongoing management of coordinated cancer research studies involving multiple institutions, to provide the ability to manage and securely share information and analytic resources, and to spur a new generation of research applications that empower researchers to take a more integrative, trans-domain approach to data mining and analysis. AVAILABILITY: The caGrid version 0.5 release can be downloaded from https://cabig.nci.nih.gov/workspaces/Architecture/caGrid/. The operational test bed Grid can be accessed through the client included in the release, or through the caGrid-browser web application http://cagrid-browser.nci.nih.gov.
Joel H. Saltz, Scott Oster, Shannon Hastings, Stephen Langella, Tahsin M. Kurç, William Sanchez, Manav Kher, Arumani Manisundaram, Krishnakant Shanbhag, Peter A. Covitz
Bioinform.1
2006 Application of Information Technology: An XML-based System for Synthesis of Data from Disparate Databases
abstract
Diverse data sets have become key building blocks of translational biomedical research. Data types captured and referenced by sophisticated research studies include high throughput genomic and proteomic data, laboratory data, data from imagery, and outcome data. In this paper, the authors present the application of an XML-based data management system to support integration of data from disparate data sources and large data sets. This system facilitates management of XML schemas and on-demand creation and management of XML databases that conform to these schemas. They illustrate the use of this system in an application for genotype-phenotype correlation analyses. This application implements a method of phenotype-genotype correlation based on phylogenetic optimization of large data sets of mouse SNPs and phenotypic data. The application workflow requires the management and integration of genomic information and phenotypic data from external data repositories and from the results of phenotype-genotype correlation analyses. Our implementation supports the process of carrying out a complex workflow that includes large-scale phylogenetic tree optimizations and application of Maddison's concentrated changes test to large phylogenetic tree data sets. The data management system also allows collaborators to share data in a uniform way and supports complex queries that target data sets.
Tahsin M. Kurç, Daniel Janies, Andrew D. Johnson, Stephen Langella, Scott Oster, Shannon Hastings, Farhat Habib, Terry Camerlengo, David Ervin, Ümit V. Çatalyürek, Joel H. Saltz
J. Am. Medical Informatics Assoc.11
2005 The GPU on irregular computing: performance issues and contributions
abstract
The paper describes a set of strategies for mapping irregular codes onto commodity graphics hardware. We start identifying the resources that current GPUs contain for solving indirect array accesses entirely on hardware, like vertices, textures and color tables. We then show how multiple indirections can be mapped onto the graphics pipeline, basically taking advantage of its streaming architecture for sequencing the indirections through subsequent pipeline stages. Our techniques are applied over typical irregular kernels like the sparse matrix-vector multiply and the Euler solver. Execution times on the GeForce Series consistently outperform the Pentium 4 and Athlon 64 processors, with performance depending on floating-point precision. 1.
Manuel Ujaldon, Joel H. Saltz
CAD/Graphics2
2005 A hypergraph partitioning based approach for scheduling of tasks with batch-shared I/O
abstract
This paper proposes a novel, hypergraph partitioning based strategy to schedule multiple data analysis tasks with batch-shared I/O behavior. This strategy formulates the sharing of files among tasks as a hypergraph to minimize the I/O overheads due to transferring of the same set of files multiple times and employs a dynamic scheme for file transfers to reduce contention on the storage system. We experimentally evaluate the proposed approach using application emulators from two application domains; analysis of remotely-sensed data and biomedical imaging.
Gaurav Khanna 0002, Nagavijayalakshmi Vydyanathan, Tahsin M. Kurç, Ümit V. Çatalyürek, Pete Wyckoff, Joel H. Saltz, P. Sadayappan
CCGRID6
2005 Servicing range queries on multidimensional datasets with partial replicas
abstract
Partial replication is one type of optimization to speed up execution of queries submitted to large datasets. In partial replication, a portion of the dataset is extracted, re-organized, and re-distributed across the storage system. The objective is to reduce the volume of I/O and increase I/O parallelism for different types of queries and for the portions of the dataset that are likely to be accessed frequently. When multiple partial replicas of a dataset exist, query execution plan should be generated so as to use the best combination of subsets of partial replicas (and possibly the original dataset) to minimize query execution time. In this paper, we present a compiler and runtime approach for range queries submitted against distributed scientific datasets. A heuristic algorithm is proposed to choose the set of replicas to reduce query execution. We show the efficiency of the proposed method using datasets and queries in oil reservoir simulation studies on a cluster machine.
Li Weng, Ümit V. Çatalyürek, Tahsin M. Kurç, Gagan Agrawal, Joel H. Saltz
CCGRID5
2005 Distributed Out-of-Core Preprocessing of Very Large Microscopy Images for Efficient Querying
abstract
We present a combined task- and data-parallel approach for distributed execution of pre-processing operations to support efficient evaluation of polygonal aggregation queries on digitized microscopy images. Our approach targets out-of-core, pipelined processing of very large images on active storage clusters. Our experimental results show that the proposed approach is scalable both in terms of number of processors and the size of images
Benjamin Rutt, Vijay S. Kumar, Tony Pan, Tahsin M. Kurç, Ümit V. Çatalyürek, Joel H. Saltz
CLUSTER6
2005 Design of a next generation sampling service for large scale data analysis applications
abstract
Advances in data collection and storage technologies have resulted in large and dynamically growing data sets at many organizations. Database and data mining researchers often use sampling with great effect to scale up performance on these data sets with small cost to accuracy. However, existing techniques often ignore the cost of computing a sample. This cost is often linear in the size of the data set, not the sample, which is expensive. Furthermore, for data mining applications that leverage progressive sampling or bootstrapping-based techniques, this cost can be prohibitive, since they require the generation of multiple samples.To address this problem, we present a solution in the context of a state-of-the-art data analysis center. Specifically, we propose a scalable service that supports sample generation with cost linear in the size of the sample. We then present an efficient parallelization of this service. Our solution leverages high speed interconnects (e.g. Myrinet, Infini-band) for parallel I/O operations with pipelined data transfers. We export an interface that supports both ad-hoc SQL-like querying for database applications, as well as a stand-alone service for data mining applications. We then evaluate our work using queries abstracted from a network monitoring and analysis application, which uses both database and progressive sampling queries. We demonstrate that our implementation achieves good load balance and realizes up to an order of magnitude speedup when compared with extant approaches.
Huai Wang, Srinivasan Parthasarathy 0001, Amol Ghoting, Shirish Tatikonda, Gregory Buehrer, Tahsin M. Kurç, Joel H. Saltz
ICS7
2005 A simulation and data analysis system for large-scale, data-driven oil reservoir simulation studies
abstract
Abstract The main goal of oil reservoir management is to provide more efficient, cost‐effective and environmentally safer production of oil from reservoirs. Numerical simulations can aid in the design and implementation of optimal production strategies. However, traditional simulation‐based approaches to optimizing reservoir management are rapidly overwhelmed by data volume when large numbers of realizations are sought using detailed geologic descriptions. In this paper, we describe a software architecture to facilitate large‐scale simulation studies, involving ensembles of long‐running simulations and analysis of vast volumes of output data. Copyright © 2005 John Wiley & Sons, Ltd.
Tahsin M. Kurç, Ümit V. Çatalyürek, Joel H. Saltz, Ryan Martino, Mary F. Wheeler, Malgorzata Peszynska, Alan Sussman, Christian Hansen 0002, Mrinal K. Sen, Roustam Seifoullaev, Paul L. Stoffa, Carlos Torres-Verdín, Manish Parashar
Concurr. Pract. Exp.4
2005 Application of Grid-enabled technologies for solving optimization problems in data-driven reservoir studies
Manish Parashar, Hector Klie, Ümit V. Çatalyürek, Tahsin M. Kurç, Wolfgang Bangerth, Vincent Matossian, Joel H. Saltz, Mary F. Wheeler
Future Gener. Comput. Syst.7
2005 Application of Information Technology: A Grid-Based Image Archival and Analysis System
abstract
Here the authors present a Grid-aware middleware system, called GridPACS, that enables management and analysis of images in a massive scale, leveraging distributed software components coupled with interconnected computation and storage platforms. The need for this infrastructure is driven by the increasing biomedical role played by complex datasets obtained through a variety of imaging modalities. The GridPACS architecture is designed to support a wide range of biomedical applications encountered in basic and clinical research, which make use of large collections of images. Imaging data yield a wealth of metabolic and anatomic information from macroscopic (e.g., radiology) to microscopic (e.g., digitized slides) scale. Whereas this information can significantly improve understanding of disease pathophysiology as well as the noninvasive diagnosis of disease in patients, the need to process, analyze, and store large amounts of image data presents a great challenge.
Shannon Hastings, Scott Oster, Stephen Langella, Tahsin M. Kurç, Tony Pan, Ümit V. Çatalyürek, Joel H. Saltz
J. Am. Medical Informatics Assoc.7
2004 Serving queries to multi-resolution datasets on disk-based storage clusters
abstract
This paper is concerned with efficient querying of very large multi-resolution datasets on storage and compute clusters. We present a suite of services that support storage, indexing, and data processing (data sampling and data aggregation) on datasets that consist of a collection of multi-resolution Grids. We empirically evaluate the performance impact of different data declustering, indexing, and query processing strategies. The experimental evaluation is carried out using a data server implemented to serve multi-terabyte multi-resolution volumetric datasets to remote visualization clients and a one-terabyte multi-resolution volumetric dataset on a PC cluster with distributed disk space.
Tony Pan, Ümit V. Çatalyürek, Tahsin M. Kurç, Joel H. Saltz
CCGRID5
2004 A distributed data management middleware for data-driven application systems
abstract
A key challenge in supporting data-driven scientific applications is the storage and management of input and output data in a distributed environment. We describe a distributed storage middleware, based on a data and metadata management framework, to address this problem. In this middleware system, applications define the structure of their input and output data using XML schemas. The system provides support for 1) registration, versioning, management of schemas, and 2) management of storage, querying, and retrieval of instance data corresponding to the schemas in distributed databases. We carry out an experimental evaluation of the system on a set of PC clusters connected over wide- (WANs) and local-area networks (LANs).
Stephen Langella, Shannon Hastings, Scott Oster, Tahsin M. Kurç, Ümit V. Çatalyürek, Joel H. Saltz
CLUSTER6
2004 An Approach for Automatic Data Virtualization
Li Weng, Gagan Agrawal, Ümit V. Çatalyürek, Tahsin M. Kurç, Sivaramakrishnan Narayanan, Joel H. Saltz
HPDC6
2004 Strategies for Using Additional Resources in Parallel Hash-Based Join Algorithms
Tahsin M. Kurç, Tony Pan, Ümit V. Çatalyürek, Sivaramakrishnan Narayanan, Pete Wyckoff, Joel H. Saltz
HPDC7
2004 A Parallel Implementation of 4-Dimensional Haralick Texture Analysis for Disk-Resident Image Datasets
abstract
Texture analysis is one possible method to detect features in biomedical images. During texture analysis, texture related information is found by examining local variations in image brightness. 4-dimensional (4D) Haralick texture analysis is a method that extracts local variations along space and time dimensions and represents them as a collection of fourteen statistical parameters. However, the application of the 4D Haralick method on large time-dependent 2D and 3D image datasets is hindered by computation and memory requirements. This paper presents a parallel implementation of 4D Haralick texture analysis on PC clusters. We present a performance evaluation of our implementation on a cluster of PCs. Our results show that good performance can be achieved for this application via combined use of task- and data-parallelism.
Brent Woods, Bradley D. Clymer, Joel H. Saltz, Tahsin M. Kurç
SC3
2004 Optimizing the Execution of Multiple Data Analysis Queries on Parallel and Distributed Environments
abstract
We investigate techniques for efficiently executing multiquery workloads from data and computation-intensive applications in parallel and/or distributed computing environments. In this context, we describe a database optimization framework that supports data and computation reuse, query scheduling, and active semantic caching to speed up the evaluation of multiquery workloads. Its most striking feature is the ability of optimizing the execution of queries in the presence of application-specific constructs by employing a customizable data and computation reuse model. Furthermore, we discuss how the proposed optimization model is flexible enough to work efficiently irrespective of the parallel/distributed environment underneath. In order to evaluate the proposed optimization techniques, we present experimental evidence using real data analysis applications. For this purpose, a common implementation for the queries under study was provided according to the database optimization framework and deployed on top of three distinct experimental configurations: a shared memory multiprocessor, a cluster of workstations, and a distributed computational Grid-like environment.
Henrique Andrade, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
IEEE Trans. Parallel Distributed Syst.4
2003 Pharmacokinetic Mapping of Breast Tumors: a New Statistical Analysis Technique for Dynamic magnetic resonance Imaging
Catalin C. Barbacioru, Anand Arunachalam 0002, Daniel J. Cowden, Eiad B. Kahwash, Joel H. Saltz
AMIA5
2003 Order Sets Utilization in a Clinical Order Entry System
Daniel J. Cowden, Catalin C. Barbacioru, Eiad B. Kahwash, Joel H. Saltz
AMIA4
2003 Proteomic Pattern Analysis, a New Era of Screening Cancers
Eiad B. Kahwash, Catalin C. Barbacioru, Daniel J. Cowden, Joel H. Saltz
AMIA4
2003 Information Warehouse as a Tool to Analyze Computerized Physician Order Entry Order Set Utilization: Opportunities for Improvement
Jyoti Kamal, Patrick Rogers, Joel H. Saltz, Hagop S. Mekhjian
AMIA3
2003 Impact of CPOE Order Sets on Lab Orders
Hagop S. Mekhjian, Joel H. Saltz, Patrick Rogers, Jyoti Kamal
AMIA2
2003 Image Processing or the Grid: A Toolkit or Building Grid-enabled Image Processing Applications
abstract
Analyzing large and distributed image datasets is a crucial step in understanding the structural and functional characteristics of biological systems. In this paper, we present the design and implementation of a toolkit that allows rapid and efficient development of biomedical image analysis applications in a distributed environment. This toolkit employs the Insight Segmentation and Registration Toolkit (ITK) and Visualization Toolkit (VTK) layered on a component-based framework. We present experimental results on a cluster of workstations.
Shannon Hastings, Tahsin M. Kurç, Stephen Langella, Ümit V. Çatalyürek, Tony Pan, Joel H. Saltz
CCGRID6
2003 A Slacker Coherence rotocol for Pull-based Monitoring of On-line Data Source
abstract
An increasing number of online applications operate on data from disparate, and often wide-spread, data sources. This paper studies the design of a system for the automated monitoring of on-line data sources. In this system a number of ad-hoc data warehouses, which maintain client-specified views, are interposed between clients and data sources. We present a model of coherence, referred to here as slacker coherence, to address the freshness problem in the context of pull-based protocols. We experimentally examine various techniques for estimating update rates and polling adaptively. We also look at the impact on the coherence model performance of the request scheduling algorithm at the source.
Radhakrishnan Sundaresan, Tahsin M. Kurç, Mario Lauria, Srinivasan Parthasarathy 0001, Joel H. Saltz
CCGRID5
2003 Impact of High Performance Sockets on Data Intensive Applications
abstract
The challenging issues in supporting data intensive applications on clusters include efficient movement of large volumes of data between processor memories and efficient coordination of data movement and processing by a runtime support to achieve high performance. Such applications have several requirements such as guarantees in performance, scalability with these guarantees and adaptability to heterogeneous environments. With the advent of user-level protocols like the Virtual Interface Architecture (VIA) and the modern InfiniBand Architecture, the latency and bandwidth experienced by applications has approached to that of the physical network on clusters. In order to enable applications written on top of TCP/IP to take advantage of the high performance of these user-level protocols, researchers have come up with a number of techniques including User Level Sockets Layers over high performance protocols. In this paper, we study the performance and limitations of such substrate, referred to here as SocketVIA, using a component framework designed to provide runtime support for data intensive applications. The experimental results show that by reorganizing certain components of an application (in our case, the partitioning of a dataset into smaller data chunks), we can make significant improvements in application performance. This leads to a higher scalability of applications with performance guarantees. It also allows fine grained load balancing, hence making applications more adaptable to heterogeneity in resource availability. The experimental results also show that the different performance characteristics of SocketVIA allow a more efficient partitioning of data at the source nodes, thus improving the performance of the application up to an order of magnitude in some cases.
Pavan Balaji, Jiesheng Wu, Tahsin M. Kurç, Ümit V. Çatalyürek, Dhabaleswar K. Panda 0001, Joel H. Saltz
HPDC6
2003 Adaptive Polling of Grid Resource Monitors Using a Slacker Coherence Model
abstract
As data and computational grids grow in size and complexity, the crucial task of identifying, monitoring and utilizing available resources in an efficient manner is becoming increasingly difficult. The design of monitoring systems that are scalable both in the number of sources being monitored and in the number of clients served is a challenging issue. In this paper we investigate the trade-offs of different polling strategies that can be used to monitor resource availability on machines in a distributed environment. We show how adaptive polling protocols can substantially increase scalability with a less than proportional loss of precision, and how these protocols can be personalized for different types of resource usage patterns.
Radhakrishnan Sundaresan, Mario Lauria, Tahsin M. Kurç, Srinivasan Parthasarathy 0001, Joel H. Saltz
HPDC5
2003 Optimizing Reduction Computations In a Distributed Environment
abstract
We investigate runtime strategies for data-intensive applications that invovle generalized reductions on large, distributed datasets.Our set of strategies includes replicated filter state, partitioned filter state, and hybrid options between these two extremes.We evaluate these strategies using emulators of three real applications, different query and output sizes, and a number of configurations.We consider execution in a homogeneous cluster and in a distributed environment where only a subset of nodes hst the data.Our results show replicating the filter state scales well and outperforms other schemes, if sufficient memory is available and sufficient computation is involved to offset the cost of global merge step.In other cases, hybrid is usually the best.Moreover, in almost all cases, the performance of the hybrid strategy is quite close to the best strategy. Thus, we believe that hybrid is an attractive approach when the relative performance of different schemes cannot be predicted.
Tahsin M. Kurç, Feng Lee, Gagan Agrawal, Ümit V. Çatalyürek, Renato Ferreira 0001, Joel H. Saltz
SC6
2003 Identifying parallelism in programs with cyclic graphs
Yuan-Shin Hwang, Joel H. Saltz
J. Parallel Distributed Comput.2
2003 The virtual microscope
abstract
We present the design and implementation of the Virtual Microscope, a software system employing a client/server architecture to provide a realistic emulation of a high power light microscope. The system provides a form of completely digital telepathology, allowing simultaneous access to archived digital slide images by multiple clients. The main problem the system targets is storing and processing the extremely large quantities of data required to represent a collection of slides. The Virtual Microscope client software runs on the end user's PC or workstation, while database software for storing, retrieving and processing the microscope image data runs on a parallel computer or on a set of workstations at one or more potentially remote sites. We have designed and implemented two versions of the data server software. One implementation is a customization of a database system framework that is optimized for a tightly coupled parallel machine with attached local disks. The second implementation is component-based, and has been designed to accommodate access to and processing of data in a distributed, heterogeneous environment. We also have developed caching client software, implemented in Java, to achieve good response time and portability across different computer platforms. The performance results presented show that the Virtual Microscope systems scales well, so that many clients can be adequately serviced by an appropriately configured data server.
Ümit V. Çatalyürek, Michael D. Beynon, Chialin Chang, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
IEEE Trans. Inf. Technol. Biomed.6
2002 Multiple Query Optimization for Data Analysis Applications on Clusters of SMPs
abstract
This paper is concerned with the efficient execution of multiple query workloads on a cluster of SMPs. We target applications that access and manipulate large scientific datasets. Queries in these applications involve user-defined processing operations and distributed data structures to hold intermediate and final results. Our goal is to implement system components to leverage previously computed query results and to effectively utilize processing power and aggregated I/O bandwidth on SMP nodes so that both single queries and multi-query batches can be efficiently executed.
Henrique Andrade, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
CCGRID4
2002 Low-Cost Non-Intrusive Debugging Strategies for Distributed Parallel Programs
abstract
We show how five low-cost and nonintrusive techniques that work using free commodity tools such as GDB can be used to improve the debugging process of multi-threaded and/or distributed parallel programs. These techniques have been used in the development of major software middleware - DataCutter and MQO - and have proven their value by lowering the time necessary to detect and correct bugs.
Michael D. Beynon, Henrique Andrade, Joel H. Saltz
CLUSTER3
2002 Compiler supported high-level abstractions for sparse disk-resident datasets
abstract
Processing and analyzing large volumes of data plays an increasingly important role in many domains of scientific research. The complexity and irregularity of datasets in many domains make the task of developing such processing applications tedious and error-prone.We propose use of high-level abstractions for hiding the irregularities in these datasets and enabling rapid development of correct data processing applications. We present two execution strategies and a set of compiler analysis techniques for obtaining high performance from applications written using our proposed high-level abstractions. Our execution strategies achieve high locality in disk accesses. Once a disk block is read from the disk, all iterations that access any of the elements from this disk block are performed. To support our execution strategies and improve the performance, we have developed static analysis techniques for: 1) computing the set of iterations that access a particular right-hand-side element, 2) generating a function that can be applied to the meta-data associated with each disk block, for determining if that disk block needs to be read, and 3) performing code hoisting of conditionals.We present experimental results from a prototype compiler implementing our techniques to demonstrate the effectiveness of our approach.
Renato Ferreira 0001, Gagan Agrawal, Joel H. Saltz
ICS3
2002 Active Proxy-G: optimizing the query execution process in the grid
abstract
The Grid environment facilitates collaborative work and allows many users to query and process data over geographically dispersed data repositories. Over the past several years, there has been a growing interest in developing applications that interactively analyze datasets, potentially in a collaborative setting. We describe the Active Proxy-G service that is able to cache query results, use those results for answering new incoming queries, generate subqueries for the parts of a query that cannot be produced from the cache, and submit the subqueries for final processing at application servers that store the raw datasets. We present an experimental evaluation to illustrate the effects of various design tradeoffs. We also show the benefits that two real applications gain from using the middleware.
Henrique Andrade, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
SC4
2002 Executing multiple pipelined data analysis operations in the grid
abstract
Processing of data in many data analysis applications can be represented as an acyclic, coarse grain data flow, from data sources to the client. This paper is concerned with scheduling of multiple data analysis operations, each of which is represented as a pipelined chain of processing on data. We define the scheduling problem for effectively placing components onto Grid resources, and propose two scheduling algorithms. Experimental results are presented using a visualization application.
Matthew Spencer, Renato Ferreira 0001, Michael D. Beynon, Tahsin M. Kurç, Ümit V. Çatalyürek, Alan Sussman, Joel H. Saltz
SC7
2002 Optimizing execution of component-based applications using group instances
Michael D. Beynon, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
Future Gener. Comput. Syst.4
2002 Processing large-scale multi-dimensional data in parallel and distributed environments
Michael D. Beynon, Chialin Chang, Ümit V. Çatalyürek, Tahsin M. Kurç, Alan Sussman, Henrique Andrade, Renato Ferreira 0001, Joel H. Saltz
Parallel Comput.8
2002 Data parallel language and compiler support for data intensive applications
Renato Ferreira 0001, Gagan Agrawal, Joel H. Saltz
Parallel Comput.3
2001 Optimizing Execution of Component-based Applications using Group Instances
abstract
Research on programming models for developing applications in the Grid has proposed component-based models as a viable approach, in which an application is composed of multiple interacting computational objects. We have been developing a framework, called filter-stream programming, for building data-intensive applications that query, analyze and manipulate very large data sets in a distributed environment. In this model, the processing structure of an application is represented as a set of processing units, referred to as filters. We develop the problem of scheduling instances of a filter group. A filter group is a set of filters collectively performing a computation for an application. In particular we seek the answer to the following question: should a new instance be created, or an existing one reused? We experimentally investigate the effects of instantiating multiple filter groups on performance under varying application characteristics.
Michael D. Beynon, Alan Sussman, Tahsin M. Kurç, Joel H. Saltz
CCGRID4
2001 Efficient execution of multiple query workloads in data analysis applications
abstract
Applications that analyze, mine, and visualize large datasets are considered an important class of applications in many areas of science, engineering, and business. Queries commonly executed in data analysis applications often involve user-defined processing of data and application-specific data structures. If data analysis is employed in a collaborative environment, the data server should execute multiple such queries simultaneously to minimize the response time to clients. In this paper we present the design of a runtime system for executing multiple query workloads on a shared-memory machine. We describe experimental results using an application for browsing digitized microscopy images.
Henrique Andrade, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
SC4
2001 Distributed processing of very large datasets with DataCutter
Michael D. Beynon, Tahsin M. Kurç, Ümit V. Çatalyürek, Chialin Chang, Alan Sussman, Joel H. Saltz
Parallel Comput.6
2001 Analysis of the Clustering Properties of the Hilbert Space-Filling Curve
abstract
AbstractÐSeveral schemes for the linear mapping of a multidimensional space have been proposed for various applications, such as access methods for spatio-temporal databases and image compression. In these applications, one of the most desired properties from such linear mappings is clustering, which means the locality between objects in the multidimensional space being preserved in the linear space. It is widely believed that the Hilbert space-filling curve achieves the best clustering [1], [14]. In this paper, we analyze the clustering property of the Hilbert space-filling curve by deriving closed-form formulas for the number of clusters in a given query region of an arbitrary shape (e.g., polygons and polyhedra). Both the asymptotic solution for the general case and the exact solution for a special case generalize previous work [14]. They agree with the empirical results that the number of clusters depends on the hypersurface area of the query region and not on its hypervolume. We also show that the Hilbert curve achieves better clustering than the z curve. From a practical point of view, the formulas given in this paper provide a simple measure that can be used to predict the required disk access behaviors and, hence, the total access time.
Bongki Moon, H. V. Jagadish, Christos Faloutsos, Joel H. Saltz
IEEE Trans. Knowl. Data Eng.4
2000 Evaluation of Active Disks for Decision Support Databases
abstract
Growth and usage trends for large decision support databases indicate that there is a need for architectures that scale the processing power as the dataset grows. To meet this need, several researchers have recently proposed active disk architectures which integrate substantial processing power and memory into disk units. In this paper, we evaluate Active Disks for decision support databases. First, we compare the performance of Active Disks with that of existing scalable server architectures: SMP-based conventional disk farms and commodity clusters of PCs. Second, we evaluate the impact of several design choices on the performance of Active Disks. We focus on the performance impact of interconnect bandwidth, amount of disk memory and disk-to-disk communication architecture on decision support workloads. Our results show that for identical disks, number of processors and I/O interconnect, Active Disks provide better price/performance than both SMP-based conventional disk farms and commodity cluster. Experiments evaluating the impact of design alternatives in Active Disk architectures indicate that: (1) for configurations up to 64 disks, a dual fibre channel arbitrated loop interconnect is sufficient even for the most communication-intensive decision support tasks: (2) most decision support task do not require a large amount of memory: and (3) direct disk-to-disk communication is necessary for achieving good performance on tasks that repartition all (or a large fraction of) their dataset.
Mustafa Uysal, Anurag Acharya 0001, Joel H. Saltz
HPCA3
2000 Identifying Parallelism in Programs with Cyclic Graphs
abstract
Dependence analysis algorithms have been proposed to identify parallelism in programs with tree-like data structures. However, they can not analyze the dependence of statements if recursive data structures of programs are cyclic. This paper presents a technique to identify parallelism in programs with cyclic graphs. The technique consists of three steps: (1) Traversal patterns that loops or recursive procedures traverse graphs are identified, and the statements that construct the links of traversal patterns are located by definition-use chains of recursive data structures; (2) Shape analysis is performed to estimate possible shapes of traversal patterns; (3) Dependence analysis is performed to identify parallelism using the result of shape analysis. This approach can identify parallelism in programs with cyclic data structures due to the facts that many programs follow acyclic structures (i.e. traversal patterns) to access all nodes on the cyclic data structures. Once the traversal patterns are isolated from the overall data structures, dependence analysis can be applied to identify parallelism.
Yuan-Shin Hwang, Joel H. Saltz
ICPP2
2000 Compiling object-oriented data intensive applications
abstract
Processing and analyzing large volumes of data plays an increasingly important role in many domains of scientific research. High-level language and compiler support for developing applications that analyze and process such datasets has, however, been lacking so far. In this paper, we present a set of language extensions and a prototype compiler for supporting high-level object-oriented programming of data intensive reduction operations over multidimensional data. We have chosen a dialect of Java with data-parallel extensions for specifying collection of objects, a parallel for loop, and reduction variables as our source high-level language. Our compiler analyzes parallel loops and optimizes the processing of datasets through the use of an existing run-time system, called Active Data Repository (ADR). We show how loop fission followed by interprocedural static program slicing can be used by the compiler to extract required information for the run-time system. We present the design of a co...
Renato Ferreira 0001, Gagan Agrawal, Joel H. Saltz
ICS3
2000 Optimizing Retrieval and Processing of Multi-Dimensional Scientific Datasets
abstract
We have developed the Active Data Repository (ADR), an infrastructure that integrates storage, retrieval, and processing of large multi-dimensional scientific datasets on distributed memory parallel machines with multiple disks attached to each node. In earlier work, we proposed three strategies for processing range queries within the ADR framework. Our experimental results show that the relative performance of the strategies changes under varying application characteristics and machine configurations. In this work we investigate approaches to guide and automate the selection of the best strategy for a given application and machine configuration. We describe analytical models to predict the relative performance of the strategies where input data elements are uniformly distributed in the attribute space of the output dataset, restricting the output dataset to be a regular d-dimensional array.
Chialin Chang, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
IPDPS4
1999 A High-Performance Database System for Managing Large Multi-resolution Medical Images
Tahsin M. Kurç, Michael D. Beynon, Chialin Chang, Renato Ferreira 0001, Benjamin B. Bederson, Joel H. Saltz, Alan Sussman
AMIA6
1999 Performance impact of proxies in data intensive client-server applications
abstract
Large client-server data intensive applications can place high demands on system and network resources.This is especially true when the connection between the client and server spans a widearea intemet link.In this paper, we describe our experience changing the typical client-server architecture for a class of data intensive applications.We show that given sufficient common interest among multiple clients, our enhancements reduce the response time per-client, the response-time variation and the amount of data sent across the wide-area link.In addition, we also see a reduction in server utilization which helps to improve server scalability.
Michael D. Beynon, Alan Sussman, Joel H. Saltz
International Conference on Supercomputing3
1999 Querying Very Large Multi-dimensional Datasets in ADR
abstract
Applications that make use of very large scientific datasets have become an increasingly important subset of scientific applications.In these applications, datasets are often multi-dimensional, i.e., data items are associated with points in a multi-dimensional attribute space, and access to data items is described by range queries.The basic processing involves mapping input data items to output data items, and some form of aggregation of all the input data items that project to the each output data item.We have developed an infrastructure, called the Active Data Repository (ADR), that integrates storage, retrieval and processing of multi-dimensional datasets on distributed-memory parallel architectures with multiple disks attached to each node.In this paper we address efficient execution of range queries on distributed memory parallel machines within ADR framework.We present three potential strategies, and evaluate them under different application scenarios and machine configurations.We present experimental results on the scalability and performance of the strategies on a 128-node IBM SP.
Tahsin M. Kurç, Chialin Chang, Renato Ferreira 0001, Alan Sussman, Joel H. Saltz
SC5
1999 Active Storage Hierarchy, Database Systems and Applications - Socratic Exegesis
Felipe Cariño, William O'Connell, John G. Burgess, Joel H. Saltz
VLDB4
1998 Digital dynamic telepathology-the Virtual Microscope
Asmara Afework, Michael D. Beynon, Fabián E. Bustamante, Soon Cho, Angelo Demarzo, Renato Ferreira 0001, Mark Silberman, Joel H. Saltz, Alan Sussman, Hubert Tsang
AMIA9
1998 A graphical tool for ad hoc query generation
Kilian Stoffel, John D. Davis, Gerald Rottman, Joel H. Saltz, James Dick, William Merz
AMIA4
1998 Active Disks: Programming Model, Algorithms and Evaluation
abstract
Several application and technology trends indicate that it might be both profitable and feasible to move computation closer to the data that it processes. In this paper, we evaluate Active Disk architectures which integrate significant processing power and memory into a disk drive and allow application-specific code to be downloaded and executed on the data that is being read from (written to) disk. The key idea is to offload bulk of the processing to the diskresident processors and to use the host processor primarily for coordination, scheduling and combination of results from individual disks. To program Active Disks, we propose a stream-based programming model which allows disklets to be executed efficiently and safely. Simulation results for a suite of six algorithms from three application domains (commercial data warehouses, image processing and satellite data processing) indicate that for these algorithms, Active Disks outperform conventional-disk architectures.
Anurag Acharya 0001, Mustafa Uysal, Joel H. Saltz
ASPLOS3
1998 Adapting to Bandwidth Variations in Wide-Area Data Combination
abstract
Efficient data combination over wide area networks is hard as these networks have large variations in available bandwidth. We examine the utility of changing the location of combination operations as a technique to adapt to variations in network bandwidth. We try to answer the following questions. First, does relocation of operators provide a significant performance improvement? Second, is online relocation useful or does a one-time positioning at start-up time provide most if not all the benefits? If online relocation is useful, how frequently should it be done and is global knowledge of network performance required or can local knowledge and local relocation of operators be sufficient? Fourth, does the effectiveness of operator relocation depend on the ordering of the combination operations. That is, are certain ways of ordering more amenable to adaptation than others? Finally, how do the results change as the number of data sources changes?.
M. Ranganathan, Anurag Acharya 0001, Joel H. Saltz
ICDCS3
1998 The Design and Evaluation of a High-Performance Earth Science Database
Carter Shock, Chialin Chang, Bongki Moon, Anurag Acharya 0001, Larry Davis 0001, Joel H. Saltz, Alan Sussman
Parallel Comput.6
1998 Scalability Analysis of Declustering Methods for Multidimensional Range Queries
abstract
Efficient storage and retrieval of multi-attribute data sets has become one of the essential requirements for many data-intensive applications. The Cartesian product file has been known as an effective multi-attribute file structure for partial-match and best-match queries. Several heuristic methods have been developed to decluster Cartesian product files across multiple disks to obtain high performance for disk accesses. Although the scalability of the declustering methods becomes increasingly important for systems equipped with a large number of disks, no analytic studies have been done so far. The authors derive formulas describing the scalability of two popular declustering methods-Disk Module and Fieldwise Xor-for range queries, which are the most common type of queries. These formulas disclose the limited scalability of the declustering methods, and this is corroborated by extensive simulation experiments. From the practical point of view, the formulas given in the paper provide a simple measure that can be used to predict the response time of a given range query and to guide the selection of a declustering method under various conditions.
Bongki Moon, Joel H. Saltz
IEEE Trans. Knowl. Data Eng.2
1997 The Virtual Microscope
Renato Ferreira 0001, Bongki Moon, Jim Humphries, Alan Sussman, Joel H. Saltz, Angelo Demarzo
AMIA5
1997 Semantic Indexing for Complex Patient Grouping
Kilian Stoffel, Joel H. Saltz, James A. Hendler, James Dick, William Merz
AMIA2
1997 Titan: A High-Performance Remote Sensing Database
abstract
There are two major challenges for a high performance remote sensing database. First, it must provide low latency retrieval of very large volumes of spatio temporal data. This requires effective declustering and placement of a multidimensional dataset onto a large disk farm. Second, the order of magnitude reduction in data size due to post processing makes it imperative, from a performance perspective, that the post processing be done on the machine that holds the data. This requires careful coordination of computation and data retrieval. The paper describes the design, implementation and evaluation of Titan, a parallel shared nothing database designed for handling remote sensing data. The computational platform for Titan is a 16 processor IBM SP-2 with four fast disks attached to each processor. Titan is currently operational and contains about 24 GB of AVHRR data from the NOAA-7 satellite. The experimental results show that Titan provides good performance for global queries and interactive response times for local queries.
Chialin Chang, Bongki Moon, Anurag Acharya 0001, Carter Shock, Alan Sussman, Joel H. Saltz
ICDE6
1997 Page Replacement Using Marginal Loss Functions
abstract
This paper describes a new technique to reduce page-faults in multiprocessing systems by supplying compile-time information about application's access patterns to the kernel. Runtime support is used to determine parameters unknown at compile-time and to summarize all this information as a marginal-loss function. Marginal loss functions describe the extra number of page faults that a process incurs when one of its pages is removed from main memory. System calls convey the marginal loss function of each active memory segment to the kernel which uses this information to determine the allocation and replacement of pages. We outline how the compiler analyzes access patterns and how the marginal loss function can be computed for common patterns. Simulation results demonstrate a significant reduction in page faults compared to typical approaches used by current operating systems.
Manuel Ujaldon, Shamik D. Sharma, Joel H. Saltz
SC3
1997 The Utility of Exploiting Idle Workstations for Parallel Computation
abstract
In this paper, we examine the utility of exploiting idle workstations for parallel computation. We attempt to answer the following questions. First, given a workstation pool, for what fraction of time can we expect to find a cluster of $k$ workstations available? This provides an estimate of the opportunity for parallel computation. Second, how stable is a cluster of free machines and how does the stability vary with the size of the cluster? This indicates how frequently a parallel computation might have to stop for adapting to changes in processor availability. Third, what is the distribution of workstation idle-times? This information is useful for selecting workstations to place computation on. Fourth, how much benefit can a user expect? To state this in concrete terms, if I have a pool of size S, how big a parallel machine should I expect to get for free by harvesting idle machines. Finally, how much benefit can be achieved on a real machine and how hard does a parallel programmer have to work to make this happen? To answer the workstation-availability questions, we have analyzed 14-day traces from three workstation pools. To determine the equivalent parallel machine, we have simulated the execution of a group of well-known parallel programs on these workstation pools. To gain an understanding of the practical problems, we have developed the system support required for adaptive parallel programs as well as an adaptive parallel CFD application. (Also cross-referenced as UMIACS-TR-96-80)
Anurag Acharya 0001, Guy Edjlali, Joel H. Saltz
SIGMETRICS3
1997 Resource-aware metacomputing
abstract
In this paper we outline some potential applications of resource-aware scheduling to high-performance metacomputing applications and describe requirements associated with the use of mobility for resource-aware scheduling. Programs that use mobility as a mechanism to adapt to resource changes have three requirements that are not shared with other mobile programs. Firstly, they need to monitor the level and quality of resources in their operating environment. Secondly, they need to be able to react to changes in resource availability. Thirdly, they need to be able to control the way in which resources are used on their behalf (by libraries and other support code). In this paper, we describe the design and implementation of Sumatra, an extension of Java that supports resource-aware mobile programs. We also describe the design and implementation of a distributed resource monitor that provides the information required by Sumatra programs. Finally, we present a prototype resource-aware data intensive program that combines and composes weather images from multiple geographically distributed sources. © 1997 John Wiley & Sons, Ltd.
Anurag Acharya 0001, M. Ranganathan, Joel H. Saltz
Concurr. Pract. Exp.3
1997 Interprocedural Data Flow Based Optimizations for Distributed Memory Compilation
abstract
Data parallel languages like High Performance Fortran (HPF) are emerging as the architecture independent mode of programming distributed memory parallel machines. In this paper, we present the interprocedural optimizations required for compiling applications having irregular data access patterns, when coded in such data parallel languages. We have developed an Interprocedural Partial Redundancy Elimination (IPRE) algorithm for optimized placement of runtime preprocessing routine and collective communication routines inserted for managing communication in such codes. We also present two new interprocedural optimizations: placement of scatter routines and use of coalescing and incremental routines. We then describe how program slicing can be used for further applying IPRE in more complex scenarios. We have done a preliminary implementation of the schemes presented here using the Fortran D compilation system as the necessary infrastructure. We present experimental results from two codes compiled usng our system to demonstrate the efficacy of the presented schemes. ©1997 John Wiley & Sons, Ltd.
Gagan Agrawal, Joel H. Saltz
Softw. Pract. Exp.2
1996 An Interprocedural Framework for Placement of Asynchronous I/O Operations
abstract
Overlapping memory accesses with computations is a standard technique for improving performance on modern architectures, which have deep memory hierarchies. In this paper, we present a compiler technique for overlapping accesses to secondary memory (disks) with computation. We have developed an Interprocedural Balanced Code Placement (IBCP) framework, which performs analysis on arbitrary recursive procedures and arbitrary control flow and replaces synchronous I/O operations with a balanced pair of asynchronous operations. We demonstrate how this analysis is useful for applications which perform frequent and large accesses to secondary memory, including applications which snapshot or checkpoint their computations or out-of-core applications. 1 Introduction Modern architectures have large number of memory hierarchies. Processors have one or two levels of cache, followed by primary memory (RAM), secondary memory (disks) and tertiary memory. The cost of data access increases rapidly with ...
Gagan Agrawal, Anurag Acharya 0001, Joel H. Saltz
International Conference on Supercomputing3
1996 Runtime Coupling of Data-Parallel Programs
abstract
\f e conslcier the problem of efficiently couphrrg multiple cfataparakl programs at Irrntime.We propose an approach that establishes mappings between data structures m different flat a-parallel programs and implements a user-specified com wmencv model.Mappings are established at rnntlrne and can be added and deleted while the programs being coupled are in execution.Mappings, or the icleutity of the processors involved, do not ha~,r TObe kno~vn at compile-time or even link-time.Pro,gramh ran be m adc to interact tvith different granularities of interaction without requiring any re-cociing.A-priori knowledge of consistency requirements allows buffering of data as well as concurrent execution of the coupled applications.Efficient data movement is achii=w=d b~-p] e-computing an optimized schedule.~~e describe our ])lorotype mrpiernent,atiou and evaluate lts performance us-LMga wt of svnthetrc "benchmarks.\Ve examme the varlat]on of performance with varlatlon in t he conslstenc} reqrrme-[nent W+ demonstrate That the cost of' the tlexibiht}-pro-~,lded hl our coupling scheme is not prolubltl~,e whel L ronpared with a monolithic program that performs the sam~ computatlou.1
M. Ranganathan, Anurag Acharya 0001, Guy Edjlali, Alan Sussman, Joel H. Saltz
International Conference on Supercomputing5
1996 Experimental Evaluation of Efficient Sparse Matrix Distributions
abstract
Sparse matrix problems are difficult to parallelize efficiently on distributed memory machines since non-zero elements are unevenly scattered and are accessed via multiple levels of indirection. Distributions that achieve good load balance and locality are hard to compute and also lead to further indirection in locating distributed data. This paper evaluates alternative distribution strategies which trade off the quality of load-balance and locality for lower decomposition costs and efficient lookup. The proposed techniques are compared with previous strategies for parallelizing sparse matrix problems and the relative merits of each method is outlined. 1 Introduction Sparse matrices are used in a large number of important scientific codes, such as molecular dynamics, CFD solvers, finite element methods and climate modelling. Unfortunately, these applications are hard to parallelize efficiently, particularly using automated compiler techniques. This is because sparse matrices are repre...
Manuel Ujaldon, Shamik D. Sharma, Emilio L. Zapata, Joel H. Saltz
International Conference on Supercomputing4
1996 Parallelization Techniques for Sparse Matrix Applications
Manuel Ujaldon, Emilio L. Zapata, Shamik D. Sharma, Joel H. Saltz
J. Parallel Distributed Comput.4
1995 Interprocedural Partial Redundancy Elimination and its Application to Distributed Memory Compilation
abstract
Partial Redundancy Elimination (PRE) is a general scheme for suppressing partial redundancies which encompasses traditional optimizations like loop invariant code motion and redundant code elimination. In this paper we address the problem of performing this optimization interprocedurally. We use interprocedural partial redundancy elimination for placement of communication and communication preprocessing statements while compiling for distributed memory parallel machines.
Gagan Agrawal, Joel H. Saltz, Raja Das
PLDI2
1995 Efficient Support for Irregular Applications on Distributed-Memory Machines
abstract
Irregular computation problems underlie many important scientific applications. Although these problems are computationally expensive, and so would seem appropriate for parallel machines, their irregular and unpredictable run-time behavior makes this type of parallel program difficult to write and adversely affects run-time performance.
Shubhendu S. Mukherjee, Shamik D. Sharma, Mark D. Hill, James R. Larus, Anne Rogers, Joel H. Saltz
PPoPP6
1995 Interprocedural Compilation of Irregular Applications for Distributed Memory Machines
abstract
Data parallel languages like High Performance Fortran (HPF) are emerging as the architecture independent mode of programming distributed memory parallel machines. In this paper, we present the interprocedural optimizations required for compiling applications having irregular data access patterns, when coded in such data parallel languages. We have developed an Interprocedural Partial Redundancy Elimination (IPRE) algorithm for optimized placement of runtime preprocessing routine and collective communication routines inserted for managing communication in such codes. We also present two new interprocedural optimizations: placement of scatter routines and use of coalescing and incremental routines. We then describe how program slicing can be used for further applying IPRE in more complex scenarios. We have done a preliminary implementation of the schemes presented here using the Fortran D compilation system as the necessary infrastructure. We present experimental results from two codes c...
Gagan Agrawal, Joel H. Saltz
SC2
1995 Index Array Flattening Through Program Transformation
abstract
This paper presents techniques for compiling loops with complex, indirect array accesses into loops whose array references have at most one level of indirection. The transformation allows prefetching of array indices for more efficient structuring of communication on distributed-memory machines. It can also improve performance on other architectures by enabling prefetching of data between levels of the memory hierarchy or exploitation of hardware support for vectorized gather/scatter. Our techniques are implemented in a compiler for Fortran D and execution speed improvements are given for multiprocessor and vector machines.
Raja Das, Paul Havlak, Joel H. Saltz, Ken Kennedy
SC3
1995 Runtime and Language Support for Compiling Adaptive Irregular Programs on Distributed-memory Machines
abstract
Abstract In many scientific applications, arrays containing data are indirectly indexed through indirection arrays. Such scientific applications are called irregular programs and are a distinct class of applications that require special techniques for parallelization. This paper presents a library called CHAOS, which helps users implement irregular programs on distributed‐memory message‐passing machines, such as the Paragon, Delta, CM‐5 and SP‐1. The CHAOS library provides efficient runtime primitives for distributing data and computation over processors; it supports efficient index translation mechanisms and provides users high‐level mechanisms for optimizing communication. CHAOS subsumes the previous PARTI library and supports a larger class of applications. In particular, it provides efficient support for parallelization of adaptive irregular programs where indirection arrays are modified during the course of computation. To demonstrate the efficacy of CHAOS, two challenging real‐life adaptive applications were parallelized using CHAOS primitives: a molecular dynamics code, CHARMM, and a particle‐in‐cell code, DSMC. Besides providing runtime support to users, CHAOS can also be used by compilers to automatically parallelize irregular applications. This paper demonstrates how CHAOS can be effectively used in such a framework. By embedding CHAOS primitives in the Syracuse Fortran 90D/HPF compiler, kernels taken from the CHARMM and DSMC codes have been automatically parallelized.
Yuan-Shin Hwang, Bongki Moon, Shamik D. Sharma, Ravi Ponnusamy, Raja Das, Joel H. Saltz
Softw. Pract. Exp.6
1995 Distributed Memory Compiler Design for Sparse Problems
abstract
This paper addresses the issue of compiling concurrent loop nests in the presence of complicated array references and irregularly distributed arrays. Arrays accessed within loops may contain accesses that make it impossible to precisely determine the reference pattern at compile time. This paper proposes a run time support mechanism that is used effectively by a compiler to generate efficient code in these situations. The compiler accepts as input a Fortran 77 program enhanced with specifications for distributing data, and outputs a message passing program that runs on the nodes of a distributed memory machine. The runtime support for the compiler consists of a library of primitives designed to support irregular patterns of distributed array accesses and irregularly distributed array partitions. A variety of performance results on the Intel iPSC/860 are presented.>
Janet Wu, Raja Das, Joel H. Saltz, Harry Berryman, Seema Hiranandani
IEEE Trans. Computers3
1995 Implementation of a parallel unstructured Euler solver on shared- and distributed-memory architectures
Dimitri J. Mavriplis, Raja Das, Joel H. Saltz, R. E. Vermeland
J. Supercomput.3
1995 An Integrated Runtime and Compile-Time Approach for Parallelizing Structured and Block Structured Applications
abstract
In compiling applications for distributed memory machines, runtime analysis is required when data to be communicated cannot be determined at compile-time. One such class of applications requiring runtime analysis is block structured codes. These codes employ multiple structured meshes, which may be nested (for multigrid codes) and/or irregularly coupled (called multiblock or irregularly coupled regular mesh problems). In this paper, we present runtime and compile-time analysis for compiling such applications on distributed memory parallel machines in an efficient and machine-independent fashion. We have designed and implemented a runtime library which supports the runtime analysis required. The library is currently implemented on several different systems. We have also developed compiler analysis for determining data access patterns at compile time and inserting calls to the appropriate runtime routines. Our methods can be used by compilers for HPF-like parallel programming languages in compiling codes in which data distribution, loop bounds and/or strides are unknown at compile-time. To demonstrate the efficacy of our approach, we have implemented our compiler analysis in the Fortran 90D/HPF compiler developed at Syracuse University. We have experimented with a multi-bloc Navier-Stokes solver template and a multigrid code. Our experimental results show that our primitives have low runtime communication overheads and the compiler parallelized codes perform within 20% of the codes parallelized by manually inserting calls to the runtime library.>
Gagan Agrawal, Alan Sussman, Joel H. Saltz
IEEE Trans. Parallel Distributed Syst.3
1995 Runtime Support and Compilation Methods for User-Specified Irregular Data Distributions
abstract
This paper describes two new ideas by which a High Performance Fortran compiler can deal with irregular computations effectively. The first mechanism invokes a user specified mapping procedure via a set of proposed compiler directives. The directives allow use of program arrays to describe graph connectivity, spatial location of array elements, and computational load. The second mechanism is a conservative method for compiling irregular loops in which dependence arises only due to reduction operations. This mechanism in many cases enables a compiler to recognize that it is possible to reuse previously computed information from inspectors (e.g., communication schedules, loop iteration partitions, and information that associates off-processor data copies with on-processor buffer locations). This paper also presents performance results for these mechanisms from a Fortran 90D compiler implementation.>
Ravi Ponnusamy, Joel H. Saltz, Alok N. Choudhary, Yuan-Shin Hwang, Geoffrey C. Fox
IEEE Trans. Parallel Distributed Syst.2
1994 High performance computing for land cover dynamics
abstract
Presents the overall goals of the authors' research program on the application of high performance computing to remote sensing applications, specifically applications in land cover dynamics. This involves developing scalable and portable programs for a variety of image and map data processing applications, eventually integrated with new models for parallel I/O of large scale images and maps. After an overview of the multiblock PARTI run time support system, the authors explain extensions made to that system to support image processing applications, and then present an example involving multiresolution image processing. Results of running the parallel code on both a TMC CM5 and an Intel Paragon are discussed.
Rahul Parulekar, Larry Davis 0001, Rama Chellappa, Joel H. Saltz, Alan Sussman, John Townshend
ICPR (3)4
1994 Run-time and compile-time support for adaptive irregular problems
abstract
In adaptive irregular problems, data arrays are accessed via indirection arrays, and data access patterns change during computation. Parallelizing such problems on distributed memory machines requires support for dynamic data partitioning, efficient preprocessing and fast data migration. This paper describes CHAOS, a library of efficient runtime primitives that provides such support. To demonstrate the effectiveness of the runtime support, two adaptive irregular applications have been parallelized using CHAOS primitives: a molecular dynamics code (CHARMM) and a code for simulating gas flows (DSMC). We have also proposed minor extensions to Fortran D which would enable compilers to parallelize irregular for all loops in such adaptive applications by embedding calls to primitives provided by a runtime library. We have implemented our proposed extensions in the Syracuse Fortran 90D/HPF prototype compiler, and have used the compiler to parallelize kernels from two adaptive applications.>
Shamik D. Sharma, Ravi Ponnusamy, Bongki Moon, Yuan-Shin Hwang, Raja Das, Joel H. Saltz
SC6
1994 Communication Optimizations for Irregular Scientific Computations on Distributed Memory Architectures
Raja Das, Mustafa Uysal, Joel H. Saltz, Yuan-Shin Hwang
J. Parallel Distributed Comput.3
1993 Compiler and runtime support for structured and block structured applications
abstract
Scientific and engineering applications often involve structured meshes. These meshes may be nested (for multigrid or adaptive codes) and/or irregularly coupled(called Irregularly CoupledRegular Meshes). We have designed and implemented a runtime library for parallelizing this general class of applications on distributed memory parallel machines in an efficient and machine independent manner. In this paper we present how this runtime library can be integrated with compilers for High PerformanceFortran (HPF) style parallel programming languages. We discuss how we have integrated this runtime library with the Fortran 90D compiler being developed at Syracuse University and provide experimental data on a block structured NavierStokes solver template and a small multigrid example parallelized using this compiler and run on an Intel iPSC/860. We show that the compiler parallelizedcode performs within 20% of the code parallelized by inserting calls to the runtime library manually.
Gagan Agrawal, Alan Sussman, Joel H. Saltz
SC3
1993 Common runtime support for high-performance parallel languages
abstract
No abstract available.
Geoffrey C. Fox, Sanjay Ranka, Michael L. Scott, Allen D. Malony, James C. Browne, Marina C. Chen, Alok N. Choudhary, Thomas E. Cheatham, Janice E. Cuny, Rudolf Eigenmann, Amr F. Fahmy, Ian T. Foster, Dennis Gannon, Tomasz Haupt, Carl Kesselman, Charles Koelbel, Wei Li 0015, Monica S. Lam, Thomas J. LeBlanc, Jim Openshaw, David A. Padua, Constantine D. Polychronopoulos, Joel H. Saltz, Alan Sussman, Gil Weigand, Katherine A. Yelick
SC23
1993 Runtime compilation techniques for data partitioning and communication schedule reuse
abstract
In this paper, we describe two new ideas by which HPF compiler can deal with irregular computations effectively. The first mechanism invokes a user specified mapping procedure via a set of compiler directives. The directives allow the user to use program arrays to describe graph connectivity, spatial location of array elements and computational load. The second is a simple conservative method that in many cases enables a compiler to recognize that it is possible to reuse previously computed results from inspectors (e.g. communication schedules, loop iteration partitions, information that associates off-processor data copies with on-processor buffer locations). We present performance results for these mechanisms from a Fortran 90D compiler implementation. 1 Introduction In sparse and unstructured problems the data access pattern is determined by variable values known only at runtime. In these cases, programmers carry out preprocessig to partition work, map data structures and schedule ...
Ravi Ponnusamy, Joel H. Saltz, Alok N. Choudhary
SC2
1992 Compiler and runtime support for irregularly coupled regular meshes
abstract
Regular meshes are frequently used for modeling physical phenomena on both serial and parallel computers. One advantage of regular meshes is that efficient discretization schemes can be implemented in a straightforward manner. However, geometrically-complex objects, such as aircraft, cannot be easily described using a single regular mesh. Multiple interacting regular meshes are frequently used to describe complex geometries. Each mesh models a subregion of the physical domain. The meshes, or subdomains, can be processed in parallel, with periodic updates carried out to move information between the coupled meshes. In many cases, there are a relatively small number (one to a few dozen) subdomains, so that each subdomain may also be partitioned among several processors.
Craig M. Chase, Kay Crowley, Joel H. Saltz, Anthony P. Reeves
ICS3
1992 Implementation of a Parallel Unstructured Euler Solver on Shared and Distributed Memory Architectures
abstract
An efficient three-dimensional unstructured Euler solver has been parallelized on a Cray Y-MP C90 shared memory computer and on an Intel Touchstone Delta distributed memory computer. Both machines yield comparable performance rates. However, the availability of sophisticated software tools enabled the parallelization of EUL3D on the shared memory vector/parallel CRAY Y-MP C90 with minimal user input. On the other hand, the implementation on the distributed memory massively parallel architecture of the Intel Touchstone Delta machine is considerably more involved. As massively parallel software tools become more mature, the task of developing or porting software to such machines should diminish. It has also been shown that with today's supercomputers, and with efficient codes such as EUL3D, the aerodynamic characteristics of complex vehicles can be computed in a matter of minutes, making design use feasible.>
Dimitri J. Mavriplis, Raja Das, R. E. Vermeland, Joel H. Saltz
SC4
1992 Performance of Distributed Sparse Cholesky Factorization with Pre-Scheduling
abstract
The authors propose a Cholesky factorization scheme based on a static task and communication schedule generated by symbolic preprocessing of the intercolumn dependencies. This information is used to reduce the overheads of maintaining data structures and the communication costs incurred in message passing during the numerical factorization step. The authors introduce three primitives which encapsulate these optimizations and alleviate the user's programming effort. Performance results obtained on the iPSC/860 show that an implementation using these primitives results in 30% to 40% savings over one that does not use any static communication structure information.>
Sesh Venugopal, Vijay K. Naik, Joel H. Saltz
SC3
1991 The Preprocessed Doacross Loop
Joel H. Saltz, Ravi Mirchandaney
ICPP (2)1
1991 Runtime Compilation Methods for Multicomputers
Janet Wu, Joel H. Saltz, Seema Hiranandani, Harry Berryman
ICPP (2)2
1991 Execution time support for adaptive scientific algorithms on distributed memory machines
abstract
Abstract We consider optimizations that are required for efficient execution of code segments that consist of loops over distributed data structures. The PARTI execution time primitives are designed to perform these optimizations and can be used to Implement a wide range of scientific algorithms on distributed memory machines. These primitives allow the user to control array mappings in a way that gives an appearance of shared memory. Computations can be based on a global index set. Primitives are used to perform gather and scatter operations on distributed arrays. Communications patterns are derived at run time, and the appropriate send and receive messages are automatically generated.
Harry Berryman, Joel H. Saltz, Jeffrey S. Scroggs
Concurr. Pract. Exp.2
1991 Multiprocessors and run-time compilation
abstract
Abstract Run‐time preprocessing plays a major role in many efficient algorithms in computer science, as well as playing an important role in exploiting multiprocessor architectures. We give examples that elucidate the importance of run‐time preprocessing and show how these optimizations can be integrated into compilers. To support our arguments, we describe transformations implemented in prototype multiprocessor compilers and present benchmarks from the iPSC2/860, the CM‐2 and the Encore Multimax/320.
Joel H. Saltz, Harry Berryman, Janet Wu
Concurr. Pract. Exp.1
1991 Performance of Hashed Cache Data Migration Schemes on Multicomputers
Seema Hiranandani, Joel H. Saltz, Piyush Mehrotra, Harry Berryman
J. Parallel Distributed Comput.2
1991 Performance Effects of Irregular Communication Patterns on Massively Parallel Multiprocessors
Joel H. Saltz, Serge G. Petition, Harry Berryman, Adam Rifkin
J. Parallel Distributed Comput.1
1991 Run-Time Parallelization and Scheduling of Loops
abstract
The authors study run-time methods to automatically parallelize and schedule iterations of a do loop in certain cases where compile-time information is inadequate. The methods presented involve execution time preprocessing of the loop. At compile-time, these methods set up the framework for performing a loop dependency analysis. At run-time, wavefronts of concurrently executable loop iterations are identified. Using this wavefront information, loop iterations are reordered for increased parallelism. The authors utilize symbolic transformation rules to produce: inspector procedures that perform execution time preprocessing, and executors or transformed versions of source code loop structures. These transformed loop structures carry out the calculations planned in the inspector procedures. The authors present performance results from experiments conducted on the Encore Multimax. These results illustrate that run-time reordering of loop indexes can have a significant impact on performance.>
Joel H. Saltz, Ravi Mirchandaney, Kay Crowley
IEEE Trans. Computers1
1990 Krylov Methods Preconditioned with Incompletely Factored Matrices on the CM-2
abstract
The performance is measured of the components of the key interative kernel of a preconditioned Krylov space interative linear system solver. In some sense, these numbers can be regarded as best case timings for these kernels. Sweeps were timed over meshes, sparse triangular solves, and inner products on a large 3-D model problem over a cube shaped domain discretized with a seven point template. The performance of the CM-2 is highly dependent on the use of very specialized programs. These programs mapped a regular problem domain onto the processor topology in a careful manner and used the optimized local NEWS communications network. The rather dramatic deterioration in performance was documented when these ideal conditions no longer apply. A synthetic workload generator was developed to produce and solve a parameterized family of increasingly irregular problems.
Harry Berryman, Joel H. Saltz, William Gropp, Ravi Mirchandaney
J. Parallel Distributed Comput.2
1990 Run-Time Scheduling and Execution of Loops on Message Passing Machines
abstract
We examine the effectiveness of optimizations aimed to allowing distributed machine to efficiently compute inner loops over globally defined data structures. Our optimizations are specifically targeted toward loops in which some array references are made through a level of indirection. Unstructured mesh codes and sparse matrix solvers are examplese of programs with kernels of this sort. Experimental data that quantify the performance obtainable using the methods discussed here are included.
Joel H. Saltz, Kathleen Crowley, Ravi Mirchandaney, Harry Berryman
J. Parallel Distributed Comput.1
1990 An Anlysis of Scatter Decomposition
abstract
A formal analysis of a powerful mapping technique known as scatter decomposition is provided. Scatter decomposition divides an irregular computational domain into a large number of equally sized pieces and distributes them modularly among processors. A probabilistic model of workload in one dimension is used to formally explain why and when scatter decomposition works. The first result is that if a correlation in workload is a convex function of distance, then scattering a more finely decomposed domain yields a lower average processor workload variance. The second result shows that if the workload process is a stationary Gaussian and the correlation function decreases linearly in distance until becoming zero and then remain zero, scattering a more finely decomposed domain yields a lower expected maximum processor workload. It is shown that if the correlation function decreases linearly across the entire domain, then among all mappings that assign an equal number of domain pieces to each processor, scatter decomposition minimizes the average processor workload variance. The dependence of these results on the assumption of decreasing correlation is illustrated with situations where a coarser granularity actually achieves better load balance.>
David M. Nicol, Joel H. Saltz
IEEE Trans. Computers2
1989 The doconsider loop
abstract
Article Free Access Share on The doconsider loop Authors: Joel H. Saltz Department of Computer Science, Yale University, New Haven, CT Department of Computer Science, Yale University, New Haven, CTView Profile , Ravi Mirchandaney Department of Computer Science, Yale University, New Haven, CT Department of Computer Science, Yale University, New Haven, CTView Profile , Kathleen Crowley Department of Computer Science, Yale University, New Haven, CT Department of Computer Science, Yale University, New Haven, CTView Profile Authors Info & Claims ICS '89: Proceedings of the 3rd international conference on SupercomputingJune 1989Pages 29–40https://doi.org/10.1145/318789.318794Published:01 June 1989Publication History 17citation242DownloadsMetricsTotal Citations17Total Downloads242Last 12 Months11Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Publisher SiteeReaderPDF
Joel H. Saltz, Ravi Mirchandaney, Kathleen Crowley
ICS1
1989 Run-Time Parallelization and Scheduling of Loops
abstract
Run time methods are studied to automatically parallelize and schedule iterations of a do loop in certain cases, where compile-time information is inadequate. The methods presented involve execution time preprocessing of the loop. At compile-time, these methods set up the framework for performing a loop dependency analysis. At run time, wave fronts of concurrently executable loop iterations are identified. Using this wavefront information, loop iterations are reordered for increased parallelism. Symbolic transformation rules are used to produce: inspector procedures that perform execution time preprocessing and executors or transformed versions of source code loop structures. These transformed loop structures carry out the calculations planned in the inspector procedures. Performance results are presented from experiments conducted on the Encore Multimax. These results illustrate that run time reordering of loop indices can have a significant impact on performance. Furthermore, the overheads associated with this type of reordering are amortized when the loop is executed several times with the same dependency structure.
Doug Baxter, Ravi Mirchandaney, Joel H. Saltz
SPAA3
1988 Principles of runtime support for parallel processors
abstract
There exists substantial data level parallelism in scientific problems. The PARTY runtime system is an attempt to obtain efficient parallel implementations for scientific computations, particularly those where the data dependencies are manifest only at runtime. This can preclude compiler based detection of certain types of parallelism. The automated system is structured as follows: An appropriate level of granularity is first selected for the computations. A directed acyclic graph representation of the program is generated on which various aggregation techniques may be employed in order to generate efficient schedules. These schedules are then mapped onto the target machine. We describe some initial results from experiments conducted on the Intel Hypercube and the Encore Multimax that indicate the usefulness of our approach.
Ravi Mirchandaney, Joel H. Saltz, Roger M. Smith, David M. Nicol, Kay Crowley
ICS2
1988 Towards developing robust algorithms for solving partial differential equations on MIMD machines
Joel H. Saltz, Vijay K. Naik
Parallel Comput.1
1988 Dynamic Remapping of Parallel Computations with Varying Resource Demands
abstract
The issue of deciding when to invoke a global load remapping mechanism is studied. Such a decision policy must effectively weigh the costs of remapping against the performance benefits, and should be general enough to apply automatically to a wide range of computations. The authors propose a general mapping decision heuristic, then study its effectiveness and its anticipated behavior on two very different models of load evolution. Assuming only that the remapping cost is known, this policy dynamically minimizes system degradation (including the cost of remapping) for each computation step. This policy is quite simple, choosing to remap when the first local minimum in the degradation function is detected. Simulations show that the decision obtained provides significantly better performance than that achieved by never remapping. The authors also observe that the average intermapping frequency is quite close to the optimal fixed remapping frequency.>
David M. Nicol, Joel H. Saltz
IEEE Trans. Computers2
1986 A Comparative Analysis of Static and Dynamic Load Balancing Strategies
Mohammad Ashraf Iqbal, Joel H. Saltz, Shahid H. Bokhari
ICPP2