Rick L. Stevens

dblp:s/RickLStevens · also Rick Stevens · DBLP profile ↗
← Back
61ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0002-4268-4020ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorTheory of computation · 2 · 1 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Benchmarking community drug response prediction models: datasets, models, tools, and metrics for cross-dataset generalization analysis
abstract
Deep learning and machine learning models have shown promise in drug response prediction (DRP), yet their ability to generalize across datasets remains an open question, raising concerns about their real-world applicability. Due to the lack of standardized benchmarking approaches, model evaluations and comparisons often rely on inconsistent datasets and evaluation criteria, making it difficult to assess true predictive capabilities. In this work, we introduce a benchmarking framework for evaluating cross-dataset prediction generalization in DRP models. Our framework incorporates five publicly available drug screening datasets, seven standardized DRP models, and a scalable workflow for systematic evaluation. To assess model generalization, we introduce a set of evaluation metrics that quantify both absolute performance (e.g. predictive accuracy across datasets) and relative performance (e.g. performance drop compared to within-dataset results), enabling a more comprehensive assessment of model transferability. Our results reveal substantial performance drops when models are tested on unseen datasets, underscoring the importance of rigorous generalization assessments. While several models demonstrate relatively strong cross-dataset generalization, no single model consistently outperforms across all datasets. Furthermore, we identify CTRPv2 as the most effective source dataset for training, yielding higher generalization scores across target datasets. By sharing this standardized evaluation framework with the community, our study aims to establish a rigorous foundation for model comparison, and accelerate the development of robust DRP models for real-world applications.
Alexander Partin, Priyanka Vasanthakumari, Oleksandr Narykov, Andreas Wilke, Natasha Koussa, Sara E. Jones, Yitan Zhu, Jamie C. Overbeek, Rajeev Jain, Gayara Demini Fernando, Cesar Sanchez-Villalobos, Cristina Garcia-Cardona, Jamaludin Mohd-Yusof, Nicholas Chia, Justin M. Wozniak, Souparno Ghosh, Ranadip Pal, Thomas S. Brettin, M. Ryan Weil, Rick L. Stevens
Briefings Bioinform.20
2025 Data imbalance in drug response prediction: multi-objective optimization approach in deep learning setting
abstract
Drug response prediction (DRP) methods tackle the complex task of associating the effectiveness of small molecules with the specific genetic makeup of the patient. Anti-cancer DRP is a particularly challenging task requiring costly experiments as underlying pathogenic mechanisms are broad and associated with multiple genomic pathways. The scientific community has exerted significant efforts to generate public drug screening datasets, giving a path to various machine learning models that attempt to reason over complex data space of small compounds and biological characteristics of tumors. However, the data depth is still lacking compared to application domains like computer vision or natural language processing domains, limiting current learning capabilities. To combat this issue and improves the generalizability of the DRP models, we are exploring strategies that explicitly address the imbalance in the DRP datasets. We reframe the problem as a multi-objective optimization across multiple drugs to maximize deep learning model performance. We implement this approach by constructing Multi-Objective Optimization Regularized by Loss Entropy loss function and plugging it into a Deep Learning model. We demonstrate the utility of proposed drug discovery methods and make suggestions for further potential application of the work to achieve desirable outcomes in the healthcare field.
Oleksandr Narykov, Yitan Zhu, Thomas S. Brettin, Yvonne A. Evrard, Alexander Partin, Fangfang Xia, Maulik Shukla, Priyanka Vasanthakumari, James H. Doroshow, Rick L. Stevens
Briefings Bioinform.10
2024 Blending Imitation and Reinforcement Learning for Robust Policy Improvement
abstract
While reinforcement learning (RL) has shown promising performance, its sample complexity continues to be a substantial hurdle, restricting its broader application across a variety of domains. Imitation learning (IL) utilizes oracles to improve sample efficiency, yet it is often constrained by the quality of the oracles deployed. To address the demand for robust policy improvement in real-world scenarios, we introduce a novel algorithm, Robust Policy Improvement (RPI), which actively interleaves between IL and RL based on an online estimate of their performance. RPI draws on the strengths of IL, using oracle queries to facilitate exploration—an aspect that is notably challenging in sparse-reward RL—particularly during the early stages of learning. As learning unfolds, RPI gradually transitions to RL, effectively treating the learned policy as an improved oracle. This algorithm is capable of learning from and improving upon a diverse set of black-box oracles. Integral to RPI are Robust Active Policy Selection (RAPS) and Robust Policy Gradient (RPG), both of which reason over whether to perform state-wise imitation from the oracles or learn from its own value function when the learner’s performance surpasses that of the oracles in a specific state. Empirical evaluations and theoretical analysis validate that RPI excels in comparison to existing state-of-the-art methodologies, demonstrating superior performance across various benchmark domains.
Takuma Yoneda, Rick L. Stevens, Matthew R. Walter, Yuxin Chen 0001
ICLR3
2024 Entropy-Reinforced Planning with Large Language Models for Drug Discovery
abstract
The objective of drug discovery is to identify chemical compounds that possess specific pharmaceutical properties toward a binding target. Existing large language models (LLMS) can achieve high token matching scores in terms of likelihood for molecule generation. However, relying solely on LLM decoding often results in the generation of molecules that are either invalid due to a single misused token, or suboptimal due to unbalanced exploration and exploitation as a consequence of the LLM’s prior experience. Here we propose ERP, Entropy-Reinforced Planning for Transformer Decoding, which employs an entropy-reinforced planning algorithm to enhance the Transformer decoding process and strike a balance between exploitation and exploration. ERP aims to achieve improvements in multiple properties compared to direct sampling from the Transformer. We evaluated ERP on the SARS-CoV-2 virus (3CLPro) and human cancer cell target protein (RTCB) benchmarks and demonstrated that, in both benchmarks, ERP consistently outperforms the current state-of-the-art algorithm by 1-5 percent, and baselines by 5-10 percent, respectively. Moreover, such improvement is robust across Transformer models trained with different objectives. Finally, to further illustrate the capabilities of ERP, we tested our algorithm on three code generation benchmarks and outperformed the current state-of-the-art approach as well. Our code is publicly available at: https://github.com/xuefeng-cs/ERP.
Chih-chan Tien, Songhao Jiang, Rick L. Stevens
ICML5
2024 Contextual Active Model Selection
abstract
While training models and labeling data are resource-intensive, a wealth of pre-trained models and unlabeled data exists. To effectively utilize these resources, we present an approach to actively select pre-trained models while minimizing labeling costs. We frame this as an online contextual active model selection problem: At each round, the learner receives an unlabeled data point as a context. The objective is to adaptively select the best model to make a prediction while limiting label requests. To tackle this problem, we propose CAMS, a contextual active model selection algorithm that relies on two novel components: (1) a contextual model selection mechanism, which leverages context information to make informed decisions about which model is likely to perform best for a given context, and (2) an active query component, which strategically chooses when to request labels for data points, minimizing the overall labeling cost. We provide rigorous theoretical analysis for the regret and query complexity under both adversarial and stochastic settings. Furthermore, we demonstrate the effectiveness of our algorithm on a diverse collection of benchmark classification tasks. Notably, CAMS requires substantially less labeling effort (less than 10%) compared to existing methods on CIFAR10 and DRIFT benchmarks, while achieving similar or better accuracy.
Fangfang Xia, Rick L. Stevens, Yuxin Chen 0001
NeurIPS3
2024 MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization
abstract
We present a scalable, end-to-end workflow for protein design. By augmenting protein sequences with natural language descriptions of their biochemical properties, we train generative models that can be preferentially aligned with protein fitness landscapes. Through complex experimental-and simulation-based observations, we integrate these measures as preferred parameters for generating new protein variants and demonstrate our workflow on five diverse supercomputers. We achieve >1 ExaFLOPS sustained performance in mixed precision on each supercomputer and a maximum sustained performance of 4.11 Ex-aFLOPS and peak performance of 5.57 ExaFLOPS. We establish the scientific performance of our model on two tasks: (1) across a predetermined benchmark dataset of deep mutational scanning experiments to optimize the fitness-determining mutations in the yeast protein HIS7, and (2) in optimizing the design of the enzyme malate dehydrogenase to achieve lower activation barriers (and therefore increased catalytic rates) using simulation data. Our implementation thus sets high watermarks for multimodal protein design workflows.
Gautham Dharuman, Kyle Hippe, Alex Brace, Sam Foreman, Väinö Hatanpää, Varuni Sastry 0001, Huihuo Zheng, Logan T. Ward, Servesh Muralidharan, Archit Vasan, Bharat Kale, Carla M. Mann, Yun-Hsuan Cheng, Yuliana Zamora, Shengchao Liu, Chaowei Xiao, Murali Emani, Tom Gibbs, Mahidhar Tatineni, Deepak Canchi, Jerome Mitchell, Koichi Yamada, María Jesús Garzarán, Michael E. Papka, Ian T. Foster, Rick L. Stevens, Anima Anandkumar, Venkatram Vishwanath, Arvind Ramanathan
SC27
2023 Trillion Parameter AI Serving Infrastructure for Scientific Discovery: A Survey and Vision
abstract
Deep learning methods are transforming research, enabling new techniques, and ultimately leading to new discoveries. As the demand for more capable AI models continues to grow, we are now entering an era of Trillion Parameter Models (TPM), or models with more than a trillion parameters---such as Huawei's PanGu-Σ. We describe a vision for the ecosystem of TPM users and providers that caters to the specific needs of the scientific community. We then outline the significant technical challenges and open problems in system design for serving TPMs to enable scientific research and discovery. Specifically, we describe the requirements of a comprehensive software stack and interfaces to support the diverse and flexible requirements of researchers.
Nathaniel Hudson 0001, J. Gregory Pauloski, Matt Baughman, Alok Kamatar, Mansi Sakarvadia, Logan T. Ward, Ryan Chard, André Bauer 0001, Maksim Levental, Will Engler, Owen Price Skelly, Ben Blaiszik, Rick L. Stevens, Kyle Chard, Ian T. Foster
BDCAT14
2023 An Automation Framework for Comparison of Cancer Response Models Across Configurations
abstract
Machine learning has made significant advancements in precision medicine, resulting in the development of various deep learning applications. For instance, in cancer drug response prediction, numerous deep learning models have been created. However, comparing these models across vast configurations of hyperparameters and data sets can be challenging. In this paper, we introduce a new scalable workflow suite that aims to answer questions that arise when comparing different models developed by different teams on similar or the same problems. We explain the problem in more detail and discuss our approach using near-exascale or exascale computers.
Justin M. Wozniak, Rajeev Jain, Andreas Wilke, Rylie Weaver, Alexander Partin, Thomas S. Brettin, Rick L. Stevens
e-Science7
2023 Transferable Graph Neural Fingerprint Models for Quick Response to Future Bio-Threats
abstract
Fast screening of drug molecules based on the ligand binding affinity is an important step in the drug discovery pipeline. Graph neural fingerprint is a promising method for developing molecular docking surrogates with high throughput and great fidelity. In this study, we built a COVID-19 drug docking dataset of about 300,000 drug candidates on 23 coronavirus protein targets. With this dataset, we trained graph neural fin-gerprint docking models for high-throughput virtual COVID-19 drug screening. The graph neural fingerprint models yield high prediction accuracy on docking scores with the mean squared error lower than 0.21 kcal/mol for most of the docking targets, showing significant improvement over conventional circular fin-gerprint methods. To make the neural fingerprints transferable for unknown targets, we also propose a transferable graph neural fingerprint method trained on multiple targets. With comparable accuracy to target-specific graph neural fingerprint models, the training and data efficiency of the transferable model is several times higher. We highlight that the impact of this study extends beyond COVID-19 dataset, as our approach for fast virtual ligand screening can be easily adapted and integrated into a general machine learning-accelerated pipeline to battle future bio-threats.
Wei Chen 0043, Yihui Ren 0001, Ai Kagawa, Matthew R. Carbone, Samuel Yen-Chi Chen, Xiaohui Qu, Shinjae Yoo, Austin Clyde, Arvind Ramanathan, Rick L. Stevens, Huub J. J. Van Dam, Deyu Lu
ICMLA10
2023 WordScape: a Pipeline to extract multilingual, visually rich Documents with Layout Annotations from Web Crawl Data
abstract
We introduce WordScape, a novel pipeline for the creation of cross-disciplinary, multilingual corpora comprising millions of pages with annotations for document layout detection. Relating visual and textual items on document pages has gained further significance with the advent of multimodal models. Various approaches proved effective for visual question answering or layout segmentation. However, the interplay of text, tables, and visuals remains challenging for a variety of document understanding tasks. In particular, many models fail to generalize well to diverse domains and new languages due to insufficient availability of training data. WordScape addresses these limitations. Our automatic annotation pipeline parses the Open XML structure of Word documents obtained from the web, jointly providing layout-annotated document images and their textual representations. In turn, WordScape offers unique properties as it (1) leverages the ubiquity of the Word file format on the internet, (2) is readily accessible through the Common Crawl web corpus, (3) is adaptive to domain-specific documents, and (4) offers culturally and linguistically diverse document pages with natural semantic structure and high-quality text. Together with the pipeline, we will additionally release 9.5M urls to word documents which can be processed using WordScape to create a dataset of over 40M pages. Finally, we investigate the quality of text and layout annotations extracted by WordScape, assess the impact on document understanding benchmarks, and demonstrate that manual labeling costs can be substantially reduced.
Maurice Weber, Carlo Siebenschuh, Rory Butler, Anton Alexandrov, Valdemar Thanner, Georgios Tsolakis, Haris Jabbar, Ian T. Foster, Bo Li 0026, Rick L. Stevens, Ce Zhang 0001
NeurIPS10
2023 ChemoGraph: Interactive Visual Exploration of the Chemical Space
abstract
Abstract Exploratory analysis of the chemical space is an important task in the field of cheminformatics. For example, in drug discovery research, chemists investigate sets of thousands of chemical compounds in order to identify novel yet structurally similar synthetic compounds to replace natural products. Manually exploring the chemical space inhabited by all possible molecules and chemical compounds is impractical, and therefore presents a challenge. To fill this gap, we present ChemoGraph, a novel visual analytics technique for interactively exploring related chemicals. In ChemoGraph, we formalize a chemical space as a hypergraph and apply novel machine learning models to compute related chemical compounds. It uses a database to find related compounds from a known space and a machine learning model to generate new ones, which helps enlarge the known space. Moreover, ChemoGraph highlights interactive features that support users in viewing, comparing, and organizing computationally identified related chemicals. With a drug discovery usage scenario and initial expert feedback from a case study, we demonstrate the usefulness of ChemoGraph.
Bharat Kale, Austin Clyde, Maoyuan Sun, Arvind Ramanathan, Rick L. Stevens, Michael E. Papka
Comput. Graph. Forum5
2022 Spatial Graph Attention and Curiosity-driven Policy for Antiviral Drug Discovery
Nicholas Choma, Andrew Deru Chen, Mikaela Cashman, Érica T. Prates, Verónica G. Vergara Larrea, Manesh Shah, Austin Clyde, Thomas S. Brettin, Bert de Jong, Martha S. Head, Rick L. Stevens, Peter Nugent, Daniel A. Jacobson, James B. Brown
ICLR13
2022 A cross-study analysis of drug response prediction in cancer cell lines
abstract
To enable personalized cancer treatment, machine learning models have been developed to predict drug response as a function of tumor and drug features. However, most algorithm development efforts have relied on cross-validation within a single study to assess model accuracy. While an essential first step, cross-validation within a biological data set typically provides an overly optimistic estimate of the prediction performance on independent test sets. To provide a more rigorous assessment of model generalizability between different studies, we use machine learning to analyze five publicly available cell line-based data sets: National Cancer Institute 60, ancer Therapeutics Response Portal (CTRP), Genomics of Drug Sensitivity in Cancer, Cancer Cell Line Encyclopedia and Genentech Cell Line Screening Initiative (gCSI). Based on observed experimental variability across studies, we explore estimates of prediction upper bounds. We report performance results of a variety of machine learning models, with a multitasking deep neural network achieving the best cross-study generalizability. By multiple measures, models trained on CTRP yield the most accurate predictions on the remaining testing data, and gCSI is the most predictable among the cell line data sets included in this study. With these experiments and further simulations on partial data, two lessons emerge: (1) differences in viability assays can limit model generalizability across studies and (2) drug diversity, more than tumor diversity, is crucial for raising model generalizability in preclinical screening.
Fangfang Xia, Jonathan E. Allen, Prasanna Balaprakash, Thomas S. Brettin, Cristina Garcia-Cardona, Austin Clyde, Judith D. Cohn, James H. Doroshow, Xiaotian Duan, Veronika Dubinkina, Yvonne A. Evrard, Ya-Ju Fan, Jason Gans, Stewart He, Pinyi Lu, Sergei Maslov, Alexander Partin, Maulik Shukla, Eric A. Stahlberg, Justin M. Wozniak, Hyun Seung Yoo, George F. Zaki, Yitan Zhu, Rick L. Stevens
Briefings Bioinform.24
2021 IMPECCABLE: Integrated Modeling PipelinE for COVID Cure by Assessing Better LEads
abstract
The drug discovery process currently employed in the pharmaceutical industry typically requires about 10 years and $2–3 billion to deliver one new drug. This is both too expensive and too slow, especially in emergencies like the COVID-19 pandemic. In silico methodologies need to be improved both to select better lead compounds, so as to improve the efficiency of later stages in the drug discovery protocol, and to identify those lead compounds more quickly. No known methodological approach can deliver this combination of higher quality and speed. Here, we describe an Integrated Modeling PipEline for COVID Cure by Assessing Better LEads (IMPECCABLE) that employs multiple methodological innovations to overcome this fundamental limitation. We also describe the computational framework that we have developed to support these innovations at scale, and characterize the performance of this framework in terms of throughput, peak performance, and scientific results. We show that individual workflow components deliver 100 × to 1000 × improvement over traditional methods, and that the integration of methods, supported by scalable infrastructure, speeds up drug discovery by orders of magnitudes. IMPECCABLE has screened ∼ 1011 ligands and has been used to discover a promising drug candidate. These capabilities have been used by the US DOE National Virtual Biotechnology Laboratory and the EU Centre of Excellence in Computational Biomedicine.
Aymen Alsaadi, Dario Alfè, Yadu N. Babuji, Agastya Bhati, Ben Blaiszik, Alex Brace, Thomas S. Brettin, Kyle Chard, Ryan Chard, Austin Clyde, Peter V. Coveney, Ian T. Foster, Tom Gibbs, Shantenu Jha, Kristopher Keipert, Dieter Kranzlmüller, Thorsten Kurth, Hyungro Lee, Zhuozhao Li, Gerald Mathias, André Merzky, Alexander Partin, Arvind Ramanathan, Ashka Shah, Abraham C. Stern, Rick L. Stevens, Mikhail Titov, Anda Trifan, Aristeidis Tsaris, Matteo Turilli, Huub J. J. Van Dam, Shunzhou Wan, David Wifling, Junqi Yin
ICPP27
2021 AgEBO-tabular: joint neural architecture and hyperparameter search with autotuned data-parallel training for tabular data
abstract
Developing high-performing predictive models for large tabular data sets is a challenging task. Neural architecture search (NAS) is an AutoML approach that generates and evaluates multiple neural networks with different architectures concurrently to automatically discover an high performing model. A key issue in NAS, particularly for large data sets, is the large computation time required to evaluate each generated architecture. While data-parallel training has the potential to address this issue, a straightforward approach can result in significant loss of accuracy. To that end, we develop AgEBO-Tabular, which combines Aging Evolution (AE) to search over neural architectures and asynchronous Bayesian optimization (BO) to search over hyperparameters to adapt data-parallel training. We evaluate the efficacy of our approach on two large predictive modeling tabular data sets from the Exascale Computing Project-CANcer Distributed Learning Environment (ECP-CANDLE).
Romain Egele, Prasanna Balaprakash, Isabelle Guyon, Venkatram Vishwanath, Fangfang Xia, Rick L. Stevens, Zhengying Liu
SC6
2021 A genomic data resource for predicting antimicrobial resistance from laboratory-derived antimicrobial susceptibility phenotypes
abstract
Antimicrobial resistance (AMR) is a major global health threat that affects millions of people each year. Funding agencies worldwide and the global research community have expended considerable capital and effort tracking the evolution and spread of AMR by isolating and sequencing bacterial strains and performing antimicrobial susceptibility testing (AST). For the last several years, we have been capturing these efforts by curating data from the literature and data resources and building a set of assembled bacterial genome sequences that are paired with laboratory-derived AST data. This collection currently contains AST data for over 67 000 genomes encompassing approximately 40 genera and over 100 species. In this paper, we describe the characteristics of this collection, highlighting areas where sampling is comparatively deep or shallow, and showing areas where attention is needed from the research community to improve sampling and tracking efforts. In addition to using the data to track the evolution and spread of AMR, it also serves as a useful starting point for building machine learning models for predicting AMR phenotypes. We demonstrate this by describing two machine learning models that are built from the entire dataset to show where the predictive power is comparatively high or low. This AMR metadata collection is freely available and maintained on the Bacterial and Viral Bioinformatics Center (BV-BRC) FTP site ftp://ftp.bvbrc.org/RELEASE_NOTES/PATRIC_genomes_AMR.txt.
Margo VanOeffelen, Marcus Nguyen, Derya Aytan-Aktug, Thomas S. Brettin, Emily M. Dietrich, Ron Kenyon, Dustin Machi, Chunhong Mao, Robert Olson, Gordon D. Pusch, Maulik Shukla, Rick L. Stevens, Veronika Vonstein, Andrew S. Warren, Alice R. Wattam, Hyun Seung Yoo, James J. Davis 0002
Briefings Bioinform.12
2021 Learning curves for drug response prediction in cancer cell lines
abstract
BACKGROUND: Motivated by the size and availability of cell line drug sensitivity data, researchers have been developing machine learning (ML) models for predicting drug response to advance cancer treatment. As drug sensitivity studies continue generating drug response data, a common question is whether the generalization performance of existing prediction models can be further improved with more training data. METHODS: We utilize empirical learning curves for evaluating and comparing the data scaling properties of two neural networks (NNs) and two gradient boosting decision tree (GBDT) models trained on four cell line drug screening datasets. The learning curves are accurately fitted to a power law model, providing a framework for assessing the data scaling behavior of these models. RESULTS: The curves demonstrate that no single model dominates in terms of prediction performance across all datasets and training sizes, thus suggesting that the actual shape of these curves depends on the unique pair of an ML model and a dataset. The multi-input NN (mNN), in which gene expressions of cancer cells and molecular drug descriptors are input into separate subnetworks, outperforms a single-input NN (sNN), where the cell and drug features are concatenated for the input layer. In contrast, a GBDT with hyperparameter tuning exhibits superior performance as compared with both NNs at the lower range of training set sizes for two of the tested datasets, whereas the mNN consistently performs better at the higher range of training sizes. Moreover, the trajectory of the curves suggests that increasing the sample size is expected to further improve prediction scores of both NNs. These observations demonstrate the benefit of using learning curves to evaluate prediction models, providing a broader perspective on the overall data scaling characteristics. CONCLUSIONS: A fitted power law learning curve provides a forward-looking metric for analyzing prediction performance and can serve as a co-design tool to guide experimental biologists and computational scientists in the design of future experiments in prospective research studies.
Alexander Partin, Thomas S. Brettin, Yvonne A. Evrard, Yitan Zhu, Hyun Seung Yoo, Fangfang Xia, Songhao Jiang, Austin Clyde, Maulik Shukla, Michael Fonstein, James H. Doroshow, Rick L. Stevens
BMC Bioinform.12
2020 Overview of HPC and AI Computing for COVID-19 in the US
abstract
In this talk I'll describe some of the ongoing work in the US applying HPC and AI to COVID-19 related research. I will discuss two activities launched in the spring of 2020, the first is the COVID-19 HPC consortium that joins US supercomputing centers, computing and technology vendors and federal agencies to provide HPC cycles to the SARS-CoV-2/COVID-19 research community and to streamline access to resources via a single proposal mechanism. Currently the HPC consortium has 40 members, access to over 136K nodes, 5M CPUS and 50K GPUs totaling more than 558 Petaflops and is supporting 66 peer reviewed research projects related to the pandemic. The second topic I'll discuss is the nine DOE laboratory collaboration (Argonne, Berkeley, Brookhaven, Oak Ridge, Livermore, Los Alamos, Pacific Northwest, Sandia and SLAC) formed to apply advanced computing to the problem of developing molecular therapeutics for COVID-19. This project is one of part of the DOE sponsored National Virtual BioTechnology Laboratory (NVBTL) formed to coordinate national laboratory efforts related to the pandemic. For each of these I'll give a brief overview of the science, the state of play and how HPC and AI are being used and progress towards solutions.
Rick L. Stevens
PACT1
2019 Performance, Energy, and Scalability Analysis and Improvement of Parallel Cancer Deep Learning CANDLE Benchmarks
abstract
Training scientific deep learning models requires the significant compute power of high-performance computing systems. In this paper, we analyze the performance characteristics of the benchmarks from the exploratory research project CANDLE (Cancer Distributed Learning Environment) with a focus on the hyperparameters epochs, batch sizes, and learning rates. We present the parallel methodology that uses the distributed deep learning framework Horovod to parallelize the CANDLE benchmarks. We then use scaling strategies for both epochs and batch size with linear learning rate scaling to investigate how they impact the execution time and accuracy as well as the power, energy, and scalability of the parallel CANDLE benchmarks under conditions of strong scaling and weak scaling on the IBM Power9 heterogeneous system Summit at Oak Ridge National Laboratory and the Cray XC40 Theta at Argonne National Laboratory. This study provides insights into how to set the proper numbers of epochs, batch sizes, and compute resources for these benchmarks to preserve the high accuracy and to reduce the execution time of the benchmarks. We identify the data-loading performance bottleneck and then improve the performance and energy for better scalability. Results with the modified benchmarks on Summit indicate up to 78.25% in performance improvement and up to 78% in energy saving under strong scaling on up to 384 GPUs, and up to 79.5% in performance improvement and up to 77.11% in energy saving under weak scaling on up to 3,072 GPUs. On Theta, we achieve up to 45.22% performance improvement and up to 41.78% in energy saving under strong scaling on up to 384 nodes. Moreover, the modification dramatically reduces the broadcast overhead.
Xingfu Wu, Valerie Taylor 0001, Justin M. Wozniak, Rick L. Stevens, Thomas S. Brettin, Fangfang Xia
ICPP4
2019 Scalable reinforcement-learning-based neural architecture search for cancer deep learning research
abstract
Cancer is a complex disease, the understanding and treatment of which are being aided through increases in the volume of collected data and in the scale of deployed computing power. Consequently, there is a growing need for the development of data-driven and, in particular, deep learning methods for various tasks such as cancer diagnosis, detection, prognosis, and prediction. Despite recent successes, however, designing high-performing deep learning models for nonimage and nontext cancer data is a time-consuming, trial-and-error, manual task that requires both cancer domain and deep learning expertise. To that end, we develop a reinforcement-learning-based neural architecture search to automate deep-learning-based predictive model development for a class of representative cancer data. We develop custom building blocks that allow domain experts to incorporate the cancer-data-specific characteristics. We show that our approach discovers deep neural network architectures that have significantly fewer trainable parameters, shorter training time, and accuracy similar to or higher than those of manually designed architectures. We study and demonstrate the scalability of our approach on up to 1,024 Intel Knights Landing nodes of the Theta supercomputer at the Argonne Leadership Computing Facility.
Prasanna Balaprakash, Romain Egele, Misha Salim, Stefan M. Wild, Venkatram Vishwanath, Fangfang Xia, Thomas S. Brettin, Rick L. Stevens
SC8
2019 PATRIC as a unique resource for studying antimicrobial resistance
abstract
The Pathosystems Resource Integration Center (PATRIC, www.patricbrc.org) is designed to provide researchers with the tools and services that they need to perform genomic and other 'omic' data analyses. In response to mounting concern over antimicrobial resistance (AMR), the PATRIC team has been developing new tools that help researchers understand AMR and its genetic determinants. To support comparative analyses, we have added AMR phenotype data to over 15 000 genomes in the PATRIC database, often assembling genomes from reads in public archives and collecting their associated AMR panel data from the literature to augment the collection. We have also been using this collection of AMR metadata to build machine learning-based classifiers that can predict the AMR phenotypes and the genomic regions associated with resistance for genomes being submitted to the annotation service. Likewise, we have undertaken a large AMR protein annotation effort by manually curating data from the literature and public repositories. This collection of 7370 AMR reference proteins, which contains many protein annotations (functional roles) that are unique to PATRIC and RAST, has been manually curated so that it projects stably across genomes. The collection currently projects to 1 610 744 proteins in the PATRIC database. Finally, the PATRIC Web site has been expanded to enable AMR-based custom page views so that researchers can easily explore AMR data and design experiments based on whole genomes or individual genes.
Dionysios A. Antonopoulos, Rida Assaf, Ramy K. Aziz, Thomas S. Brettin, Christopher Bun, Neal Conrad, James J. Davis 0002, Emily M. Dietrich, Terry Disz, Svetlana Gerdes, Ron Kenyon, Dustin Machi, Chunhong Mao, Daniel E. Murphy-Olson, Eric K. Nordberg, Gary J. Olsen, Robert Olson, Ross A. Overbeek, Bruce D. Parrello, Gordon D. Pusch, John Santerre, Maulik Shukla, Rick L. Stevens, Margo VanOeffelen, Veronika Vonstein, Andrew S. Warren, Alice R. Wattam, Fangfang Xia, Hyun Seung Yoo
Briefings Bioinform.23
2018 CANDLE/Supervisor: a workflow framework for machine learning applied to cancer research
abstract
BACKGROUND: Current multi-petaflop supercomputers are powerful systems, but present challenges when faced with problems requiring large machine learning workflows. Complex algorithms running at system scale, often with different patterns that require disparate software packages and complex data flows cause difficulties in assembling and managing large experiments on these machines. RESULTS: This paper presents a workflow system that makes progress on scaling machine learning ensembles, specifically in this first release, ensembles of deep neural networks that address problems in cancer research across the atomistic, molecular and population scales. The initial release of the application framework that we call CANDLE/Supervisor addresses the problem of hyper-parameter exploration of deep neural networks. CONCLUSIONS: Initial results demonstrating CANDLE on DOE systems at ORNL, ANL and NERSC (Titan, Theta and Cori, respectively) demonstrate both scaling and multi-platform execution.
Justin M. Wozniak, Rajeev Jain, Prasanna Balaprakash, Jonathan Ozik, Nicholson T. Collier, John Bauer, Fangfang Xia, Thomas S. Brettin, Rick L. Stevens, Jamaludin Mohd-Yusof, Cristina Garcia-Cardona, Brian Van Essen, Matt Baughman
BMC Bioinform.9
2018 Predicting tumor cell line response to drug pairs with deep learning
abstract
BACKGROUND: The National Cancer Institute drug pair screening effort against 60 well-characterized human tumor cell lines (NCI-60) presents an unprecedented resource for modeling combinational drug activity. RESULTS: We present a computational model for predicting cell line response to a subset of drug pairs in the NCI-ALMANAC database. Based on residual neural networks for encoding features as well as predicting tumor growth, our model explains 94% of the response variance. While our best result is achieved with a combination of molecular feature types (gene expression, microRNA and proteome), we show that most of the predictive power comes from drug descriptors. To further demonstrate value in detecting anticancer therapy, we rank the drug pairs for each cell line based on model predicted combination effect and recover 80% of the top pairs with enhanced activity. CONCLUSIONS: We present promising results in applying deep learning to predicting combinational drug response. Our feature analysis indicates screening data involving more cell lines are needed for the models to make better use of molecular features.
Fangfang Xia, Maulik Shukla, Thomas S. Brettin, Cristina Garcia-Cardona, Judith D. Cohn, Jonathan E. Allen, Sergei Maslov, Susan L. Holbeck, James H. Doroshow, Yvonne A. Evrard, Eric A. Stahlberg, Rick L. Stevens
BMC Bioinform.12
2017 Deep Learning in Cancer and Infectious Disease: Novel Driver Problems for Future HPC Architecture
abstract
The adoption of machine learning is proving to be an amazingly successful strategy in improving predictive models for cancer and infectious disease. In this talk I will discuss two projects my group is working on to advance biomedical research through the use of machine learning and HPC. In cancer, machine learning and in deep learning in particular, is used to advance our ability to diagnosis and classify tumors. Recently demonstrated automated systems are routinely out performing human expertise. Deep learning is also being used to predict patient response to cancer treatments and to screen for new anti-cancer compounds. In basic cancer research its being use to supervise large-scale multi-resolution molecular dynamics simulations used to explore cancer gene signaling pathways. In public health it's being used to interpret millions of medical records to identify optimal treatment strategies. In infectious disease research machine learning methods are being used to predict antibiotic resistance and to identify novel antibiotic resistance mechanisms that might be present. More generally machine learning is emerging as a general tool to augment and extend mechanistic models in biology and many other fields. It's becoming an important component of scientific workloads. From a computational architecture standpoint, deep neural network (DNN) based scientific applications have some unique requirements. They require high compute density to support matrix-matrix and matrix-vector operations, but they rarely require 64bit or even 32bits of precision, thus architects are creating new instructions and new design points to accelerate training. Most current DNNs rely on dense fully connected networks and convolutional networks and thus are reasonably matched to current HPC accelerators. However future DNNs may rely less on dense communication patterns. Like simulation codes, power efficient DNNs require high-bandwidth memory be physically close to arithmetic units to reduce costs of data motion and a high-bandwidth communication fabric between (perhaps modest scale) groups of processors to support network model parallelism. DNNs in general do not have good strong scaling behavior, so to fully exploit large-scale parallelism they rely on a combination of model, data and search parallelism. Deep learning problems also require large-quantities of training data to be made available or generated at each node, thus providing opportunities for NVRAM. Discovering optimal deep learning models often involves a large-scale search of hyperparameters. It's not uncommon to search a space of tens of thousands of model configurations. Naïve searches are outperformed by various intelligent searching strategies, including new approaches that use generative neural networks to manage the search space. HPC architectures that can support these large-scale intelligent search methods as well as efficient model training are needed.
Rick L. Stevens
HPDC1
2012 Real Time Metagenomics: Using k-mers to annotate metagenomes
abstract
Abstract Summary: Annotation of metagenomes involves comparing the individual sequence reads with a database of known sequences and assigning a unique function to each read. This is a time-consuming task that is computationally intensive (though not computationally complex). Here we present a novel approach to annotate metagenomes using unique k-mer oligopeptide sequences from 7 to 12 amino acids long. We demonstrate that k-mer-based annotations are faster and approach the sensitivity and precision of blastx-based annotations without loosing accuracy. A last-common ancestor approach was also developed to describe the members of the community. Availability and implementation: This open-source application was implemented in Perl and can be accessed via a user-friendly website at http://edwards.sdsu.edu/rtmg. In addition, code to access the annotation servers is available for download from http://www.theseed.org/. FIGfams and k-mers are available for download from ftp://ftp.theseed.org/FIGfams/. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Robert A. Edwards, Robert Olson, Terry Disz, Gordon D. Pusch, Veronika Vonstein, Rick L. Stevens, Ross A. Overbeek
Bioinform.6
2010 Accessing the SEED genome databases via Web services API: tools for programmers
abstract
BACKGROUND: The SEED integrates many publicly available genome sequences into a single resource. The database contains accurate and up-to-date annotations based on the subsystems concept that leverages clustering between genomes and other clues to accurately and efficiently annotate microbial genomes. The backend is used as the foundation for many genome annotation tools, such as the Rapid Annotation using Subsystems Technology (RAST) server for whole genome annotation, the metagenomics RAST server for random community genome annotations, and the annotation clearinghouse for exchanging annotations from different resources. In addition to a web user interface, the SEED also provides Web services based API for programmatic access to the data in the SEED, allowing the development of third-party tools and mash-ups. RESULTS: The currently exposed Web services encompass over forty different methods for accessing data related to microbial genome annotations. The Web services provide comprehensive access to the database back end, allowing any programmer access to the most consistent and accurate genome annotations available. The Web services are deployed using a platform independent service-oriented approach that allows the user to choose the most suitable programming platform for their application. Example code demonstrate that Web services can be used to access the SEED using common bioinformatics programming languages such as Perl, Python, and Java. CONCLUSIONS: We present a novel approach to access the SEED database. Using Web services, a robust API for access to genomics data is provided, without requiring large volume downloads all at once. The API ensures timely access to the most current datasets available, including the new genomes as soon as they come online.
Terry Disz, Sajia Akhter, Daniel Cuevas, Robert Olson, Ross A. Overbeek, Veronika Vonstein, Rick L. Stevens, Robert A. Edwards
BMC Bioinform.7
2010 Oscillation in a network model of neocortex
Jennifer Dwyer, Hyong Lee, Amber Martell, Rick L. Stevens, Mark Hereld, Wim van Drongelen
Neurocomputing4
2009 Oscillation in a network model of neocortex
Wim van Drongelen, Hyong Lee, Amber Martell, Jennifer Dwyer, Rick L. Stevens, Mark Hereld
ESANN5
2009 The GAAS Metagenomic Tool and Its Estimations of Viral and Microbial Average Genome Size in Four Major Biomes
abstract
Metagenomic studies characterize both the composition and diversity of uncultured viral and microbial communities. BLAST-based comparisons have typically been used for such analyses; however, sampling biases, high percentages of unknown sequences, and the use of arbitrary thresholds to find significant similarities can decrease the accuracy and validity of estimates. Here, we present Genome relative Abundance and Average Size (GAAS), a complete software package that provides improved estimates of community composition and average genome length for metagenomes in both textual and graphical formats. GAAS implements a novel methodology to control for sampling bias via length normalization, to adjust for multiple BLAST similarities by similarity weighting, and to select significant similarities using relative alignment lengths. In benchmark tests, the GAAS method was robust to both high percentages of unknown sequences and to variations in metagenomic sequence read lengths. Re-analysis of the Sargasso Sea virome using GAAS indicated that standard methodologies for metagenomic analysis may dramatically underestimate the abundance and importance of organisms with small genomes in environmental systems. Using GAAS, we conducted a meta-analysis of microbial and viral average genome lengths in over 150 metagenomes from four biomes to determine whether genome lengths vary consistently between and within biomes, and between microbial and viral communities from the same environment. Significant differences between biomes and within aquatic sub-biomes (oceans, hypersaline systems, freshwater, and microbialites) suggested that average genome length is a fundamental property of environments driven by factors at the sub-biome level. The behavior of paired viral and microbial metagenomes from the same environment indicated that microbial and viral average genome sizes are independent of each other, but indicative of community responses to stressors and environmental conditions.
Florent E. Angly, Dana Willner, Alejandra Prieto-Davó, Robert A. Edwards, Robert Schmieder, Rebecca Vega-Thurber, Dionysios A. Antonopoulos, Katie Barott, Matthew T. Cottrell, Christelle Desnues, Elizabeth A. Dinsdale, Mike Furlan, Matthew Haynes, Matthew R. Henn, Yongfei Hu, David L. Kirchman, Tracey McDole, John D. McPherson, Folker Meyer, R. Michael Miller, Egbert Mundt, Robert K. Naviaux, Beltran Rodriguez-Mueller, Rick L. Stevens, Linda Wegley, Lixin Zhang 0006, Baoli Zhu, Forest Rohwer
PLoS Comput. Biol.24
2008 The metagenomics RAST server - a public resource for the automatic phylogenetic and functional analysis of metagenomes
abstract
BACKGROUND: Random community genomes (metagenomes) are now commonly used to study microbes in different environments. Over the past few years, the major challenge associated with metagenomics shifted from generating to analyzing sequences. High-throughput, low-cost next-generation sequencing has provided access to metagenomics to a wide range of researchers. RESULTS: A high-throughput pipeline has been constructed to provide high-performance computing to all researchers interested in using metagenomics. The pipeline produces automated functional assignments of sequences in the metagenome by comparing both protein and nucleotide databases. Phylogenetic and functional summaries of the metagenomes are generated, and tools for comparative metagenomics are incorporated into the standard views. User access is controlled to ensure data privacy, but the collaborative environment underpinning the service provides a framework for sharing datasets between multiple users. In the metagenomics RAST, all users retain full control of their data, and everything is available for download in a variety of formats. CONCLUSION: The open-source metagenomics RAST service provides a new paradigm for the annotation and analysis of metagenomes. With built-in support for multiple data sources and a back end that houses abstract data types, the metagenomics RAST is stable, extensible, and freely available to all researchers. This service has removed one of the primary bottlenecks in metagenome sequence analysis - the availability of high-performance computing for annotating the data. http://metagenomics.nmpdr.org.
Folker Meyer, Daniel Paarmann, Mark D'Souza, Robert Olson, Elizabeth M. Glass, Michael Kubal, Tobias Paczian, Alexis A. Rodriguez, Rick L. Stevens, Andreas Wilke, Jared Wilkening, Robert A. Edwards
BMC Bioinform.9
2007 ParalleX: A Study of A New Parallel Computation Model
abstract
This paper proposes the study of a new computation model that attempts to address the underlying sources of performance degradation (e.g. latency, overhead, and starvation) and the difficulties of programmer productivity (e.g. explicit locality management and scheduling, performance tuning, fragmented memory, and synchronous global barriers) to dramatically enhance the broad effectiveness of parallel processing for high end computing. In this paper, we present the progress of our research on a parallel programming and execution model - mainly, ParalleX. We describe the functional elements of ParalleX, one such model being explored as part of this project. We also report our progress on the development and study of a subset of ParalleX $the LITL-X at University of Delaware. We then present a novel architecture model - Gilgamesh II - as a ParalleX processing architecture. A design point study of Gilgamesh II and the architecture concept strategy are presented.
Guang R. Gao, Thomas L. Sterling, Rick L. Stevens, Mark Hereld, Weirong Zhu
IPDPS3
2006 Extending Multicast Communications by Hybrid Overlay Network
abstract
Recently scalable streaming applications have been developed by the new emerging networking technologies. One interesting challenge is how to build an efficient and reliable network topology over heterogeneous network systems. In this paper, we argue that hybrid network overlay, via both IP multicast and overlay mesh multicast, is an important technology for group communications. In this paper, we propose a hybrid network overlay construction with three performance goals: high performance (low end-to-end transmission latency, high bandwidth mesh overlay links), low end-to-end hop count and high reliability. This paper compares the performance penalties from IP multicast, overlay multicast and hybrid overlay multicast via various topology simulations.
Michael E. Papka, Rick L. Stevens
ICC3
2006 Hierarchical multithreading: programming model and system software
abstract
This paper addresses the underlying sources of performance degradation (e.g. latency, overhead, and starvation) and the difficulties of programmer productivity (e.g. explicit locality management and scheduling, performance tuning, fragmented memory, and synchronous global barriers) to dramatically enhance the broad effectiveness of parallel processing for high end computing. We are developing a hierarchical threaded virtual machine (HTVM) that defines a dynamic, multithreaded execution model and programming model, providing an architecture abstraction for HEC system software and tools development. We are working on a prototype language, LITL-X (pronounced "little-X") for latency intrinsic-tolerant language, which provides the application programmers with a powerful set of semantic constructs to organize parallel computations in a way that hides/manages latency and limits the effects of overhead. This is quite different from locality management, although the intent of both strategies is to minimize the effect of latency on the efficiency of computation. We work on a dynamic compilation and runtime model to achieve efficient LITL-X program execution. Several adaptive optimizations were studied. A methodology of incorporating domain-specific knowledge in program optimization was studied. Finally, we plan to implement our method in an experimental testbed for a HEC architecture and perform a qualitative and quantitative evaluation on selected applications
Guang R. Gao, Thomas L. Sterling, Rick L. Stevens, Mark Hereld, Weirong Zhu
IPDPS3
2006 CupHolder: A Multi-Person Interactive High-Resolution Workstation
abstract
Demand for high resolution visualization, large pixel real estate collaborative workspaces, and interactive computer interfaces continue to drive researchers to develop new physical portals connecting them to their computational tools, to their data, and to their colleagues around the world. In this paper we describe CupHolder, a high performance workstation designed to support interactive collaboration and research activities. It is configured to enable display of high resolution imagery while enabling a comfortable interactive environment for one to several co-located researchers. Moreover it is driven by a high performance commodity cluster that provides substantial local rendering muscle as well as a high performance interface to Grid-based computational tasks. CupHolder is comprised of commodity components. It is primarily novel because it represents an integration of these components into a new form factor that we believe is a useful precursor and testbed for future integrated workspaces.
Mark Hereld, Michael E. Papka, Justin Binns, Rick L. Stevens
VR4
2005 An Infrastructure of Network Services for Seamless Integration in Advanced Collaborative Computing Environments
abstract
Advanced collaborative computing environments are one of the most important tools for integrating high-performance computers and computations and for interacting with colleagues around the world. However, heterogeneous characteristics such as network transfer rates, computational abilities, and hierarchical systems make the seamless integration of distributed resources a challenge. In this paper, we argue that advanced collaborative computing environments need an infrastructure of network services to support distributed and quality guaranteed multimedia applications. Accordingly, we propose the design of network services for high-performance collaborative computing. We present a collaborative environment network service infrastructure (CENSI) to embed network services into various systems intelligently and elastically. We also discuss three management modules: a three-party matching module (resources, requests, and network services), a module for performance monitoring and evaluation of group communications, and a module for distribution topology analysis
Ivan R. Judson, Thomas D. Uram, S. Lefvert, Terry Disz, Michael E. Papka, Rick L. Stevens
CLUSTER7
2005 Performance Metrics of IP Multicast Sessions
abstract
Most research on design and implementation of multicast network systems has concentrated mainly on the capacities of network devices, without completely considering the impact from multicast sessions on network systems. This impact heavily depends on the activities during the life span of multicast sessions, such as the session initialization, communication, termination, and membership management, and on the influence within the overall architecture context of multicast sessions, such as routing devices and protocols. Based on the fundamental rules and principles of multicast sessions, this paper defines performance metrics for IP multicast sessions. The paper summarizes the metrics in four categories: latency, quality of service, group characteristics, and link bandwidth consumptions. Also presented are results from our simulations.
Michael E. Papka, Rick L. Stevens
ISM3
2005 Real Time Change Detection and Alerts from Highway Traffic Data
abstract
We developed a testbed containing: real time data from over 830 highway traffic sensors in the Chicago region, data about weather, and text data about events that might affect traffic. The goal was to detect in real time interesting changes in traffic conditions. Given the size and complexity of the data, we choose to build a large number of separate baseline models. We built a separate baseline for each hour in the day, for each day in the week, and for every 2 or 3 traffic sensors, resulting in over 42,000 separate baseline models. We also built a baseline engine to build the necessary baselines automatically. We modified an open source scoring engine to process in real time each new sensor reading, update the appropriate feature vectors, score the updated feature vectors using the baseline models, and send out real time alerts when deviations from the baselines were detected.
Robert L. Grossman, Michal Sabala, Anushka Anand, Steve Eick, Leland Wilkinson, John Chaves, Steve Vejcik, John F. Dillenburg, Peter C. Nelson, Doug Rorem, Javid Alimohideen, Jason Leigh, Michael E. Papka, Rick L. Stevens
SC15
2005 Perceptual photometric seamlessness in projection-based tiled displays
abstract
Arguably, the most vexing problem remaining for multi-projector displays is that of photometric (brightness) seamlessness within and across different projectors. Researchers have strived for strict photometric uniformity that achieves identical response at every pixel of the display. However, this goal typically results in displays with severely compressed dynamic range and poor image quality. In this article, we show that strict photometric uniformity is not a requirement for achieving photometric seamlessness. We introduce a general goal for photometric seamlessness by defining it as an optimization problem, balancing perceptual uniformity with display quality. Based on this goal, we present a new method to achieve perceptually seamless high quality displays. We first derive a model that describes the photometric response of projection-based displays. Then we estimate the model parameters and modify them using perception-driven criteria. Finally, we use the graphics hardware to reproject the image computed using the modified model parameters by manipulating only the projector inputs at interactive rates. Our method has been successfully demonstrated on three different practical display systems at Argonne National Laboratory, made of 2 × 2 array of four projectors, 2 × 3 array of six, projectors, and 3 × 5 array of fifteen projectors. Our approach is efficient, automatic and scalable---requiring only a digital camera and a photometer. To the best of our knowledge, this is the first approach and system that addresses the photometric variation problem from a perceptual stand point and generates truly seamless displays with high dynamic range.
Aditi Majumder, Rick L. Stevens
ACM Trans. Graph.2
2004 Capability matching of data streams with network services
abstract
Distributed computing middleware needs to support a wide range of resources, such as diverse software components, various hardware devices, and heterogeneous operating systems and architectures. Current technologies are unable to implement a maintenance-free platform to be compatible with such different computing environments. This situation is presenting an increasing challenge as Grid computing becomes more widespread. The infrastructure of network services (CENSA and CENSI) has been proposed to address this challenge. A seamless Grid computing environment, supported by network services, is composed of various streams such as data, video, audio, and text. We define a mathematical model of capability matching for three-party agreements: requests from users, resources, and network services. Based on the mathematical model, we provide a general approach for capability matching. We also present a new language schema for capability description. As an example, we embed the general matchmaker in the architecture of the access Grid. Several tests of accuracy and performance are discussed.
Ivan R. Judson, Thomas D. Uram, Terry Disz, Michael E. Papka, Rick L. Stevens
CCGRID6
2004 Isocoupling: Reusing Kernel Coupling Values to Predict the Performance of Parallel Applications
abstract
Summary form only given. Kernel coupling quantifies the interaction between adjacent and chains of kernels in an application. A kernel can be a loop, procedure or file. In our previous work, we used the kernel coupling values to identify how to combine the execution times of the individual kernels that compose the application to predict the execution time of the full application. The results of this previous work using the NAS Parallel Benchmark SP demonstrated that the use of coupling values resulted in very good predictions with average errors in the range of only 1.18% in contrast to simply summing the execution times of the kernels that resulted in average errors in the range of 20.54%. The major concern with the coupling values is the fact that values are needed for each different problem size, number of processors and machine. We explore the ability to reuse coupling values. In particular, we explore the reuse in terms of the three dimensional space consisting of the following axes: number of processors, problem size and system architecture. The experimental results indicate that when considering parallel systems, with increasing number of processors and problem sizes, we found clear transitions with the coupling values resulting in the ability to reuse values. Further, reusing coupling values is feasible on classes of systems such as clusters, distributed shared memory and other distributed memory systems.
Xingfu Wu, Jonathan Geisler, Rick L. Stevens
IPDPS3
2004 Simulation of neocortical epileptiform activity using parallel computing
Wim van Drongelen, Hyong Lee, Mark Hereld, Matthew Cohoon, Frank Elsen, Michael E. Papka, Rick L. Stevens
Neurocomputing8
2004 Color Nonuniformity in Projection-Based Displays: Analysis and Solutions
abstract
Large-area displays made up of several projectors show significant variation in color. In this paper, we identify different projector parameters that cause the color variation and study their effects on the luminance and chrominance characteristics of the display. This work leads to the realization that luminance varies significantly within and across projectors, while chrominance variation is relatively small, especially across projectors of same model. To address this situation, we present a method to achieve luminance matching across all pixels of a multiprojector display that results in photometrically uniform displays. We use a camera as a measurement device for this purpose. Our method comprises a one-time calibration step that generates a per channel per projector luminance attenuation map (LAM), which is then used to correct any image projected on the display at interactive rates on commodity graphics hardware. To the best of our knowledge, this is the first effort to match luminance across all the pixels of a multiprojector display.
Aditi Majumder, Rick L. Stevens
IEEE Trans. Vis. Comput. Graph.2
2002 Using Kernel Couplings to Predict Parallel Application Performance
abstract
Performance models provide significant insight into the performance relationships between an application and the system used for execution. The major obstacle to developing performance models is the lack of knowledge about the performance relationships between the different functions that compose an application. This paper addresses the issue by using a coupling parameter, which quantifies the interaction between kernels, to develop performance predictions. The results, using three NAS parallel application benchmarks, indicate that the predictions using the coupling parameter were greatly improved over a traditional technique of summing the execution times of the individual kernels in an application. In one case the coupling predictor had less than 1% relative error in contrast the summation methodology that had over 20% relative error. Further, as the problem size and number of processors scale, the coupling values go through a finite number of major value changes that is dependent on the memory subsystem of the processor architecture.
Valerie Taylor 0001, Xingfu Wu, Jonathan Geisler, Rick L. Stevens
HPDC4
2002 LAM: luminance attenuation map for photometric uniformity in projection based displays
abstract
Large-area multi-projector displays show significant spatial variation in color, both within a single projector's field of view and across different projectors. Recent research in this area has shown that the color variation is primarily due to luminance variation. Luminance varies within a single projector's field of view, across different brands of projectors and with the variation in projector parameters. Luminance variation is also introduced by overlap between adjacent projectors. On the other hand, chrominance remains constant throughout a projector's field of view and varies little with the change in projector parameters, especially for projectors of the same brand. Hence, matching luminance response of all the pixels of a multi-projector display should help us to achieve photometric uniformity.In this paper, we present a method to do a per channel per pixel luminance matching. Our method consists of a one-time calibration procedure when a luminance attenuation map (LAM) is generated. This LAM is then used to correct any image to achieve photometric uniformity. In the one-time calibration step, we first use a camera to measure the per channel luminance response of a multi-projector display and find the pixel with the most luminance response. Then, for each projector, we generate a per channel LAM that assigns a weight to every pixel of the projector to scale the luminance response of that pixel to match with the most limited response. This LAM is then used to attenuate any image projected by the projector.This method can be extended to do the image correction in real time on traditional graphics pipeline by using alpha blending and color look-up-tables. To the best of our knowledge, this is the first effort to match luminance across all the pixels of a multi-projector display. Our results show that luminance matching can indeed achieve photometric uniformity.
Aditi Majumder, Rick L. Stevens
VRST2
2000 Prophesy: An Infrastructure for Analyzing and Modeling the Performance of Parallel and Distributed Applications
abstract
Efficient execution of applications requires insight into how the system features impact the performance of the application. For distributed systems, the task of gaining this insight is complicated by the complexity of the system features. This insight generally results from significant experimental analysis and possibly the development of performance models. This paper presents the Prophesy project, an infrastructure that aids in gaining this needed insight based upon experience. The core component of Prophesy is a relational database that allows for the recording of performance data, system features and application details.
Xingfu Wu, Valerie Taylor 0001, Jonathan Geisler, Zhiling Lan, Rick L. Stevens, Mark Hereld, Ivan R. Judson
HPDC6
2000 The Ten Hottest Topics in Parallel and Distributed Computing for the Next Millennium
Ian T. Foster, David E. Culler, Deborah Estrin, Harvey B. Newman, Rick L. Stevens
IPDPS5
1999 Capstone Address: ActiveSpaces - The Access Grid, Active Mural and Advanced Visualization Systems
Rick L. Stevens
IEEE Visualization1
1997 Performance model of the Argonne Voyager multimedia server
abstract
The Argonne Voyager Multimedia Server is being developed in the Futures Lab of the Mathematics and Computer Science Division at Argonne National Laboratory. As a network based service for recording and playing multimedia streams, it is important that the Voyager system be capable of sustaining certain minimal levels of performance in order for it to be a viable system. In this article, we examine the performance characteristics of the server. As we examine the architecture of the system, we try to determine where bottlenecks lie, show actual vs potential performance, and recommend areas for improvement through custom architectures and system tuning.
Terry Disz, Robert Olson, Rick L. Stevens
ASAP3
1997 The Argonne Voyager Multimedia Server
abstract
With the growing presence of multimedia-enabled systems, we will see an integration of collaborative computing concepts into future scientific and technical workplaces. Desktop teleconferencing is common today, while more complex teleconferencing technology that relies on the availability of multipoint-enabled tools is starting to become available on PCs. A critical problem when using these collaborative tools is archiving multistream, multipoint meetings and making the content available to others. Ideally, one would like the ability to capture, record, play back, index, annotate, and distribute multimedia stream data as easily as we currently handle text or still-image data. The Argonne Voyager project is exploring and developing media server technology needed to provide such a flexible, virtual multipoint recording/playback capability. In this article we describe the motivating requirements, architecture, implementation, operation, performance, and related work.
Terry Disz, Ivan R. Judson, Robert Olson, Rick L. Stevens
HPDC4
1997 Parallel Molecular Dynamics: Implications for Massively Parallel Machines
Valerie Taylor 0001, Rick L. Stevens, Kathryn E. Arnold
J. Parallel Distributed Comput.2
1996 A Decomposition Method For Efficient Use Of Distributed Supercomputers For Finite Element Applications
abstract
The interconnection of geographically distributed supercomputers via highspeed networks makes available the needed compute power for large-scale scientific applications, such as finite element applications. In this paper we propose a two-level data decomposition method for efficient execution of finite element applications on a network of supercomputers. Our method exploits the following features that may be different for each supercomputer in the system: processor speed, number of processors used from each supercomputer, local network performance, wide area network performance and wide area topology. Preliminary experiments involving a nonlinear, finite element application executed on a network of two supercomputers, one located at Argonne National Laboratory and the other one at the Cornell Theory Center, demonstrate a 20% reduction in execution time when the proposed decomposition is used as compared with naively applying conventional decompositions that are applicable to single supercomputers.
Valerie Taylor 0001, Jian Chen 0044, Thomas Canfield, Rick L. Stevens
ASAP4
1996 Tools for Distributed Collaborative Environments: A Research Agenda
abstract
Argues that future computing environments will be collaboration-oriented, globally distributed and computation/information-rich. These environments will be accessed via multiple interface devices. As we move from a desktop-centric computing model to a network-centric model, new approaches in the way software and data are handled will need to be developed. In this article, we outline the requirements for enabling the technological infrastructure, describe some first steps that we have taken toward building this infrastructure, and sketch directions for future development.
Ian T. Foster, Michael E. Papka, Rick L. Stevens
HPDC3
1996 UbiWorld: An Environment Integrating Virtual Reality, Supercomputing and Design
abstract
Summary form only given. UbiWorld is a concept that ties together the notion of ubiquitous computing (Ubicomp) with that of using virtual reality for rapid prototyping. The goal is to develop an environment where one can explore Ubicomp-type concepts without having to build real Ubicomp hardware. The basic notion is to extend object models in a virtual world using distributed wide-area heterogeneous computing technology to provide complex networking and processing capabilities to virtual reality objects. Starting with the CAVE/sup TM/ family of display devices, we integrate tools for the construction of 30 objects into the existing library. Then, using these objects as models, we can embed new information technology within them. The plan is then to couple the virtual objects to remote computers via fine-grain heterogenous computing technology to provide Ubicomp behavior and functionality to the modeled objects. We tightly couple the process-defined behavior with the 30 objects and place these objects into rooms, creating a shared virtual world where users can experiment with using the virtual devices. Each object in the world has its behavior controlled by a program running some place on the network. This behavior could be one that in the real object would be provided by a local computer or by a combination of local computer and network connection to remote processors or databases. These "behavior" processes are able to communicate with each other using a shared protocol (UbiWorldcomm). These object also react and are influenced directly by interactions with the virtual world and users.
Michael E. Papka, Rick L. Stevens
HPDC2
1995 Distributed Information Management in the National HPCC Software Exchange
Shirley Browne, Jack J. Dongarra, Geoffrey C. Fox, Kenneth A. Hawick, Ken Kennedy, Rick L. Stevens, Robert Olson, Tom Rowan
SC6
1994 Multimedia Supercomputing: The Use of Supercomputers to Drive High-Performance Multimedia Systems and Virtual Environments
abstract
Summary form only given, as follows. The author describes a project underway at Argonne National Laboratory (ANL) to merge multimedia capability with large-scale parallel supercomputing. He describes the IBM SP parallel system currently installed at ANL and the project to extend its capability with ATM networks and video processor hardware/software, in order to provide scientific computing users with state-of-the-art multimedia capability. We are developing ways that video and audio can be used to augment traditional scientific visualization, including ideas for multimedia I/O libraries that will enable existing programs to directly produce video output for visualization and/or debugging. The author also describes work to develop a collaborative virtual environment based upon linking the CAVE virtual reality environment, via high-performance networking, to the IBM SP and connecting this system to similar facilities at the Electronic Visualization Laboratory at UIC and the National Center for Supercomputing Applications at UIC. These linked virtual environments will be used for collaborative scientific visualization and experiments in the development of high-end collaborative tools.>
Rick L. Stevens
HPDC1
1994 High-performance computing and communications
Rick L. Stevens
Future Gener. Comput. Syst.1
1990 Automated Reasoning Contributed to Mathematics and Logic
Larry Wos, Steven K. Winker, William McCune, Ross A. Overbeek, Ewing L. Lusk, Rick L. Stevens, Ralph M. Butler
CADE6
1990 Parallel Programming with Algorithmic Motifs
Ian T. Foster, Rick L. Stevens
ICPP (2)2
1989 Solving Open Problems in Right Alternative Rings with Z-Module Reasoning
Tie-Cheng Wang, Rick L. Stevens
J. Autom. Reason.2
1988 Challenge Problems from Nonassociative Rings for Theorem Provers
Rick L. Stevens
CADE1
1987 Some Experiments in Nonassociative Ring Theory with an Automated Theorem Prover
Rick L. Stevens
J. Autom. Reason.1