Robert L. Grossman

dblp:g/RobertLGrossman · DBLP profile ↗
← Back
68ranked-venue papers
22as first author
7since 2021 · last 2024
0000-0003-3741-5739ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 9 first-author · 1 since 2021Databases, data management, data science and information retrieval · 15 · 7 first-authorApplied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 2 since 2021Theory of computation · 9 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-authorComputer networks · 2 · 1 first-authorSecurity and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2024 Enhancing Instance-Level Image Classification with Set-Level Labels
abstract
Instance-level image classification tasks have traditionally relied on single-instance labels to train models, e.g., few-shot learning and transfer learning. However, set-level coarse-grained labels that capture relationships among instances can provide richer information in real-world scenarios. In this paper, we present a novel approach to enhance instance-level image classification by leveraging set-level labels. We provide a theoretical analysis of the proposed method, including recognition conditions for fast excess risk rate, shedding light on the theoretical foundations of our approach. We conducted experiments on two distinct categories of datasets: natural image datasets and histopathology image datasets. Our experimental results demonstrate the effectiveness of our approach, showcasing improved classification performance compared to traditional single-instance label-based methods. Notably, our algorithm achieves 13\% improvement in classification accuracy compared to the strongest baseline on the histopathology image classification benchmarks. Importantly, our experimental findings align with the theoretical analysis, reinforcing the robustness and reliability of our proposed method. This work bridges the gap between instance-level and set-level image classification, offering a promising avenue for advancing the capabilities of image classification models with set-level coarse-grained labels.
Aly Azeem Khan, Yuxin Chen 0001, Robert L. Grossman
ICLR4
2023 CNT: Semi-Automatic Translation from CWL to Nextflow for Genomic Workflows
abstract
With the rise of advanced workflow languages for scientific computations, Nextflow has gained increased attention from the bioinformatics community. Nextflow offers native support for advanced parallelism, which can greatly enhance resource utilization and throughput. Still, a significant portion of bioinformatics workflows are developed with the Common Workflow Language (CWL). Transitioning from CWL to Nextflow poses a significant challenge due to the differences in programming models, scripting language compatibilities, and the prerequisite for in-depth knowledge in both languages. To address this challenge, we present CNT, a novel, semi-automated translator converting CWL workflows into Nextflow ones. At its core, CNT uses an automated translation mechanism that converts the CommandLineTool, the most basic unit of CWL, into Nextflow's Process class. This component integrates tool-level conversion, graph dependency analysis, and correctness checks to provide highly automated translation coverage, significantly reducing the development time while satisfying language-specific requirements like building a proper dataflow model when creating workflows. Furthermore, CNT incorporates a module for aiding manual translation. Specifically, it can identify three common JavaScript patterns in CWL workflows, offering further guidance for developers during the translation phase. We evaluated CNT with production-grade workflows and found that it can cover up to 81% of the original workflows, substantially reducing development time. Additionally, transitioning from a cwltool-based system to Nextflow with CNT can result in a 72% speedup and 85% increased CPU utilization.
Martin L. Putra, In Kee Kim, Haryadi S. Gunawi, Robert L. Grossman
BIBE4
2023 Scalable Batch-Mode Deep Bayesian Active Learning via Equivalence Class Annealing
Aly Azeem Khan, Robert L. Grossman, Yuxin Chen 0001
ICLR3
2023 Building a collaborative cloud platform to accelerate heart, lung, blood, and sleep research
abstract
Research increasingly relies on interrogating large-scale data resources. The NIH National Heart, Lung, and Blood Institute developed the NHLBI BioData CatalystⓇ (BDC), a community-driven ecosystem where researchers, including bench and clinical scientists, statisticians, and algorithm developers, find, access, share, store, and compute on large-scale datasets. This ecosystem provides secure, cloud-based workspaces, user authentication and authorization, search, tools and workflows, applications, and new innovative features to address community needs, including exploratory data analysis, genomic and imaging tools, tools for reproducibility, and improved interoperability with other NIH data science platforms. BDC offers straightforward access to large-scale datasets and computational resources that support precision medicine for heart, lung, blood, and sleep conditions, leveraging separately developed and managed platforms to maximize flexibility based on researcher needs, expertise, and backgrounds. Through the NHLBI BioData Catalyst Fellows Program, BDC facilitates scientific discoveries and technological advances. BDC also facilitated accelerated research on the coronavirus disease-2019 (COVID-19) pandemic.
Stanley C. Ahalt, Paul Avillach, Rebecca R. Boyles, Kira Bradford, Steven Cox 0001, Brandi Davis-Dusenbery, Robert L. Grossman, Ashok K. Krishnamurthy 0001, Alisa Manning, Benedict Paten, Anthony Philippakis, Ingrid Borecki, Shu Hui Chen, Jon Kaltman, Sweta Ladwa, Chip Schwartz, Alastair Thomson, Sarah Davis, Alison Leaf, Jessica Lyons, Elizabeth Sheets, Joshua C. Bis, Matthew P. Conomos, Alessandro Culotti, Thomas N. Desain, Jack DiGiovanna, Milan Domazet, Stephanie M. Gogarten, Alba Gutiérrez-Sacristán, Tim Harris 0003, Benjamin D. Heavner, Deepti Jain, Brian O'Connor, Kevin Osborn, Danielle Pillion, Jacob Pleiness, Ken Rice, Garrett Rupp, Arnaud Serret-Larmande, Albert Smith, Jason Stedman, Adrienne Stilp, Teresa Barsanti, John B. Cheadle, Christopher Erdmann, Brandy Farlow, Allie Gartland-Gray, Julie Hayes, Hannah Hiles, Paul Kerr, W. Christopher Lenhardt, Tom Madden, Joanna O. Mieczkowska, Amanda Miller, Patrick Patton, Marcie Rathbun, Stephanie Suber, Joe Asare
J. Am. Medical Informatics Assoc.7
2023 Towards self-describing and FAIR bulk formats for biomedical data
abstract
We introduce a self-describing serialized format for bulk biomedical data called the Portable Format for Biomedical (PFB) data. The Portable Format for Biomedical data is based upon Avro and encapsulates a data model, a data dictionary, the data itself, and pointers to third party controlled vocabularies. In general, each data element in the data dictionary is associated with a third party controlled vocabulary to make it easier for applications to harmonize two or more PFB files. We also introduce an open source software development kit (SDK) called PyPFB for creating, exploring and modifying PFB files. We describe experimental studies showing the performance improvements when importing and exporting bulk biomedical data in the PFB format versus using JSON and SQL formats.
Michael Lukowski, Andrew Prokhorenkov, Robert L. Grossman
PLoS Comput. Biol.3
2022 The Biomedical Research Hub: a federated platform for patient research data
abstract
OBJECTIVE: The objective was to develop and operate a cloud-based federated system for managing, analyzing, and sharing patient data for research purposes, while allowing each resource sharing patient data to operate their component based upon their own governance rules. The federated system is called the Biomedical Research Hub (BRH). MATERIALS AND METHODS: The BRH is a cloud-based federated system built over a core set of software services called framework services. BRH framework services include authentication and authorization, services for generating and assessing findable, accessible, interoperable, and reusable (FAIR) data, and services for importing and exporting bulk clinical data. The BRH includes data resources providing data operated by different entities and workspaces that can access and analyze data from one or more of the data resources in the BRH. RESULTS: The BRH contains multiple data commons that in aggregate provide access to over 6 PB of research data from over 400 000 research participants. DISCUSSION AND CONCLUSION: With the growing acceptance of using public cloud computing platforms for biomedical research, and the growing use of opaque persistent digital identifiers for datasets, data objects, and other entities, there is now a foundation for systems that federate data from multiple independently operated data resources that expose FAIR application programming interfaces, each using a separate data model. Applications can be built that access data from one or more of the data resources.
Craig Barnes, Binam Bajracharya, Matthew Cannalte, Zakir Gowani, Will Haley, Taha A. Kass-Hout, Kyle Hernandez, Michael Ingram, Hara Prasad Juvvala, Gina Kuffel, Plamen Martinov, J. Montgomery Maxwell, John McCann, Ankit Malhotra, Noah Metoki-Shlubsky, Chris Meyer, Andre Paredes, Jawad Qureshi, Xenia Ritter, Philip Schumm, Mingfei Shao, Urvi Sheth, Trevar Simmons, Alexander Vantol, Zhenyu Zhang 0016, Robert L. Grossman
J. Am. Medical Informatics Assoc.26
2021 Experiences in Managing the Performance and Reliability of a Large-Scale Genomics Cloud Platform
Michael Hao Tong, Robert L. Grossman, Haryadi S. Gunawi
USENIX ATC2
2020 Evaluation of Hyperbolic Attention in Histopathology Images
abstract
We bring together into a common framework three key ideas - multi-scale medical image analysis, the attention mechanism, and hyperbolic embeddings. The formulation and evaluation of hyperbolic-attention models for multi-scale medical image analysis have not been previously explored. In this paper, we evaluate a hyperbolic-attention model on two classification tasks using histopathology image datasets. The experiments show improvement compared to other commonly used models. Our method directly captures the multi-scale structure of histopathology images, and we speculate that the hyperbolic attention mechanism naturally singles out one or more structures at one or more scales that are most discriminatory.
Aly Azeem Khan, Robert L. Grossman
BIBE3
2018 The medical science DMZ: a network design pattern for data-intensive medical science
abstract
OBJECTIVE: We describe a detailed solution for maintaining high-capacity, data-intensive network flows (eg, 10, 40, 100 Gbps+) in a scientific, medical context while still adhering to security and privacy laws and regulations. MATERIALS AND METHODS: High-end networking, packet-filter firewalls, network intrusion-detection systems. RESULTS: We describe a "Medical Science DMZ" concept as an option for secure, high-volume transport of large, sensitive datasets between research institutions over national research networks, and give 3 detailed descriptions of implemented Medical Science DMZs. DISCUSSION: The exponentially increasing amounts of "omics" data, high-quality imaging, and other rapidly growing clinical datasets have resulted in the rise of biomedical research "Big Data." The storage, analysis, and network resources required to process these data and integrate them into patient diagnoses and treatments have grown to scales that strain the capabilities of academic health centers. Some data are not generated locally and cannot be sustained locally, and shared data repositories such as those provided by the National Library of Medicine, the National Cancer Institute, and international partners such as the European Bioinformatics Institute are rapidly growing. The ability to store and compute using these data must therefore be addressed by a combination of local, national, and industry resources that exchange large datasets. Maintaining data-intensive flows that comply with the Health Insurance Portability and Accountability Act (HIPAA) and other regulations presents a new challenge for biomedical research. We describe a strategy that marries performance and security by borrowing from and redefining the concept of a Science DMZ, a framework that is used in physical sciences and engineering research to manage high-capacity data flows. CONCLUSION: By implementing a Medical Science DMZ architecture, biomedical researchers can leverage the scale provided by high-performance computer and cloud storage facilities and national high-speed research networks while preserving privacy and meeting regulatory requirements.
Sean Peisert, Eli Dart, William K. Barnett, Edward Balas, James A. Cuff, Robert L. Grossman, Ari Berman, Anurag Shankar, Brian Tierney
J. Am. Medical Informatics Assoc.6
2016 Deploying Analytics with the Portable Format for Analytics (PFA)
abstract
We introduce a new language for deploying analytic models into products, services and operational systems called the Portable Format for Analytics (PFA). PFA is an example of what is sometimes called a model interchange format, a language for describing analytic models that is independent of specific tools, applications or systems. Model interchange formats allow one application (the model producer) to export models and another application (the model consumer or scoring engine) to import models. The core idea behind PFA is to support the safe execution of statistical functions, mathematical functions, and machine learning algorithms and their compositions within a safe execution environment. With this approach, the common analytic models used in data science can be implemented, as well as the data transformations and data aggregations required for pre- and post-processing data. PFA compliant scoring engines can be extended by adding new user defined functions described in PFA. We describe the design of PFA. A Data Mining Group (DMG) Working Group is developing the PFA standard. The current version is 0.8.1 and contains many of the commonly used statistical and machine learning models, including regression, clustering, support vector machines, neural networks, etc. We also describe two implementations of Hadrian, one in Scala and one in Python. We discuss four case studies that use PFA and Hadrian to specify analytic models, including two that are deployed in operations at client sites.
James Pivarski, Collin Bennett, Robert L. Grossman
KDD3
2016 The Medical Science DMZ
abstract
OBJECTIVE: We describe use cases and an institutional reference architecture for maintaining high-capacity, data-intensive network flows (e.g., 10, 40, 100 Gbps+) in a scientific, medical context while still adhering to security and privacy laws and regulations. MATERIALS AND METHODS: High-end networking, packet filter firewalls, network intrusion detection systems. RESULTS: We describe a "Medical Science DMZ" concept as an option for secure, high-volume transport of large, sensitive data sets between research institutions over national research networks. DISCUSSION: The exponentially increasing amounts of "omics" data, the rapid increase of high-quality imaging, and other rapidly growing clinical data sets have resulted in the rise of biomedical research "big data." The storage, analysis, and network resources required to process these data and integrate them into patient diagnoses and treatments have grown to scales that strain the capabilities of academic health centers. Some data are not generated locally and cannot be sustained locally, and shared data repositories such as those provided by the National Library of Medicine, the National Cancer Institute, and international partners such as the European Bioinformatics Institute are rapidly growing. The ability to store and compute using these data must therefore be addressed by a combination of local, national, and industry resources that exchange large data sets. Maintaining data-intensive flows that comply with HIPAA and other regulations presents a new challenge for biomedical research. Recognizing this, we describe a strategy that marries performance and security by borrowing from and redefining the concept of a "Science DMZ"-a framework that is used in physical sciences and engineering research to manage high-capacity data flows. CONCLUSION: By implementing a Medical Science DMZ architecture, biomedical researchers can leverage the scale provided by high-performance computer and cloud storage facilities and national high-speed research networks while preserving privacy and meeting regulatory requirements.
Sean Peisert, William K. Barnett, Eli Dart, James A. Cuff, Robert L. Grossman, Edward B. Talbot, Ari Berman, Anurag Shankar, Brian Tierney
J. Am. Medical Informatics Assoc.5
2014 Bionimbus: a cloud for managing, analyzing and sharing large genomics datasets
abstract
BACKGROUND: As large genomics and phenotypic datasets are becoming more common, it is increasingly difficult for most researchers to access, manage, and analyze them. One possible approach is to provide the research community with several petabyte-scale cloud-based computing platforms containing these data, along with tools and resources to analyze it. METHODS: Bionimbus is an open source cloud-computing platform that is based primarily upon OpenStack, which manages on-demand virtual machines that provide the required computational resources, and GlusterFS, which is a high-performance clustered file system. Bionimbus also includes Tukey, which is a portal, and associated middleware that provides a single entry point and a single sign on for the various Bionimbus resources; and Yates, which automates the installation, configuration, and maintenance of the software infrastructure required. RESULTS: Bionimbus is used by a variety of projects to process genomics and phenotypic data. For example, it is used by an acute myeloid leukemia resequencing project at the University of Chicago. The project requires several computational pipelines, including pipelines for quality control, alignment, variant calling, and annotation. For each sample, the alignment step requires eight CPUs for about 12 h. BAM file sizes ranged from 5 GB to 10 GB for each sample. CONCLUSIONS: Most members of the research community have difficulty downloading large genomics datasets and obtaining sufficient storage and computer resources to manage and analyze the data. Cloud computing platforms, such as Bionimbus, with data commons that contain large genomics datasets, are one choice for broadening access to research data in genomics.
Allison P. Heath, Matthew Greenway, Ray Powell, Jonathan Spring, Rafael D. Suarez, David Hanley, Chai Bandlamudi, Megan E. McNerney, Kevin P. White, Robert L. Grossman
J. Am. Medical Informatics Assoc.10
2012 The Namibia Early Flood Warning System, a CEOS pilot project
abstract
This paper describes a pilot project effort under the auspices of the Namibian Ministry of Agriculture Water and Forestry (MAWF)/Department of Water Affairs, the Committee on Earth Observing Satellites (CEOS) /Working Group on Information Systems and Services (WGISS) and originally moderated by the United Nations Platform for Space-based Information for Disaster Management and Emergency Response (UN-SPIDER). The effort began by identifying and prototyping technologies which enabled the rapid gathering and dissemination of both space-based and ground sensor data and data products for the purpose of flood disaster management. This was followed by an international collaboration to build small portions of the identified system which was prototyped during the past few years during the flood seasons which occurred in the February through May timeframe of 2010 and 2011 with further prototyping to ongoing in 2012. The pilot effort has been fostered by CEOS to facilitate international efforts to promote satellite sensor data interoperability. In particular, the group has been making use of a technology effort call SensorWeb being developed at NASA which leverages Open Geospatial Consortium (OGC) Sensor Web Enablement (SWE) standards to facilitate various satellite and ground sensor interoperability. The group has made use of such satellites such as Earth Observing 1, Terra/Aqua MODIS and the Canadian Space Agency (CSA) Radarsat together with various ground sensors such as river gauges in Namibia and models such as Global Disaster Alert and Coordination System (GDACS) from Joint Research Center (JRC) from the European Commission. Finally, the group has been experimenting with integrating a large Cloud Computing service provided by the Open Cloud Consortium (OCC) with the SensorWeb to provide management and distribution of the large data sets for emergency workers.
Dan Mandl, Stuart Frye, Robert A. Sohlberg, Patrice Cappelaere, Matthew Handy, Robert L. Grossman
IGARSS6
2011 Toward Efficient and Simplified Distributed Data Intensive Computing
abstract
While the capability of computing systems has been increasing at Moore's Law, the amount of digital data has been increasing even faster. There is a growing need for systems that can manage and analyze very large data sets, preferably on shared-nothing commodity systems due to their low expense. In this paper, we describe the design and implementation of a distributed file system called Sector and an associated programming framework called Sphere that processes the data managed by Sector in parallel. Sphere is designed so that the processing of data can be done in place over the data whenever possible. Sometimes, this is called data locality. We describe the directives Sphere supports to improve data locality. In our experimental studies, the Sector/Sphere system has consistently performed about 2-4 times faster than Hadoop, the most popular system for processing very large data sets.
Yunhong Gu, Robert L. Grossman
IEEE Trans. Parallel Distributed Syst.2
2010 An overview of the Open Science Data Cloud
abstract
The Open Science Data Cloud is a distributed cloud based infrastructure for managing, analyzing, archiving and sharing scientific datasets. We introduce the Open Science Data Cloud, give an overview of its architecture, provide an update on its current status, and briefly describe some research areas of relevance.
Robert L. Grossman, Yunhong Gu, Joe Mambretti, Michal Sabala, Alex Szalay, Kevin P. White
HPDC1
2010 Malstone: towards a benchmark for analytics on large data clouds
abstract
Developing data mining algorithms that are suitable for cloud computing platforms is currently an active area of research, as is developing cloud computing platforms appropriate for data mining. Currently, the most common benchmark for cloud computing is the Terasort (and related) benchmarks. Although the Terasort Benchmark is quite useful, it was not designed for data mining per se. In this paper, we introduce a benchmark called MalStone that is specifically designed to measure the performance of cloud computing middleware that supports the type of data intensive computing common when building data mining models. We also introduce MalGen, which is a utility for generating data on clouds that can be used with MalStone.
Collin Bennett, Robert L. Grossman, David Locke, Jonathan Seidman, Steve Vejcik
KDD2
2010 Sector: A high performance wide area community data storage and sharing system
Yunhong Gu, Robert L. Grossman
Future Gener. Comput. Syst.2
2009 Flynet: a genomic resource for Drosophila melanogaster transcriptional regulatory networks
abstract
MOTIVATION: The highly coordinated expression of thousands of genes in an organism is regulated by the concerted action of transcription factors, chromatin proteins and epigenetic mechanisms. High-throughput experimental data for genome wide in vivo protein-DNA interactions and epigenetic marks are becoming available from large projects, such as the model organism ENCyclopedia Of DNA Elements (modENCODE) and from individual labs. Dissemination and visualization of these datasets in an explorable form is an important challenge. RESULTS: To support research on Drosophila melanogaster transcription regulation and make the genome wide in vivo protein-DNA interactions data available to the scientific community as a whole, we have developed a system called Flynet. Currently, Flynet contains 101 datasets for 38 transcription factors and chromatin regulator proteins in different experimental conditions. These factors exhibit different types of binding profiles ranging from sharp localized peaks to broad binding regions. The protein-DNA interaction data in Flynet was obtained from the analysis of chromatin immunoprecipitation experiments on one color and two color genomic tiling arrays as well as chromatin immunoprecipitation followed by massively parallel sequencing. A web-based interface, integrated with an AJAX based genome browser, has been built for queries and presenting analysis results. Flynet also makes available the cis-regulatory modules reported in literature, known and de novo identified sequence motifs across the genome, and other resources to study gene regulation. AVAILABILITY: Flynet is available at https://www.cistrack.org/flynet/.
Parantu K. Shah, Nicolas N. Negre, Oleksiy Karpenko, Kevin P. White, Robert L. Grossman
Bioinform.8
2009 Compute and storage clouds using wide area high performance networks
Robert L. Grossman, Yunhong Gu, Michal Sabala, Wanzhi Zhang
Future Gener. Comput. Syst.1
2008 Data mining using high performance data clouds: experimental studies using sector and sphere
abstract
We describe the design and implementation of a high performance cloud that we have used to archive, analyze and mine large distributed data sets. By a cloud, we mean an infrastructure that provides resources and/or services over the Internet. A storage cloud provides storage services, while a compute cloud provides compute services. We describe the design of the Sector storage cloud and how it provides the storage services required by the Sphere compute cloud. We also describe the programming paradigm supported by the Sphere compute cloud. Sector and Sphere are designed for analyzing large data sets using computer clusters connected with wide area high performance networks (for example, 10+ Gb/s). We describe a distributed data mining application that we have developed using Sector and Sphere. Finally, we describe some experimental studies comparing Sector/Sphere to Hadoop.
Robert L. Grossman, Yunhong Gu
KDD1
2007 An Algorithm for Assigning Unique Keys to Metabolic Pathways
abstract
Different databases of metabolic pathways assign pathways different keys. For this reason, it is difficult to automatically compare pathways across databases. We introduce an algorithm called the Universal Pathway Key algorithm or UPK that assigns essentially unique keys to metabolic pathways. We show that the UPK algorithm assigns unique keys to the pathways in the MetaCyc database and can also be used to detect duplicate pathways. The UPK algorithm is a simple generalization of the UCK algorithm introduced in [7] that assigns essentially unique keys to chemical compounds.
Robert L. Grossman
BIBM2
2007 Detecting changes in large data sets of payment card data: a case study
abstract
An important problem in data mining is detecting changes in large datasets. Although there are a variety of change detection algorithms that have been developed, in practice it can be a problem to scale these algorithms to large data sets due to the heterogeneity of the data. In this paper, we describe a case study involving payment card data in which we built and monitored a separate change detection model for each cell in a multi-dimensional data cube. We describe a system that has been in operation for the past two years that builds and monitors over 15,000 separate baseline models and the process that isused for generating and investigating alerts using these baselines.
Chris Curry, Robert L. Grossman, David Locke, Steve Vejcik, Joseph Bugajski
KDD2
2007 UDT: UDP-based data transfer for high-speed wide area networks
Yunhong Gu, Robert L. Grossman
Comput. Networks2
2006 SDCS: Simplified Data Communications in Parallel/Distributed Applications
abstract
This paper presents SDCS (Simple Data Communication and Sharing), a programming model for data communications in parallel/distributed applications. With SDCS, developers can define data communications in shared memory style and have the model translate the declarations into corresponding message passing code. The translation from data sharing declarations to message passing code is based on simple mapping rules to lower runtime overhead and increase understandability of the model. Some frequently seen data communication modes are well supported to enhance its usability. SDCS can effectively reduce the difficulty in programming process communications.
Yong Mao, Yunhong Gu, Robert L. Grossman
CCGRID4
2006 Distributing the Sloan Digital Sky Survey Using UDT and Sector
abstract
In this paper, we describe a peer-to-peer storage system called Sector that is designed to access and transport large data sets over wide area high performance networks. We also describe our recent experience using Sector to distribute the Sloan Digital Sky Survey BESTDR4 catalog data.
Yunhong Gu, Robert L. Grossman, Alex Szalay, Ani Thakar
e-Science2
2006 A Service Oriented Architecture Supporting Data Interoperability for Payments Card Processing Systems
Joseph Bugajski, Robert L. Grossman, Steve Vejcik
ICSOC2
2006 Bandwidth challenge - Transporting sloan digital sky survey data using SECTOR
abstract
National Center for Data Mining at UICIn our SC06 BWC entry, we will transfer SDSS (Sloan Digital Sky Survey) Data Release 5 (DR5) between the SC06 show floor in Tampa and one of the NCDM labs on the UIC campus. We will use SECTOR, our newly developed distributed data space management system, to transfer DR5 in parallel between two Linux clusters in Tampa and Chicago, respectively. SECTOR transparently manages the file locating and data moving, while it employs UDT for actual data transfer. The data transfer will be from disk to disk over a 10Gb/s shared, router link between SC06 and UIC, via StarLight. We expect to reach 5Gb/s disk-to-disk data transfer rate between the two sites.
Robert L. Grossman, Yunhong Gu, Michal Sabala, Shirley Connelly, David Hanley, Joe Mambretti, Alex Szalay, Ani Thakar, Jan vandenBerg, Alainna Wonders
SC1
2006 Data mining middleware for wide-area high-performance networks
Robert L. Grossman, Yunhong Gu, David Hanley, Michal Sabala, Joe Mambretti, Alex Szalay, Ani Thakar, Kazumi Kumazoe, Yuji Oie, Yoonjoo Kwon, Woojin Seok
Future Gener. Comput. Syst.1
2006 High-Dimensional Visual Analytics: Interactive Exploration Guided by Pairwise Views of Point Distributions
abstract
We introduce a method for organizing multivariate displays and for guiding interactive exploration through high-dimensional data. The method is based on nine characterizations of the 2D distributions of orthogonal pairwise projections on a set of points in multidimensional Euclidean space. These characterizations include such measures as density, skewness, shape, outliers, and texture. Statistical analysis of these measures leads to ways for 1) organizing 2D scatterplots of points for coherent viewing, 2) locating unusual (outlying) marginal 2D distributions of points for anomaly detection, and 3) sorting multivariate displays based on high-dimensional data, such as trees, parallel coordinates, and glyphs.
Leland Wilkinson, Anushka Anand, Robert L. Grossman
IEEE Trans. Vis. Comput. Graph.3
2005 Real Time Change Detection and Alerts from Highway Traffic Data
abstract
We developed a testbed containing: real time data from over 830 highway traffic sensors in the Chicago region, data about weather, and text data about events that might affect traffic. The goal was to detect in real time interesting changes in traffic conditions. Given the size and complexity of the data, we choose to build a large number of separate baseline models. We built a separate baseline for each hour in the day, for each day in the week, and for every 2 or 3 traffic sensors, resulting in over 42,000 separate baseline models. We also built a baseline engine to build the necessary baselines automatically. We modified an open source scoring engine to process in real time each new sensor reading, update the appropriate feature vectors, score the updated feature vectors using the baseline models, and send out real time alerts when deviations from the baselines were detected.
Robert L. Grossman, Michal Sabala, Anushka Anand, Steve Eick, Leland Wilkinson, John Chaves, Steve Vejcik, John F. Dillenburg, Peter C. Nelson, Doug Rorem, Javid Alimohideen, Jason Leigh, Michael E. Papka, Rick L. Stevens
SC1
2005 Supporting Configurable Congestion Control in Data Transport Services
abstract
As wide area high-speed networks rapidly increase, new applications emerge and require new control mechanisms in data transport services to support them. In this paper, we present UDT/CCC, a data transport library that allows users to make use of a new control algorithm through simple configurations. We aim to provide a tool for fast implementation and deployment, as well as easy evaluation, of new congestion control algorithms. UDT/CCC uses an objected-oriented design. We show that our UDT/CCC library can be used to easily implement a large variety of control algorithms and can simulate the behavior of their native implementations as well. The UDT/CCC library is at the application level and it does not need root privilege to be installed. Meanwhile, it was specially developed to require very few changes to the existing applications. This paper describes its design, implementation, and evaluation.
Yunhong Gu, Robert L. Grossman
SC2
2005 Teraflows over Gigabit WANs with UDT
Robert L. Grossman, Yunhong Gu, Xinwei Hong, Antony Antony, Johan Blom, Freek Dijkstra, Cees T. A. M. de Laat
Future Gener. Comput. Syst.1
2005 Simple Available Bandwidth Utilization Library for High-Speed Wide Area Networks
Robert L. Grossman, Marco Mazzucco, Harimath Sivakumar, Y. Pan
J. Supercomput.1
2004 A Greedy Algorithm for Selecting Models in Ensembles
abstract
We are interested in ensembles of models built over k data sets. Common approaches are either to combine models by vote averaging, or to build a meta-model on the outputs of the local models. In this paper, we consider the model assignment approach, in which a meta-model selects one of the local statistical models for scoring. We introduce an algorithm called greedy data labeling (GDL) that improves the initial data partition by reallocating some data, so that when each model is built on its local data subset, the resulting hierarchical system has minimal error. We present evidence that model assignment may in certain situations be more natural than traditional ensemble learning, and if enhanced by GDL, it often outperforms traditional ensembles.
Andrei L. Turinsky, Robert L. Grossman
ICDM2
2004 Experimental Studies Using Median Polish Procedure to Reduce Alarm Rates in Data Cubes of Intrusion Data
Jorge Levera, Benjamín Barán, Robert L. Grossman
ISI3
2004 Using DataSpace Archives to Support Long-Term Stewardship of Remote and Distributed Data
Robert L. Grossman, David Hanley, Xinwei Hong, Parthasarathy Krishnaswamy
MSST1
2004 Experiences in Design and Implementation of a High Performance Transport Protocol
abstract
This paper describes our experiences in the development of the UDP-based Data Transport (UDT) protocol, an application level transport protocol used in distributed data intensive applications. The new protocol is motivated by the emergence of wide area high-speed optical networks, in which TCP is often found to fail to utilize the abundant bandwidth. UDT demonstrates good efficiency and fairness (including RTT fairness and TCP friendliness) characteristics in high performance computing applications where a small number of bulk sources share the abundant bandwidth. It combines both rate and window control and uses bandwidth estimation to determine the control parameters automatically. This paper presents the rationale behind UDT: how UDT integrates these schemes to support high performance data transfer, why these schemes are used, and what the main issues are in the design and implementation of this high performance transport protocol.
Yunhong Gu, Xinwei Hong, Robert L. Grossman
SC3
2004 GenIc: A Single-Pass Generalized Incremental Algorithm for Clustering
abstract
In this paper we introduce a new single pass clustering algorithm called GenIc designed with the objective of having low overall cost. We examine some of the properties of GenIc and compare it to windowed k-means. We also study its performance using experimental data sets obtained from network monitoring.
Chetan Gupta 0001, Robert L. Grossman
SDM2
2004 Experimental studies of data transport and data access of earth-science data over networks with high bandwidth delay products
Robert L. Grossman, Yunhong Gu, David Hanley, Xinwei Hong, Babu Krishnaswamy
Comput. Networks1
2003 Mining data records in Web pages
abstract
A large amount of information on the Web is contained in regularly structured objects, which we call data records. Such data records are important because they often present the essential information of their host pages, e.g., lists of products or services. It is useful to mine such data records in order to extract information from them to provide value-added services. Existing automatic techniques are not satisfactory because of their poor accuracies. In this paper, we propose a more effective technique to perform the task. The technique is based on two observations about data records on the Web and a string matching algorithm. The proposed technique is able to mine both contiguous and non-contiguous data records. Our experimental results show that the proposed technique outperforms existing techniques substantially.
Bing Liu 0001, Robert L. Grossman, Yanhong Zhai
KDD2
2003 Experimental studies using photonic data services at IGrid 2002
Robert L. Grossman, Yunhong Gu, Don Hamelburg, David Hanley, Xinwei Hong, Jorge Levera, David J. Lillethun, Marco Mazzucco, Joe Mambretti, Jeremy Weinberger
Future Gener. Comput. Syst.1
2003 The Photonic TeraStream: enabling next generation applications through intelligent optical networking at iGRID2002
Joe Mambretti, Jeremy Weinberger, Jim Hao Chen, Elizabeth Bacon, Fei Yeh, David J. Lillethun, Robert L. Grossman, Yunhong Gu, Marco Mazzucco
Future Gener. Comput. Syst.7
2003 TeraScope: distributed visual data mining of terascale data sets over photonic networks
Jason Leigh, Thomas A. DeFanti, Marco Mazzucco, Robert L. Grossman
Future Gener. Comput. Syst.5
2003 SABUL: A Transport Protocol for Grid Computing
Yunhong Gu, Robert L. Grossman
J. Grid Comput.2
2003 Data webs for earth science data
Asvin Ananthanarayan, Rajiv Balachandran, Robert L. Grossman, Yunhong Gu, Xinwei Hong, Jorge Levera, Marco Mazzucco
Parallel Comput.3
2002 An Algebraic Approach to Data Mining: Some Examples
abstract
We introduce an algebraic approach to the foundations of data mining. Our approach is based upon two algebras of functions defined over a common state space X and a pairing between them. One algebra is an algebra of state space observations, and the other is an algebra of labeled sets of states. We interpret H as the algebraic encoding of the data and the pairing as the misclassification rate when the classifier f is applied to the set of states X. We give a realization theorem giving conditions on formal series of data sets built from D that imply there is a realization involving a state space X, a classifier f /spl isin/ R and a set of labeled states /spl chi/ /spl isin/ R/sub 0/ that yield this series.
Robert L. Grossman, Richard G. Larson
ICDM1
2002 Merging multiple data streams on common keys over high performance networks
abstract
The model for data mining on streaming data assumes that there is a buffer of fixed length and a data stream of infinite length and the challenge is to extract patterns, changes, anomalies, and statistically significant structures by examining the data one time and storing records and derived attributes of length less than N. As data grids, data webs, and semantic webs become more common, mining distributed streaming data will become more and more important. The first step when presented with two or more distributed streams is to merge them using a common key. In this paper, we present two algorithms for merging streaming data using a common key. We also present experimental studies showing these algorithms scale in practice to OC-12 networks.
Marco Mazzucco, Asvin Ananthanarayan, Robert L. Grossman, Jorge Levera, Gokulnath Bhagavantha Rao
SC3
2000 Performance of DB2 Enterprise-Extended Edition on NT with Virtual Interface Architecture
Sivakumar Harinath, Robert L. Grossman, K. Bernhard Schiefer, Xun Xue, Sadique Syed
EDBT2
2000 PSockets: The Case for Application-level Network Striping for Data Intensive Applications using High Speed Wide Area Networks
abstract
Transmission Control Protocol (TCP) is used by various applications to achieve reliable data transfer. TCP was originally designed for unreliable networks. With the emergence of high-speed wide area networks various improvements have been applied to TCP to reduce latency and achieve improved bandwidth. The improvement is achieved by having system administrators tune the network and can take a considerable amount of time. This paper introduces PSockets (Parallel Sockets), a library that achieves an equivalent performance without manual tuning. The basic idea behind PSockets is to exploit network striping. By network striping we mean striping partitioned data across several open sockets. We describe experimental studies using PSockets over the Abilene network. We show in particular that network striping using PSockets is effective for high performance data intensive computing applications using geographically distributed data.
Harimath Sivakumar, Stuart Bailey, Robert L. Grossman
SC3
1999 A Methodology for Supporting Collaborative Exploratory Analysis of Massive Data Sets in Tele-Immersive Environments
abstract
This paper proposes a methodology for employing collaborative, immersive virtual environments as a high-end visualization interface for massive data-sets. The methodology employs feature detection, partitioning, summarization and decimation to significantly cull massive data-sets. These reduced data-sets are then distributed to the remote CAVEs, ImmersaDesks and desktop workstations for viewing. The paper also discusses novel techniques for collaborative visualization and meta-data creation.
Jason Leigh, Andrew E. Johnson 0001, Thomas A. DeFanti, Stuart Bailey, Robert L. Grossman
HPDC5
1999 Papyrus: A System for Data Mining over Local and Wide Area Clusters and Super-Clusters
abstract
Article Free Access Share on Papyrus: a system for data mining over local and wide area clusters and super-clusters Authors: S. Bailey National Center for Data Mining, University of Illinois at Chicago National Center for Data Mining, University of Illinois at ChicagoView Profile , R. Grossman National Center for Data Mining, University of Illinois at Chicago and Magnify, Inc. National Center for Data Mining, University of Illinois at Chicago and Magnify, Inc.View Profile , H. Sivakumar National Center for Data Mining, University of Illinois at Chicago National Center for Data Mining, University of Illinois at ChicagoView Profile , A. Turinsky National Center for Data Mining, University of Illinois at Chicago National Center for Data Mining, University of Illinois at ChicagoView Profile Authors Info & Claims SC '99: Proceedings of the 1999 ACM/IEEE conference on SupercomputingJanuary 1999 Pages 63–eshttps://doi.org/10.1145/331532.331595Online:01 January 1999Publication History 44citation676DownloadsMetricsTotal Citations44Total Downloads676Last 12 Months7Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Stuart Bailey, Robert L. Grossman, Harimath Sivakumar, Andrei L. Turinsky
SC2
1999 Editorial
Yike Guo, Robert L. Grossman
Data Min. Knowl. Discov.2
1999 The management and mining of multiple predictive models using the predictive modeling markup language
Robert L. Grossman, Stuart Bailey, Ashok Ramu, Balinder Malhi, Philip Hallstrom, Ivan Pulleyn, Xiao Qin 0006
Inf. Softw. Technol.1
1997 A Tutorial Introduction to High Performance Data Mining (Abstract)
Robert L. Grossman
PKDD1
1996 Optimization driven data mining and credit scoring
abstract
An optimization tree approach to the mining of very extensive and complex databases for performance optimizing opportunities is described. This methodology is based on a combination of three innovations: a data management system designed explicitly for data intensive computing; a distributed algorithm for growing classification and regression trees (CART); and a tree based stochastic programming paradigm for the selection of control attributes to optimize a specified objective function. This methodology provides a general technique for optimization in financial applications that is scalable as the number of objects in the database and as the number of attributes per object grow. This scalability allows for a complete data driven analysis of large scale data sets, without the need to restrict attention to sparsely sampled data sets that limits previous methods.
Robert L. Grossman, H. Vincent Poor
CIFEr1
1996 Data Mining and Tree-Based Optimization
Robert L. Grossman, Haim Bodek, Dave Northcutt, H. Vincent Poor
KDD1
1995 The Symbolic Computation of Differential Invariants of Polynomial Vector Field Systems Using Trees
abstract
Let K denote a field of characteristic 0, and let V = KN denote the vector space over K of dimension N. Let R denote the K-algebra of polynomials over
M. J. Doffou, Robert L. Grossman
ISSAC2
1995 PTool: A Light Weight Persistent Object Manager
abstract
No abstract available.
Robert L. Grossman, David Hanley, Xiao Qin 0006
SIGMOD Conference1
1995 An Algebraic Approach to Hybrid Systems
Robert L. Grossman, Richard G. Larson
Theor. Comput. Sci.1
1994 Ptool: A Scalable Persistent Object Manager
abstract
No abstract available.
Robert L. Grossman, Xiao Qin 0006
SIGMOD Conference1
1994 Analyzing High Energy Physics Data Using Databases: A Case Study
abstract
We describe the initial work of the PASS Project which uses techniques from distributed object management to analyze experimental data from high energy physics. At this time, we have designed two prototypes to analyze high energy physics data from the CDF experiment at Fermi Lab. The data from this experiment consists of "events" which describe particle collisions. Each event consists of several hundred numerical attributes and occupies approximately 10 K in a compressed format. We describe our experience analyzing this data using a relational database, an object oriented database, and a persistent object manager.>
Robert L. Grossman, Xiao Qin 0006, D. Valsamis, Christopher T. Day, Stewart C. Loken, J. F. MacFarlane, David R. Quarrie, Edward N. May, David Lifka, David M. Malon, L. E. Price, A. Baden, L. Cormell, Phil Leibold, U. Nixdorf, B. Scipioni, T. Song
SSDBM1
1994 Visibility with a Moving Point of View
Marshall W. Bern, David P. Dobkin, David Eppstein, Robert L. Grossman
Algorithmica4
1993 Wavelet transforms associated with finite cyclic groups
abstract
Multiresolution analysis via decomposition on wavelet bases has emerged as an important tool in the analysis of signals and images when these objects are viewed as sequences of complex or real numbers. An important class of multiresolution decompositions are the Laplacian pyramid schemes, in which the resolution is successively halved by recursively low-pass filtering the signal under analysis and decimating it by a factor of two. In general, the principal framework within which multiresolution techniques have been studied and applied is the same as that used in the discrete-time Fourier analysis of sequences of complex numbers. An analogous framework is developed for the multiresolution analysis of finite-length sequences of elements from arbitrary fields. Attention is restricted to sequences of length 2/sup n/, for n a positive integer, so that the resolution may be recursively halved to completion. As in finite-length Fourier analysis, a cyclic group structure of the index set of such sequences is exploited to characterize the transforms of interest for the particular cases of complex and finite fields.>
Giuseppe Caire, Robert L. Grossman, H. Vincent Poor
IEEE Trans. Inf. Theory2
1992 The Explicit Computation of Integration Algorithms and First Integrals for Ordinary Differential Equations with Polynomial Coefficients Using Trees
abstract
This note is concerned with the explicit symbolic computation of expressions involving differential operators and their actions on functions. The derivationof specialized numerical algorithms, the explicit symbolic computation of integrals of motion, and the explicit computation of normal forma for nonlinear systems all require such computations. More precisely, if R = It[xl,.... ZN], where k = R or C, F denotes a differential operator with coefficients from R, and g E R, we describe data structures and algorithms for efficiently computing F. g. The basic idea is to impose a multiplicative structure on the vector space with basis the set of finite rooted trees and whose nodes are labeled with the coefficients of the differential operators. Cancellations of two trees with r + 1 nodes translates into cancellation of O(Nr) expressions involving the coefficient functions and their derivatives. 1
Peter E. Crouch, Robert L. Grossman
ISSAC2
1992 Symbolic Computation of Derivations Using Labeled Trees
Robert L. Grossman, Richard G. Larson
J. Symb. Comput.1
1991 Computations Involving Differential Operators and Their Actions on Functions
abstract
The algorithms derived by Grossmann and Larson (1989) are further developed for rewriting expressions involving differential operators. The differential operators involved arise in the local analysis of nonlinear dynamical systems. These algorithms are extended in two different directions: the algorithms are generalized so that they apply to differential operators on groups and the data structures and algorithms are developed to compute symbolically the action of differential operators on functions. Both of these generalizations are needed for applications.
Peter E. Crouch, Robert L. Grossman, Richard G. Larson
ISSAC2
1990 Visibility with a Moving Point of View
Marshall W. Bern, David P. Dobkin, David Eppstein, Robert L. Grossman
SODA4
1989 Labeled Trees and the Efficient Computation of Derivations
abstract
The effective parallel symbolic computation of operators under composition is discussed. Examples include differential operators under composition and vector fields under the Lie bracket. Data structures consisting of formal linear combinations of rooted labeled trees are discussed. A multiplication on rooted labeled trees is defined, thereby making the set of these data structures into an associative algebra. An algebra homomorphism is defined from the original algebra of operators into this algebra of trees. An algebra homomorphism from the algebra of trees into the algebra of differential operators is then described. The cancellation which occurs when noncommuting operators are expressed in terms of commuting ones occurs naturally when the operators are represented using this data structure. This leads to an algorithm which, for operators which are derivations, speeds up the computation exponentially in the degree of the operator. It is shown that the algebra of trees leads naturally to a parallel version of the algorithm.
Robert L. Grossman, Richard G. Larson
ISSAC1