EDBT 2026 Demo / reviewers in the wild / expert
Chris Mattmann
dblp:m/ChrisMattmann · also Chris A. Mattmann
· DBLP profile ↗
23ranked-venue papers
8as first author
1since 2021 · last 2023
0000-0002-9414-2901ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 2 first-authorSoftware engineering, systems software and programming languages · 10 · 4 first-authorArtificial intelligence and machine learning · 6 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
5 papers |
Requirements engineering and software design · 85% Empirical software engineering · 15% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 45% Parallel and multicore computing · 34% Cloud and datacenter computing · 21% |
Topics — the 6 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Requirements engineering and software design › software architecture › software architecture analysis
software architecture recovery |
0.4 | 3 | 2013 | Obtaining ground-truth software architectures · ICSE 2013 Enhancing architectural recovery using concerns · ASE 2011 Kadre: domain-specific architectural recovery for scientific software systems · ASE 2010 |
Requirements engineering and software design
software architecture |
0.4 | 3 | 2013 | Obtaining ground-truth software architectures · ICSE 2013 Enhancing architectural recovery using concerns · ASE 2011 A software architecture-based framework for highly distributed and data intensive scientific applications · ICSE 2006 |
High-performance computing
data-intensive computing |
0.1 | 2 | 2006 | Software Connectors for Highly Distributed and Voluminous Data Intensive Systems · ASE 2006 A software architecture-based framework for highly distributed and data intensive scientific applications · ICSE 2006 |
Requirements engineering and software design › software architecture
distributed system architecture |
0.1 | 1 | 2006 | A software architecture-based framework for highly distributed and data intensive scientific applications · ICSE 2006 |
Parallel and multicore computing
data distribution |
0.1 | 1 | 2006 | Software Connectors for Highly Distributed and Voluminous Data Intensive Systems · ASE 2006 |
Computational science and engineering
scientific software |
0.0 | 1 | 2010 | Kadre: domain-specific architectural recovery for scientific software systems · ASE 2010 |
Methods — techniques the papers use, named apart from their topics
clustering · 0.2architecture recovery framework · 0.2machine learning · 0.1concern mining · 0.1software framework · 0.1classification framework · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Guest Editorial Special Issue on Deep Learning for Earth and Planetary GeosciencesabstractEarth and planetary geosciences are essential for understanding and addressing many societal challenges and scientific questions. Increased availability of geoscience data creates an opportunity for deep learning to advance the methods and scientific understanding for tackling these challenges. However, the complex characteristics of geoscience problems and datasets necessitate the development of novel approaches and frameworks. This Special Issue aims at collecting new ideas and deep learning formulations to gain new earth and planetary insights. The contributions cover a wide range of topics, such as satellite and hyperspectral imaging, land monitoring, geophysical imaging, and subsurface analysis. Antonio R. Paiva, Weichang Li, Chris Mattmann, Youzuo Lin, Maarten V. de Hoop |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Deep web crawling for insights from polar dataabstractWe describe efforts to bring new methods of search analytics, machine learning, natural language processing and data visualization to address the challenge of finding and extracting meaning from unstructured text and multimedia content. We use the Polar domain to motivate the problem and our proposed solution. However our techniques are applicable and scalable to other domains. Siri Jodha S. Khalsa, Chris Mattmann, Ruth E. Duerr |
IGARSS | 2 |
| 2017 | Scalable Hadoop-Based Pooled Time Series of Big Video Data from the Deep WebabstractWe contribute a scalable, open source implementation of the Pooled Time Series (PoT) algorithm from CVPR 2015. The algorithm is evaluated on approximately 6800 human trafficking (HT) videos collected from the deep and dark web, and on an open dataset: the Human Motion Database (HMDB). We describe PoT and our motivation for using it on larger data and the issues we encountered. Our new solution reimagines PoT as an Apache Hadoop-based algorithm. We demonstrate that our new Hadoop-based algorithm successfully identifies similar videos in the HT and HMDB datasets and we evaluate the algorithm qualitatively and quantitatively. Chris Mattmann, Madhav Sharan |
ICMR | 1 |
| 2016 | SciSpark: Highly interactive in-memory science data analyticsabstractWe present further work on SciSpark, a Big Data framework that extends Apache Spark's inmemory parallel computing to scale scientific computations. SciSpark's current architecture and design includes: time and space partitioning of highresolution geo-grids from NetCDF3/4; a sciDataset class providing N-dimensional array operations in Scala/Java and CF-style variable attributes (an update of our prior sciTensor class); parallel computation of time-series statistical metrics; and an interactive front-end using science (code & visualization) Notebooks. We demonstrate how SciSpark achieves parallel ingest and time/space partitioning of Earth science satellite and model datasets. We illustrate the usability, extensibility, and early performance of SciSpark using several Earth science Use cases, here presenting benchmarks for sciDataset Readers and parallel time-series analytics. A three-hour SciSpark tutorial was taught at an ESIP Federation meeting using a dozen “live” Notebooks. Rahul Palamuttam, Kim Whitehall, Chris Mattmann, Alex Goodman, Maziyar Boustani, Sujen Shah, Paul Zimdars, Paul M. Ramirez |
IEEE BigData | 4 |
| 2015 | SciSpark: Applying in-memory distributed computing to weather event detection and trackingabstractIn this paper we present SciSpark, a Big Data framework that extends Apache™ Spark for scaling scientific computations. The paper details the initial architecture and design of SciSpark. We demonstrate how SciSpark achieves parallel ingesting and partitioning of earth science satellite and model datasets. We also illustrate the usability and extensibility of SciSpark by implementing aspects of the Grab 'em Tag 'em Graph 'em (GTG) algorithm using SciSpark and its Map Reduce capabilities. GTG is a topical automated method for identifying and tracking Mesoscale Convective Complexes in satellite infrared datasets. Rahul Palamuttam, Renato Marroquín, Chris Mattmann, Kim Whitehall, Rishi Verma, Lewis J. McGibbney, Paul M. Ramirez |
IEEE BigData | 3 |
| 2015 | Revisiting the Anatomy and Physiology of the GridabstractA domain-specific software architecture (DSSA) represents an effective, generalized, reusable solution to constructing software systems within a given application domain. In this paper, we revisit the widely cited DSSA for the domain of grid computing. We have studied systems in this domain over the last ten years. During this time, we have repeatedly observed that, while individual grid systems are widely used and deemed successful, the grid DSSA is actually underspecified to the point where providing a precise answer regarding what makes a software system a grid system is nearly impossible. Moreover, every one of the existing purported grid technologies actually violates the published grid DSSA. In response to this, based on an analysis of the source code, documentation, and usage of eighteen of the most pervasive grid technologies, we have significantly refined the original grid DSSA. We demonstrate that this DSSA much more closely matches the grid technologies studied. Our refinements allow us to more definitively identify a software system as a grid technology, and distinguish it from software libraries, middleware, and frameworks. Chris Mattmann, Joshua Garcia, Ivo Krka, Daniel Popescu 0001, Nenad Medvidovic |
J. Grid Comput. | 1 |
| 2014 | A Laboratory-Targeted, Data Management and Processing System for the Early Detection Research NetworkabstractThe National Institutes of Health (NIH), National Cancer Institute's Early Detection Research Network (EDRN) is a cross-institutional collaborative initiative seeking to accelerate the clinical application of cancer biomarker research. Over the past decade, it has been our role, as EDRN's Informatics Center (IC), to develop a comprehensive information services infrastructure as well as a set of software tools and services to support this overall initiative. We have recently developed a novel application called the Laboratory Catalog and Archive Service (LabCAS) whose focus is to extend EDRN IC data management and processing capabilities to EDRN laboratories. By leveraging the same technologies used to manage and process NASA Earth and Planetary data sets, we offer EDRN researchers an effective way of managing their laboratory data. More specifically, LabCAS enables EDRN researchers to reliably archive their experimental data, to optionally share these data in a controlled manner with other researchers, and to gain insight into these data through highly configurable data analysis pipelines tailored to the broad range of biomarker related experiments. This particular collaboration leverages expertise from NASA's Jet Propulsion Laboratory, Vanderbilt University Medical Center, and Dartmouth Medical School, as well as builds upon existing cross-governmental collaboration between NASA and the NIH. Rishi Verma, Andrew F. Hart, Chris Mattmann, Daniel J. Crichton, Heather Kincaid, Sean C. Kelly, Michael J. Joyce, Paul Zimdars, David L. Tabb, Jay D. Holman, Matthew Chambers, Kristen Anton, Maureen Colbert, Christos Patriotis, Sudhir Srivastava |
CBMS | 3 |
| 2014 | 24 Hour near real time processing and computation for the JPL Airborne Snow ObservatoryabstractJPL's Airborne Snow Observatory is an integrated imaging spectrometer and scanning LIDAR for measuring mountain snow albedo, snow depth/snow water equivalent, and ice height (once exposed). This paper describes the first year of the project's "Snow On" campaign where over a course of 3 months, ASO flew the Tuolumne River Basin, Sierra Nevada, California above the O'Shaughnessy Dam of the Hetch Hetchy reservoir; focusing initial on the Tuolumne, and then moved to weekly flights over the Uncompahgre Basin, Colorado. To meet the needs of its customers including Water Resource managers who are keenly interested in Snow melt, the ASO team had to develop and end to end 24 hour latency capability for processing spectrometer and LIDAR data from Level 0 to Level 4 products. This paper describes the Big data processing architecture and data system for ASO. Chris Mattmann, Thomas H. Painter, Paul M. Ramirez, Cameron Goodale, Andrew F. Hart, Paul Zimdars, Maziyar Boustani, Shakeh E. Khudikyan, Rishi Verma, Felix C. Seidel, Jeffrey Deems, Amy Trangsrud, Joseph W. Boardman |
IGARSS | 1 |
| 2014 | The Earth System Grid Federation: An open infrastructure for access to distributed geospatial data
Luca Cinquini, Daniel J. Crichton, Chris Mattmann, John Harney, Galen M. Shipman, Feiyi Wang, Rachana Ananthakrishnan, Neill Miller, Sebastien Denvil, Mark Morgan, Zed Pobre, Gavin M. Bell, Charles M. Doutriaux, Bob Drach, Dean N. Williams, Philip Kershaw, Stephen Pascoe, Estanislao Gonzalez, Sandro Fiore, Roland Schweitzer |
Future Gener. Comput. Syst. | 3 |
| 2013 | Public infrastructure for cancer biomarker data capture, annotation, analysis, and distribution
David L. Tabb, Kristen Anton, Matthew Chambers, Maureen Colbert, Andrew F. Hart, Jerry D. Holman, Sean C. Kelly, Heather Kincaid, Chris Mattmann, Daniel J. Crichton |
AMIA | 9 |
| 2013 | Obtaining ground-truth software architecturesabstractUndocumented evolution of a software system and its underlying architecture drives the need for the architecture's recovery from the system's implementation-level artifacts. While a number of recovery techniques have been proposed, they suffer from known inaccuracies. Furthermore, these techniques are difficult to evaluate due to a lack of “ground-truth” architectures that are known to be accurate. To address this problem, we argue for establishing a suite of ground-truth architectures, using a recovery framework proposed in our recent work. This framework considers domain-, application-, and context-specific information about a system, and addresses an inherent obstacle in establishing a ground-truth architecture - the limited availability of engineers who are closely familiar with the system in question. In this paper, we present our experience in recovering the ground-truth architectures of four open-source systems. We discuss the primary insights gained in the process, analyze the characteristics of the obtained ground-truth architectures, and reflect on the involvement of the systems' engineers in a limited but critical fashion. Our findings suggest the practical feasibility of obtaining ground-truth architectures for large systems and encourage future efforts directed at establishing a large scale repository of such architectures. Joshua Garcia, Ivo Krka, Chris Mattmann, Nenad Medvidovic |
ICSE | 3 |
| 2012 | The Earth System Grid Federation: An open infrastructure for access to distributed geospatial dataabstractThe Earth System Grid Federation (ESGF) is a multi-agency, international collaboration that aims at developing the software infrastructure needed to facilitate and empower the study of climate change on a global scale. The ESGF's architecture employs a system of geographically distributed peer nodes, which are independently administered yet united by the adoption of common federation protocols and application programming interfaces (APIs). The cornerstones of its interoperability are the peer-to-peer messaging that is continuously exchanged among all nodes in the federation; a shared architecture and API for search and discovery; and a security infrastructure based on industry standards (OpenID, SSL, GSI and SAML). The ESGF software is developed collaboratively across institutional boundaries and made available to the community as open source. It has now been adopted by multiple Earth science projects and allows access to petabytes of geophysical data, including the entire model output used for the next international assessment report on climate change (IPCC-AR5) and a suite of satellite observations (obs4MIPs) and reanalysis data sets (ANA4MIPs). Luca Cinquini, Daniel J. Crichton, Chris Mattmann, John Harney, Galen M. Shipman, Feiyi Wang, Rachana Ananthakrishnan, Neill Miller, Sebastien Denvil, Mark Morgan, Zed Pobre, Gavin M. Bell, Bob Drach, Dean N. Williams, Philip Kershaw, Stephen Pascoe, Estanislao Gonzalez, Sandro Fiore, Roland Schweitzer |
eScience | 3 |
| 2011 | An informatics architecture for the Virtual Pediatric Intensive Care UnitabstractThe Laura P. and Leland K. Whittier Virtual Pediatric Intensive Care Unit (VPICU) is an ambitious research network focused on building online databases for improving decision-making in pediatric intensive care units. Increasingly, there is a need to unify previously distributed and heterogeneous information captured in these databases to support both traditional retrospective support ad-hoc studies, and ad-hoc analyses. VPICU and NASA's Jet Propulsion Laboratory have constructed a reference architecture and implementation framework that addresses these needs. The architecture is unobtrusive, scalable, and secure, with a strong focus on rapid deployment and integration. This paper reports on the current status of our efforts and details the strength of the framework via our recent work in unsupervised discovery of patient similarity within the hospital. Daniel J. Crichton, Chris Mattmann, Andrew F. Hart, David C. Kale, Robinder G. Khemani, Patrick Ross, Sarah Rubin, Paul Veeravatanayothin, Amy Braverman, Cameron Goodale, Randall C. Wetzel |
CBMS | 2 |
| 2011 | Workshop on software engineering for cloud computing: (SECLOUD 2011)abstractCloud computing is emerging as more than simply a technology platform but a software engineering paradigm for the future. Hordes of cloud computing technologies, techniques, and integration approaches are widely being adopted, taught at the university level, and expected as key skills in the job market. The principles and practices of the software engineering and software architecture community can serve to help guide this emerging domain. The fundamental goal of the ICSE 2011 Software Engineering for Cloud Workshop is to bring together the diverse communities of cloud computing and of software engineering and architecture research with the hopes of sharing and disseminating key tribal knowledge between these rich areas. We expect as the workshop output a set of identified key software engineering challenges and important issues in the domain of cloud computing, specifically focused on how software engineering practice and research can play a role in shaping the next five years of research and practice for clouds. Furthermore, we expect to share "war stories", best practices and lessons learned between leaders in the software engineering and cloud computing communities. Chris Mattmann, Nenad Medvidovic, T. S. Mohan, T. Owen O'Malley |
ICSE | 1 |
| 2011 | Enhancing architectural recovery using concernsabstractArchitectures of implemented software systems tend to drift and erode as they are maintained and evolved. To properly understand such systems, their architectures must be recovered from implementation-level artifacts. Many techniques for architectural recovery have been proposed, but their degrees of automation and accuracy remain unsatisfactory. To alleviate these shortcomings, we present a machine learning-based technique for recovering an architectural view containing a system's components and connectors. Our approach differs from other architectural recovery work in that we rely on recovered software concerns to help identify components and connectors. A concern is a software system's role, responsibility, concept, or purpose. We posit that, by recovering concerns, we can improve the correctness of recovered components, increase the automation of connector recovery, and provide more comprehensible representations of architectures. Joshua Garcia, Daniel Popescu 0001, Chris Mattmann, Nenad Medvidovic, Yuanfang Cai |
ASE | 3 |
| 2010 | Reuse of software assets for the NASA Earth science decadal survey missionsabstractSoftware assets from existing Earth science missions can be reused for the new decadal survey missions that are being planned by NASA in response to the 2007 Earth Science National Research Council (NRC) Study. The new missions will require the development of software to curate, process, and disseminate the data to science users of interest and to the broader NASA mission community. In this paper, we discuss new tools and a blossoming community that are being developed by the Earth Science Data System (ESDS) Software Reuse Working Group (SRWG) to improve capabilities for reusing NASA software assets. Chris Mattmann, Robert R. Downs, James J. Marshall, Neal F. Most, Shahin Samadi |
IGARSS | 1 |
| 2010 | Kadre: domain-specific architectural recovery for scientific software systemsabstractScientists today conduct new research via software-based experimentation and validation in a host of disciplines. Scientific software represents a significant investment due to its complexity and longevity yet there is little reuse of scientific software beyond small libraries which increases development and maintenance costs. To alleviate this disconnect, we have developed KADRE, a domain-specific architecture recovery approach and toolset to aid automatic and accurate identification of workflow components in existing scientific software. KADRE improves upon state of the art general cluster techniques, helping to promote component-based reuse within the domain. David Woollard, Chris Mattmann, Daniel Popescu 0001, Nenad Medvidovic |
ASE | 2 |
| 2009 | Enabling effective curation of cancer biomarker research dataabstractThe dramatic increase in data in the area of cancer research has elevated the importance of effectively managing the quality and consistency of research results from multiple providers. The U.S. National Cancer Institute's Early Detection Research Network (EDRN) is a prime example of a virtual organization, sponsoring distributed, collaborative work at dozens of institutions around the country. As part of a comprehensive informatics infrastructure, The NASA Jet Propulsion Laboratory, in collaboration with Dartmouth Medical School, has developed a web application for the curation of cancer biomarker research results. In this paper, we describe and evaluate the application in the context of the EDRN content management process, and detail our experience using the tool in an operational environment to capture and annotate biomarker research data generated by the EDRN. Andrew F. Hart, Chris Mattmann, John J. Tran, Daniel J. Crichton, J. Steven Hughes, Heather Kincaid, Sean C. Kelly, Kristen Anton, Donald Johnsey, Christos Patriotis |
CBMS | 2 |
| 2007 | A Framework for the Assessment and Selection of Software Components and Connectors in COTS-Based ArchitecturesabstractSoftware systems today are composed from prefabricated commercial components and connectors that provide complex functionality and engage in complex interactions. Unfortunately, because of the distinct assumptions made by developers of these products, successfully integrating them into a software system can be complicated, often causing budget and schedule overruns. A number of integration risks can often be resolved by selecting the 'right' set of COTS components and connectors that can be integrated with minimal effort. In this paper we describe a framework for selecting COTS software components and connectors ensuring their interoperability in software-intensive systems. Our framework is built upon standard definitions of both COTS components and connectors and is intended for use by architects and developers during the design phase of a software system. We highlight the utility of our framework using a challenging example from the data-intensive systems domain. Our preliminary experience in using the framework indicates an increase in interoperability assessment productivity by 50% and accuracy by 20%. Jesal Bhuta, Chris Mattmann, Nenad Medvidovic, Barry W. Boehm |
WICSA | 2 |
| 2006 | A Distributed Information Services Architecture to Support Biomarker Discovery in Early Detection of CancerabstractInformatics in biomedicine is becoming increasingly interconnected via distributed information services, interdisciplinary correlation, and crossinstitutional collaboration. Partnering with NASA, the Early Detection Research Network (EDRN), a program managed by the National Cancer Institute, has been defining and building an informatics architecture to support the discovery of biomarkers in their earliest stages. The architecture established by EDRN serves as a blueprint for constructing a set of services focused on the capture, processing, management and distribution of information through the phases of biomarker discovery and validation. Daniel J. Crichton, Sean C. Kelly, Chris Mattmann, J. Steven Hughes, Jane Oh, Mark Thornquist, Donald Johnsey, Sudhir Srivastava, Laura Esserman, William Bigbee |
e-Science | 3 |
| 2006 | A software architecture-based framework for highly distributed and data intensive scientific applicationsabstractModern scientific research is increasingly conducted by virtual communities of scientists distributed around the world. The data volumes created by these communities are extremely large, and growing rapidly. The management of the resulting highly distributed, virtual data systems is a complex task, characterized by a number of formidable technical challenges, many of which are of a software engineering nature. In this paper we describe our experience over the past seven years in constructing and deploying OODT, a software framework that supports large, distributed, virtual scientific communities. We outline the key software engineering challenges that we faced, and addressed, along the way. We argue that a major contributor to the success of OODT was its explicit focus on software architecture. We describe several large-scale, real-world deployments of OODT, and the manner in which OODT helped us to address the domain-specific challenges induced by each deployment. Chris Mattmann, Daniel J. Crichton, Nenad Medvidovic, Steve Hughes |
ICSE | 1 |
| 2006 | Software Connectors for Highly Distributed and Voluminous Data Intensive SystemsabstractWe describe a research agenda for selecting software connectors which quantifiably satisfy different scenarios for large volume data distribution. We outline the necessity for a framework which allows a user to select amongst the different distribution connectors available. The framework is based on a classification of distribution connectors along eight key dimensions of data distribution Chris Mattmann |
ASE | 1 |
| 2004 | Software Architecture for Large-Scale, Distributed, Data-Intensive SystemsabstractThe sheer amount of data produced by modern science research has created a need for the construction and understanding of "data-intensive systems", large-scale, distributed systems which integrate information. The formal nature of constructing such software systems; however, is relatively unstudied, and has been a large focus of the super-computing and distributed computing communicates, rather than the software engineering communities. These data-intensive systems exhibit characteristics which appear fruitful for research from a software engineering, and software architectural focus. From our experience, the methodologies and notations for design and implementation of data-intensive systems look to be a good starting point for this important research area. This paper presents our experience with OODT (object-oriented data technology), a software architectural style, and middleware-based implementation for data-intensive systems developed and maintained at the Jet Propulsion Laboratory. To date, OODT has been successfully evaluated in several different science domains including Planetary Science with NASA's Planetary Data System (PDS) and Cancer Research with the National Cancer Institute (NCI). Chris Mattmann, Daniel J. Crichton, J. Steven Hughes, Sean C. Kelly, Paul M. Ramirez |
WICSA | 1 |