Andrew J. Connolly

dblp:117/4231 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
0since 2021 · last 2020
0000-0001-5576-8189ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8Artificial intelligence and machine learning · 4Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Machine learning and data management · 76% Data stream processing · 14% Data mining · 5%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Performance modeling and evaluation · 65% High-performance computing · 30% Storage systems · 5%
Interdisciplinary, comprehensive, and emerging computing
5 papers
Computational science and engineering · 62% Bioinformatics and computational biology · 38%
Theoretical computer science
4 papers
Computational geometry · 59% Algorithms and data structures · 30% Mathematical optimization · 11%
Artificial intelligence
1 paper
Trustworthy machine learning · 100%

Topics — the 18 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management
data management for machine learning
0.412020
Toward Sampling for Deep Learning Model Diagnosis · ICDE 2020
Machine learning and data management › model evaluation
model diagnosis
0.412020
Toward Sampling for Deep Learning Model Diagnosis · ICDE 2020
High-performance computing
scientific computing systems
0.322013
A Demonstration of Iterative Parallel Array Processing in Support of Telescope Image Analysis · Proc. VLDB Endow. 2013
Optimizing the computation of n-point correlations on large-scale astronomical data · SC 2012
Performance modeling and evaluation
benchmarking
0.312017
Comparative Evaluation of Big-Data Systems on Scientific Image Analytics Workloads · Proc. VLDB Endow. 2017
Performance modeling and evaluation › benchmarking › distributed system benchmarking
big data system benchmarking
0.312017
Comparative Evaluation of Big-Data Systems on Scientific Image Analytics Workloads · Proc. VLDB Endow. 2017
Computational science and engineering
astronomy
0.232012
Fast algorithms for comprehensive n-point correlation estimates · KDD 2012
A multiple tree algorithm for the efficient association of asteroid observations · KDD 2005
Fast nonlinear regression via eigenimages applied to galactic morphology · KDD 2004
Computational science and engineering
spatial data analysis
0.112012
Fast algorithms for comprehensive n-point correlation estimates · KDD 2012
Computational geometry
spatial statistics
0.112012
Fast algorithms for comprehensive n-point correlation estimates · KDD 2012
Machine learning › Trustworthy machine learning
interpretability
0.112020
Toward Sampling for Deep Learning Model Diagnosis · ICDE 2020
Performance modeling and evaluation
workload characterization
0.112017
Comparative Evaluation of Big-Data Systems on Scientific Image Analytics Workloads · Proc. VLDB Endow. 2017
Data mining
pattern mining
0.112005
Variable KD-Tree Algorithms for Spatial Pattern Search and Discovery · NIPS 2005
Spatial and temporal data management › spatial analysis
spatial pattern matching
0.112005
Variable KD-Tree Algorithms for Spatial Pattern Search and Discovery · NIPS 2005
Computational geometry › spatial data structures
kd-tree
0.112005
Variable KD-Tree Algorithms for Spatial Pattern Search and Discovery · NIPS 2005
Algorithms and data structures › data structure design › search structures
search trees
0.112005
A multiple tree algorithm for the efficient association of asteroid observations · KDD 2005
Computational geometry
spatial data structures
0.112005
Variable KD-Tree Algorithms for Spatial Pattern Search and Discovery · NIPS 2005
Multimedia analysis and retrieval
image analysis
0.012004
Fast nonlinear regression via eigenimages applied to galactic morphology · KDD 2004
Mathematical optimization › statistical estimation › regression
nonlinear regression
0.012004
Fast nonlinear regression via eigenimages applied to galactic morphology · KDD 2004
Algorithms and data structures › similarity search
nearest neighbor search
0.012005
Variable KD-Tree Algorithms for Spatial Pattern Search and Discovery · NIPS 2005

Methods — techniques the papers use, named apart from their topics

deep learning · 0.9data sampling · 0.9comparative evaluation · 0.6parallel array processing · 0.3parallel computation · 0.3multiple tree algorithm · 0.2weighted linear regression · 0.1nearest neighbor search · 0.1eigenimages · 0.1combinatorial search · 0.1
YearPublicationVenuePosition
2020 Toward Sampling for Deep Learning Model Diagnosis
abstract
Deep learning (DL) models have achieved paradigm-changing performance in many fields with high dimensional data, such as images, audio, and text. However, the black-box nature of deep neural networks is not only a barrier to adoption in applications such as medical diagnosis, where interpretability is essential, but it also impedes diagnosis of under performing models. The task of diagnosing or explaining DL models requires the computation of additional artifacts, such as activation values and gradients. These artifacts are large in volume, and their computation, storage, and querying raise significant data management challenges. In this paper, we develop a novel data sampling technique that produces approximate but accurate results for these model debugging queries. Our sampling technique utilizes the lower dimension representation learned by the DL model and focuses on model decision boundaries for the data in this lower dimensional space.
Parmita Mehta, Stephen Portillo, Magdalena Balazinska, Andrew J. Connolly
ICDE4
2017 Comparative Evaluation of Big-Data Systems on Scientific Image Analytics Workloads
abstract
Scientific discoveries are increasingly driven by analyzing large volumes of image data. Many new libraries and specialized database management systems (DBMSs) have emerged to support such tasks. It is unclear how well these systems support real-world image analysis use cases, and how performant the image analytics tasks implemented on top of such systems are. In this paper, we present the first comprehensive evaluation of large-scale image analysis systems using two real-world scientific image data processing use cases. We evaluate five representative systems (SciDB, Myria, Spark, Dask, and TensorFlow) and find that each of them has shortcomings that complicate implementation or hurt performance. Such shortcomings lead to new research opportunities in making large-scale image analysis both efficient and easy to use.
Parmita Mehta, Sven Dorkenwald, Dongfang Zhao 0001, Tomer Kaftan, Alvin Cheung, Magdalena Balazinska, Ariel Rokem, Andrew J. Connolly, Jacob VanderPlas, Yusra AlSayyad
Proc. VLDB Endow.8
2015 Efficient iterative processing in the SciDB parallel array engine
abstract
Many scientific data-intensive applications perform iterative computations on array data. There exist multiple engines specialized for array processing. These engines efficiently support various types of operations, but none includes native support for iterative processing. In this paper, we develop a model for iterative array computations and a series of optimizations. We evaluate the benefits of an optimized, native support for iterative array processing on the SciDB engine and real workloads from the astronomy domain.
Emad Soroush, Magdalena Balazinska, K. Simon Krughoff, Andrew J. Connolly
SSDBM4
2013 A Demonstration of Iterative Parallel Array Processing in Support of Telescope Image Analysis
abstract
In this demonstration, we present AscotDB, a new tool for the analysis of telescope image data. AscotDB results from the integration of ASCOT, a Web-based tool for the collaborative analysis of telescope images and their metadata, and SciDB, a parallel array processing engine. We demonstrate the novel data exploration supported by this integrated tool on a 1 TB dataset comprising scientifically accurate, simulated telescope images. We also demonstrate novel iterative-processing features that we added to SciDB in order to support this use-case.
Matthew I. Moyers, Emad Soroush, Spencer Wallace, K. Simon Krughoff, Jacob VanderPlas, Magdalena Balazinska, Andrew J. Connolly
Proc. VLDB Endow.7
2012 Fast algorithms for comprehensive n-point correlation estimates
abstract
The n-point correlation functions (npcf) are powerful spatial statistics capable of fully characterizing any set of multidimensional points. These functions are critical in key data analyses in astronomy and materials science, among other fields, for example to test whether two point sets come from the same distribution and to validate physical models and theories. For example, the npcf has been used to study the phenomenon of dark energy, considered one of the major breakthroughs in recent scientific discoveries. Unfortunately, directly estimating the continuous npcf at a single value requires O(Nn) time for $N$ points, and n may be 2, 3, 4 or even higher, depending on the sensitivity required. In order to draw useful conclusions about real scientific problems, we must repeat this expensive computation both for many different scales in order to derive a smooth estimate and over many different subsamples of our data in order to bound the variance.
William B. March, Andrew J. Connolly, Alexander G. Gray
KDD2
2012 Optimizing the computation of n-point correlations on large-scale astronomical data
abstract
The n-point correlation functions (npcf) are powerful statistics that are widely used for data analyses in astronomy and other fields. These statistics have played a crucial role in fundamental physical breakthroughs, including the discovery of dark energy. Unfortunately, directly computing the npcf at a single value requires O(Nn) time for N points and values of n of 2, 3, 4, or even larger. Astronomical data sets can contain billions of points, and the next generation of surveys will generate terabytes of data per night. To meet these computational demands, we present a highly-tuned npcf computation code that show an order-of-magnitude speedup over current state-of-the-art. This enables a much larger 3-point correlation computation on the galaxy distribution than was previously possible. We show a detailed performance evaluation on many different architectures.
William B. March, Kenneth Czechowski, Marat Dukhan, Thomas Benson, Dongryeol Lee, Andrew J. Connolly, Richard W. Vuduc, Edmond Chow, Alexander G. Gray
SC6
2011 Towards Efficient and Precise Queries over Ten Million Asteroid Trajectory Models
Yusra AlSayyad, K. Simon Krughoff, Bill Howe, Andrew J. Connolly, Magdalena Balazinska, Lynne Jones
SSDBM4
2005 A multiple tree algorithm for the efficient association of asteroid observations
abstract
In this paper we examine the problem of efficiently finding sets of observations that conform to a given underlying motion model. While this problem is often phrased as a tracking problem, where it is called track initiation, it is useful in a variety of tasks where we want to find correspondences or patterns in spatial-temporal data. Unfortunately, this problem often suffers from a combinatorial explosion in the number of potential sets that must be evaluated. We consider the problem with respect to large-scale asteroid observation data, where the goal is to find associations among the observations that correspond to the same underlying asteroid. In this domain, it is vital that we can efficiently extract the underlying associations. We introduce a new methodology for track initiation that exhaustively considers all possible linkages. We then introduce an exact tree-based algorithm for tractably finding all compatible sets of points. Further, we extend this approach to use multiple trees, exploiting structure from several time steps at once. We compare this approach to a standard sequential approach and show how the use of multiple trees can provide a significant benefit.
Jeremy Kubica, Andrew W. Moore 0001, Andrew J. Connolly, Robert Jedicke
KDD3
2005 Variable KD-Tree Algorithms for Spatial Pattern Search and Discovery
abstract
In this paper we consider the problem of finding sets of points that conform to a given underlying model from within a dense, noisy set of observations. This problem is motivated by the task of efficiently linking faint asteroid detections, but is applicable to a range of spatial queries. We survey current tree-based approaches, showing a trade-off exists between single tree and multiple tree algorithms. To this end, we present a new type of multiple tree algorithm that uses a variable number of trees to exploit the advantages of both approaches. We empirically show that this algorithm performs well using both simulated and astronomical data.
Jeremy Kubica, Joseph Masiero, Andrew W. Moore 0001, Robert Jedicke, Andrew J. Connolly
NIPS5
2004 Fast nonlinear regression via eigenimages applied to galactic morphology
abstract
Astronomy increasingly faces the issue of massive, unwieldly data sets. The Sloan Digital Sky Survey (SDSS) [11] has so far generated tens of millions of images of distant galaxies, of which only a tiny fraction have been morphologically classified. Morphological classification in this context is achieved by fitting a parametric model of galaxy shape to a galaxy image. This is a nonlinear regression problem, whose challenges are threefold, 1) blurring of the image caused by atmosphere and mirror imperfections, 2) large numbers of local minima, and 3) massive data sets.Our strategy is to use the eigenimages of the parametric model to form a new feature space, and then to map both target image and the model parameters into this feature space. In this low-dimensional space we search for the best image-to-parameter match. To search the space, we sample it by creating a database of many random parameter vectors (prototypes) and mapping them into the feature space. The search problem then becomes one of finding the best prototype match, so the fitting process a nearest-neighbor search.In addition to the savings realized by decomposing the original space into an eigenspace, we can use the fact that the model is a linear sum of functions to reduce the prototypes further: the only prototypes stored are the components of the model function. A modified form of nearest neighbor is used to search among them.Additional complications arise in the form of missing data and heteroscedasticity, both of which are addressed with weighted linear regression. Compared to existing techniques, speed-ups ach-ieved are between 2 and 3 orders of magnitude. This should enable the analysis of the entire SDSS dataset.
Brigham S. Anderson, Andrew W. Moore 0001, Andrew J. Connolly, Robert Nichol
KDD3