EDBT 2026 Demo / reviewers in the wild / expert
Alex Szalay
dblp:s/AlexanderSSzalay · also Alexander S. Szalay
· DBLP profile ↗
53ranked-venue papers
7as first author
2since 2021 · last 2023
0000-0002-4108-3282ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 19 · 4 first-authorSystems, architecture and hardware · 17 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 12 · 2 since 2021Software engineering, systems software and programming languages · 7Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorArtificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
9 papers |
Bioinformatics and computational biology · 98% Computational science and engineering · 2% Computing education · 0% | |
| Computer architecture, parallel and distributed computing, and storage systems
14 papers |
Storage systems · 37% High-performance computing · 30% Cloud and datacenter computing · 13% | |
| Databases, data mining, and information retrieval
8 papers |
Query processing and optimization · 30% Distributed and cloud data management · 22% Indexing and storage engines · 19% | |
| Computer graphics and multimedia
2 papers |
Visualization and visual analytics · 52% Rendering · 48% | |
| Computer networks
3 papers |
Internet of things and sensor networks · 96% Transport protocols and congestion control · 4% |
Topics — the 30 heaviest of 61, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
sequence alignment |
1.2 | 2 | 2023 | Short-read aligner performance in germline variant identification · Bioinform. 2023 Performance optimization in DNA short-read alignment · Bioinform. 2022 |
Bioinformatics and computational biology › sequence analysis › read mapping
short read alignment |
1.2 | 2 | 2023 | Short-read aligner performance in germline variant identification · Bioinform. 2023 Performance optimization in DNA short-read alignment · Bioinform. 2022 |
Bioinformatics and computational biology › genomics
variant calling |
0.7 | 1 | 2023 | Short-read aligner performance in germline variant identification · Bioinform. 2023 |
Bioinformatics and computational biology › genomics
genomic data management |
0.4 | 1 | 2019 | The Terabase Search Engine: a large-scale relational database of short-read sequences · Bioinform. 2019 |
Bioinformatics and computational biology › biological database
sequence database |
0.4 | 1 | 2019 | The Terabase Search Engine: a large-scale relational database of short-read sequences · Bioinform. 2019 |
High-performance computing
scientific data management |
0.4 | 3 | 2012 | Data-intensive spatial filtering in large numerical simulation datasets · SC 2012 I/O streaming evaluation of batch queries for data-intensive computational turbulence · SC 2011 JAWS: Job-Aware Workload Scheduling for the Exploration of Turbulence Simulations · SC 2010 |
Bioinformatics and computational biology › epigenomics › DNA methylation › DNA methylation analysis
bisulfite sequencing |
0.3 | 1 | 2018 | Arioc: GPU-accelerated alignment of short bisulfite-treated reads · Bioinform. 2018 |
Bioinformatics and computational biology › sequence analysis
read mapping |
0.3 | 1 | 2018 | Arioc: GPU-accelerated alignment of short bisulfite-treated reads · Bioinform. 2018 |
High-performance computing
streaming i/o |
0.3 | 2 | 2012 | Data-intensive spatial filtering in large numerical simulation datasets · SC 2012 I/O streaming evaluation of batch queries for data-intensive computational turbulence · SC 2011 |
Parallel and multicore computing
graph processing |
0.2 | 1 | 2015 | FlashGraph: Processing Billion-Node Graphs on an Array of Commodity SSDs · FAST 2015 |
Storage systems › flash and SSD
solid-state drive |
0.2 | 1 | 2015 | FlashGraph: Processing Billion-Node Graphs on an Array of Commodity SSDs · FAST 2015 |
Bioinformatics and computational biology › genomics › variant calling
germline variant calling |
0.2 | 1 | 2023 | Short-read aligner performance in germline variant identification · Bioinform. 2023 |
Internet of things and sensor networks
time synchronization |
0.2 | 1 | 2014 | Robust time synchronization in wireless sensor networks using real time clock · SenSys 2014 |
Internet of things and sensor networks
wireless sensor network |
0.2 | 1 | 2014 | Robust time synchronization in wireless sensor networks using real time clock · SenSys 2014 |
High-performance computing
performance optimization at scale |
0.2 | 1 | 2022 | Performance optimization in DNA short-read alignment · Bioinform. 2022 |
Storage systems
flash and SSD |
0.2 | 1 | 2013 | Toward millions of file system IOPS on low-cost, commodity hardware · SC 2013 |
Storage systems › file systems › file system design
user space file system |
0.2 | 1 | 2013 | Toward millions of file system IOPS on low-cost, commodity hardware · SC 2013 |
Visualization and visual analytics
flow visualization |
0.1 | 1 | 2012 | Turbulence Visualization at the Terascale on Desktop PCs · IEEE Trans. Vis. Comput. Graph. 2012 |
Rendering › volume rendering › ray casting
GPU ray-casting |
0.1 | 1 | 2012 | Turbulence Visualization at the Terascale on Desktop PCs · IEEE Trans. Vis. Comput. Graph. 2012 |
Visualization and visual analytics › flow visualization
turbulent flow visualization |
0.1 | 1 | 2012 | Turbulence Visualization at the Terascale on Desktop PCs · IEEE Trans. Vis. Comput. Graph. 2012 |
Rendering
volume rendering |
0.1 | 1 | 2012 | Turbulence Visualization at the Terascale on Desktop PCs · IEEE Trans. Vis. Comput. Graph. 2012 |
High-performance computing
parallel i/o |
0.1 | 1 | 2012 | Data-intensive spatial filtering in large numerical simulation datasets · SC 2012 |
Hardware accelerators and domain-specific architectures
query processing |
0.1 | 1 | 2012 | Data-intensive spatial filtering in large numerical simulation datasets · SC 2012 |
Storage systems
top-k query processing |
0.1 | 1 | 2012 | Just-in-Time Analytics on Large File Systems · IEEE Trans. Computers 2012 |
Storage systems
data analytics |
0.1 | 1 | 2011 | Just-in-Time Analytics on Large File Systems · FAST 2011 |
Storage systems
file systems |
0.1 | 1 | 2011 | Just-in-Time Analytics on Large File Systems · FAST 2011 |
Cloud and datacenter computing
cloud migration |
0.1 | 1 | 2010 | Migrating a (large) science database to the cloud · HPDC 2010 |
Cloud and datacenter computing
database migration |
0.1 | 1 | 2010 | Migrating a (large) science database to the cloud · HPDC 2010 |
Storage systems
data management |
0.1 | 1 | 2010 | An overview of the Open Science Data Cloud · HPDC 2010 |
Storage systems
i/o optimization |
0.1 | 1 | 2010 | JAWS: Job-Aware Workload Scheduling for the Exploration of Turbulence Simulations · SC 2010 |
Methods — techniques the papers use, named apart from their topics
benchmarking · 1.8performance profiling · 1.1relational indexing · 0.8GPU acceleration · 0.7real time clock · 0.4offline time synchronization · 0.4summed volumes · 0.3decomposable kernel evaluation · 0.3prototype system · 0.2set-associative caching · 0.2asynchronous i/o · 0.2wavelet compression · 0.1run-length encoding · 0.1just-in-time sampling · 0.1entropy encoding · 0.1partial sums · 0.1distributed query evaluation · 0.1workload-aware batching · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Short-read aligner performance in germline variant identificationabstractMOTIVATION: Read alignment is an essential first step in the characterization of DNA sequence variation. The accuracy of variant-calling results depends not only on the quality of read alignment and variant-calling software but also on the interaction between these complex software tools. RESULTS: In this review, we evaluate short-read aligner performance with the goal of optimizing germline variant-calling accuracy. We examine the performance of three general-purpose short-read aligners-BWA-MEM, Bowtie 2, and Arioc-in conjunction with three germline variant callers: DeepVariant, FreeBayes, and GATK HaplotypeCaller. We discuss the behavior of the read aligners with regard to the data elements on which the variant callers rely, and illustrate how the runtime configurations of these software tools combine to affect variant-calling performance. AVAILABILITY AND IMPLEMENTATION: The quick brown fox jumps over the lazy dog. Richard Wilton, Alex Szalay |
Bioinform. | 2 |
| 2022 | Performance optimization in DNA short-read alignmentabstractSUMMARY: Over the past decade, short-read sequence alignment has become a mature technology. Optimized algorithms, careful software engineering and high-speed hardware have contributed to greatly increased throughput and accuracy. With these improvements, many opportunities for performance optimization have emerged. In this review, we examine three general-purpose short-read alignment tools-BWA-MEM, Bowtie 2 and Arioc-with a focus on performance optimization. We analyze the performance-related behavior of the algorithms and heuristics each tool implements, with the goal of arriving at practical methods of improving processing speed and accuracy. We indicate where an aligner's default behavior may result in suboptimal performance, explore the effects of computational constraints such as end-to-end mapping and alignment scoring threshold, and discuss sources of imprecision in the computation of alignment scores and mapping quality. With this perspective, we describe an approach to tuning short-read aligner performance to meet specific data-analysis and throughput requirements while avoiding potential inaccuracies in subsequent analysis of alignment results. Finally, we illustrate how this approach avoids easily overlooked pitfalls and leads to verifiable improvements in alignment speed and accuracy. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Appendices referenced in this article are available at Bioinformatics online. Richard Wilton, Alex Szalay |
Bioinform. | 2 |
| 2020 | Sketch and Scale Geo-distributed tSNE and UMAPabstractRunning machine learning analytics over geographically distributed datasets is a rapidly arising problem in the world of data management policies ensuring privacy and data security. Visualizing high dimensional data using tools such as t-distributed Stochastic Neighbor Embedding (tSNE) and Uniform Manifold Approximation and Projection (UMAP) became a common practice for data scientists. Both tools scale poorly in time and memory. While recent optimizations showed successful handling of 10,000 data points, scaling beyond million points is still challenging. We introduce a novel framework: Sketch and Scale (SnS). It leverages a Count Sketch data structure to compress the data on the edge nodes, aggregates the reduced size sketches on the master node, and runs vanilla tSNE or UMAP on the summary, representing the densest areas, extracted from the aggregated sketch.We show this technique to be fully parallel, scale linearly in time, logarithmically in memory and communication, making it possible to analyze datasets with many millions, potentially billions of data points, spread across several data centers around the globe. We demonstrate the power of our method on two mid-size datasets: cancer data with 52 million 35-band pixels from multiplex images of tumor biopsies; and astrophysics data of 100 million stars with multi-color photometry from the Sloan Digital Sky Survey (SDSS). Viska Wei, Nikita Ivkin, Vladimir Braverman, Alex Szalay |
IEEE BigData | 4 |
| 2020 | Arioc: High-concurrency short-read alignment on multiple GPUsabstractIn large DNA sequence repositories, archival data storage is often coupled with computers that provide 40 or more CPU threads and multiple GPU (general-purpose graphics processing unit) devices. This presents an opportunity for DNA sequence alignment software to exploit high-concurrency hardware to generate short-read alignments at high speed. Arioc, a GPU-accelerated short-read aligner, can compute WGS (whole-genome sequencing) alignments ten times faster than comparable CPU-only alignment software. When two or more GPUs are available, Arioc's speed increases proportionately because the software executes concurrently on each available GPU device. We have adapted Arioc to recent multi-GPU hardware architectures that support high-bandwidth peer-to-peer memory accesses among multiple GPUs. By modifying Arioc's implementation to exploit this GPU memory architecture we obtained a further 1.8x-2.9x increase in overall alignment speeds. With this additional acceleration, Arioc computes two million short-read alignments per second in a four-GPU system; it can align the reads from a human WGS sequencer run-over 500 million 150nt paired-end reads-in less than 15 minutes. As WGS data accumulates exponentially and high-concurrency computational resources become widespread, Arioc addresses a growing need for timely computation in the short-read data analysis toolchain. Richard Wilton, Alex Szalay |
PLoS Comput. Biol. | 2 |
| 2019 | The Terabase Search Engine: a large-scale relational database of short-read sequencesabstractMOTIVATION: DNA sequencing archives have grown to enormous scales in recent years, and thousands of human genomes have already been sequenced. The size of these data sets has made searching the raw read data infeasible without high-performance data-query technology. Additionally, it is challenging to search a repository of short-read data using relational logic and to apply that logic across samples from multiple whole-genome sequencing samples. RESULTS: We have built a compact, efficiently-indexed database that contains the raw read data for over 250 human genomes, encompassing trillions of bases of DNA, and that allows users to search these data in real-time. The Terabase Search Engine enables retrieval from this database of all the reads for any genomic location in a matter of seconds. Users can search using a range of positions or a specific sequence that is aligned to the genome on the fly. AVAILABILITY AND IMPLEMENTATION: Public access to the Terabase Search Engine database is available at http://tse.idies.jhu.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Richard Wilton, Sarah J. Wheelan, Alex Szalay, Steven Salzberg |
Bioinform. | 3 |
| 2018 | Database-Centric Scientific Computing - (In Memoriam Jim Gray)
Alex Szalay |
ADBIS | 1 |
| 2018 | Arioc: GPU-accelerated alignment of short bisulfite-treated readsabstractMotivation: The alignment of bisulfite-treated DNA sequences (BS-seq reads) to a large genome involves a significant computational burden beyond that required to align non-bisulfite-treated reads. In the analysis of BS-seq data, this can present an important performance bottleneck that can be mitigated by appropriate algorithmic and software-engineering improvements. One strategy is to modify the read-alignment algorithms by integrating the logic related to BS-seq alignment, with the goal of making the software implementation amenable to optimizations that lead to higher speed and greater sensitivity than might otherwise be attainable. Results: We evaluated this strategy using Arioc, a short-read aligner that uses GPU (general-purpose graphics processing unit) hardware to accelerate computationally-expensive programming logic. We integrated the BS-seq computational logic into both GPU and CPU code throughout the Arioc implementation. We then carried out a read-by-read comparison of Arioc's reported alignments with the alignments reported by well-known CPU-based BS-seq read aligners. With simulated reads, Arioc's accuracy is equal to or better than the other read aligners we evaluated. With human sequencing reads, Arioc's throughput is at least 10 times faster than existing BS-seq aligners across a wide range of sensitivity settings. Availability and implementation: The Arioc software is available for download at https://github.com/RWilton/Arioc. It is released under a BSD open-source license. Supplementary information: Supplementary data are available at Bioinformatics online. Richard Wilton, Xin Li 0131, Andrew P. Feinberg, Alex Szalay |
Bioinform. | 4 |
| 2016 | DOT-K: A distributed online top-K elements algorithm using extreme value statisticsabstractExtremely large (peta-scale) data collections are generally partitioned into millions of containers (disks/volumes/files) which are essentially unmovable due to their aggregate size. They are stored over a large distributed cloud of machines, with computing co-located with the data. Given this data layout, even simple tasks are difficult to perform and naive algorithms can easily become quite expensive. We present a one pass, communications-efficient technique useful for both estimating upper order quantiles and selecting the largest k elements across a highly distributed dataset or stream. Our novel approach draws its foundations from Extreme Value Statistics (EVS) to reason about the statistical relationships between the tail distributions of dataset partitions. The tail of each partition is fitted by the Generalized Pareto Distribution, which captures threshold exceedances. The obtained parameters are communicated to a central coordinator and used to estimate quantiles, or solve for a threshold above which there are approximately k elements. We discuss the computational and bandwidth costs of the algorithm, and demonstrate the accuracy of the method on both a variety of synthetic datasets and a PageRank dataset. Nicholas Carey, Tamás Budavári, Yanif Ahmad, Alex Szalay |
eScience | 4 |
| 2016 | A fast algorithm for neutrally-buoyant Lagrangian particles in numerical ocean modelingabstractAs numerical ocean simulations become more realistic, analysis of their output is increasingly time consuming. As part of a larger effort to make ocean model output more accessible for analysis, we have developed a fast particle-tracking algorithm for exploring the kinematics of the simulated flow. The algorithm is independent of operating system and fully vectorized and parallelized. Furthermore, it enables sliding of particles along solid boundaries according to the 3D along-boundary flow field. The new algorithm can easily simulate several million particle trajectories, thus opening the way for Lagrangian analysis of large-scale and multi-time scale oceanic phenomena. Renske Gelderloos, Alex Szalay, Thomas W. N. Haine, Gerard Lemson |
eScience | 2 |
| 2015 | Streaming Algorithms for Halo FindersabstractCosmological N-body simulations are essential for studies of the large-scale distribution of matter and galaxies in the Universe. This analysis often involves finding clusters of particles and retrieving their properties. Detecting such "halos" among a very large set of particles is a computationally intensive problem, usually executed on the same super-computers that produced the simulations, requiring huge amounts of memory. Recently, a new area of computer science emerged. This area, called streaming algorithms, provides new theoretical methods to compute data analytics in a scalable way using only a single pass over a data sets and logarithmic memory. The main contribution of this paper is a novel connection between the N-body simulations and the streaming algorithms. In particular, we investigate a link between halo finders and the problem of finding frequent items (heavy hitters) in a data stream, that should greatly reduce the computational resource requirements, especially the memory needs. Based on this connection, we can build a new halo finder by running efficient heavy hitter algorithms as a black-box. We implement two representatives of the family of heavy hitter algorithms, the Count-Sketch algorithm (CS) and the Pick-and-Drop sampling (PD), and evaluate their accuracy and memory usage. Comparison with other halo-finding algorithms from [1] shows that our halo finder can locate the largest haloes using significantly smaller memory space and with comparable running time. This streaming approach makes it possible to run and analyze extremely large data sets from N-body simulations on a smaller machine, rather than on supercomputers. Our findings demonstrate the connection between the halo search problem and streaming algorithms as a promising initial direction of further research. Zaoxing Liu, Nikita Ivkin, Lin Yang 0011, Mark Neyrinck, Gerard Lemson, Alex Szalay, Vladimir Braverman, Tamás Budavári, Randal C. Burns |
e-Science | 6 |
| 2015 | FlashGraph: Processing Billion-Node Graphs on an Array of Commodity SSDs
Da Zheng 0004, Disa Mhembere, Randal C. Burns, Joshua T. Vogelstein, Carey E. Priebe, Alex Szalay |
FAST | 6 |
| 2014 | Robust time synchronization in wireless sensor networks using real time clockabstractTime synchronization is an essential service in many sensor network applications. Harsh environment which causes nodes to fail, go offline, or reboot can challenge many time synchronization protocols. In this work, we first characterize this challenge and use a real time clock in one of the nodes in the network to improve robustness of time synchronization. Our experiments show that our approach improves the robustness of state-of-the-art offline time synchronization protocols. Hessam Mohammadmoradi, Omprakash Gnawali, Nir Rattner, Andreas Terzis, Alex Szalay |
SenSys | 5 |
| 2014 | Point cloud databasesabstractWe introduce the concept of the point cloud database, a new kind of database system aimed primarily towards scientific applications. Many scientific observations, experiments, feature extraction algorithms and large-scale simulations produce enormous amounts of data that are better represented as sparse (but often highly-clustered) points in a k-dimensional (k ≲ 10) metric space than on a multi-dimensional grid. Dimensionality reduction techniques, such as principal components, are also widely-used to project high dimensional data into similarly low dimensional spaces. Analysis techniques developed to work on multi-dimensional data points are usually implemented as in-memory algorithms and need to be modified to work in distributed cluster environments and on large amounts of disk-resident data. We conclude that the relational model, with certain additions, is appropriate for point clouds, but point cloud databases must also provide unique set of spatial search and proximity join operators, indexing schemes, and query language constructs that make them a distinct class of database systems. László Dobos 0001, István Csabai, János M. Szalai-Gindl, Tamás Budavári, Alex Szalay |
SSDBM | 5 |
| 2014 | Efficient classification of billions of points into complex geographic regions using hierarchical triangular meshabstractWe present a case study about the spatial indexing and regional classification of billions of geographic coordinates from geo-tagged social network data using Hierarchical Triangular Mesh (HTM) implemented for Microsoft SQL Server. Due to the lack of certain features of the HTM library, we use it in conjunction with the GIS functions of SQL Server to significantly increase the efficiency of pre-filtering of spatial filter and join queries. For example, we implemented a new algorithm to compute the HTM tessellation of complex geographic regions and precomputed the intersections of HTM triangles and geographic regions for faster false-positive filtering. With full control over the index structure, HTM-based pre-filtering of simple containment searches outperforms SQL Server spatial indices by a factor of ten and HTM-based spatial joins run about a hundred times faster. Dániel Kondor, László Dobos 0001, István Csabai, András Bodor, Gábor Vattay, Tamás Budavári, Alex Szalay |
SSDBM | 7 |
| 2013 | Toward millions of file system IOPS on low-cost, commodity hardwareabstractWe describe a storage system that removes I/O bottlenecks to achieve more than one million IOPS based on a user-space file abstraction for arrays of commodity SSDs. The file abstraction refactors I/O scheduling and placement for extreme parallelism and non-uniform memory and I/O. The system includes a set-associative, parallel page cache in the user space. We redesign page caching to eliminate CPU overhead and lock-contention in non-uniform memory architecture machines. We evaluate our design on a 32 core NUMA machine with four, eight-core processors. Experiments show that our design delivers 1.23 million 512-byte read IOPS. The page cache realizes the scalable IOPS of Linux asynchronous I/O (AIO) and increases user-perceived I/O performance linearly with cache hit rates. The parallel, set-associative cache matches the cache hit rates of the global Linux page cache under real workloads. Da Zheng 0004, Randal C. Burns, Alex Szalay |
SC | 3 |
| 2013 | The open connectome project data cluster: scalable analysis and vision for high-throughput neuroscienceabstract- neural connectivity maps of the brain-using the parallel execution of computer vision algorithms on high-performance compute clusters. These services and open-science data sets are publicly available at openconnecto.me. The system design inherits much from NoSQL scale-out and data-intensive computing architectures. We distribute data to cluster nodes by partitioning a spatial index. We direct I/O to different systems-reads to parallel disk arrays and writes to solid-state storage-to avoid I/O interference and maximize throughput. All programming interfaces are RESTful Web services, which are simple and stateless, improving scalability and usability. We include a performance evaluation of the production system, highlighting the effec-tiveness of spatial data organization. Randal C. Burns, Kunal Lillaney, Daniel R. Berger, Logan Grosenick, Karl Deisseroth, R. Clay Reid, William R. Gray Roncal, Priya Manavalan, Davi Bock, Narayanan Kasthuri, Michael M. Kazhdan, Stephen J. Smith, Dean Kleissas, Eric A. Perlman, Kwanghun Chung, Nicholas C. Weiler, Jeff Lichtman, Alex Szalay, Joshua T. Vogelstein, R. Jacob Vogelstein |
SSDBM | 18 |
| 2013 | Inverted indices for particle tracking in petascale cosmological simulationsabstractWe describe the challenges arising from tracking dark matter particles in state of the art cosmological simulations. We are in the process of running the Indra suite of simulations, with an aggregate count of more than 35 trillion particles and 1.1PB of total raw data volume. However, it is not enough just to store the particle positions and velocities in an efficient manner -- analyses also need to be able to track individual particles efficiently through the temporal history of the simulation. The required inverted indices can easily have raw sizes comparable to the original simulation. Daniel Crankshaw, Randal C. Burns, Bridget Falck, Tamás Budavári, Alex Szalay, Jie Wang 0075 |
SSDBM | 5 |
| 2013 | Graywulf: a platform for federated scientific databases and servicesabstractMany fields of science rely on relational database management systems to analyze, publish and share data. Since RDBMS are originally designed for, and their development directions are primarily driven by, business use cases they often lack features very important for scientific applications. Horizontal scalability is probably the most important missing feature which makes it challenging to adapt traditional relational database systems to the ever growing data sizes. Due to the limited support of array data types and metadata management, successful application of RDBMS in science usually requires the development of custom extensions. While some of these extensions are specific to the field of science, the majority of them could easily be generalized and reused in other disciplines. With the Graywulf project we intend to target several goals. We are building a generic platform that offers reusable components for efficient storage, transformation, statistical analysis and presentation of scientific data stored in Microsoft SQL Server. Graywulf also addresses the distributed computational issues arising from current RDBMS technologies. The current version supports load balancing of simple queries and parallel execution of partitioned queries over a set of mirrored databases. Uniform user access to the data is provided through a web based query interface and a data surface for software clients. Queries are formulated in a slightly modified syntax of SQL that offers a transparent view of the distributed data. The software library consists of several components that can be reused to develop complex scientific data warehouses: a system registry, administration tools to manage entire database server clusters, a sophisticated workflow execution framework, and a SQL parser library. László Dobos 0001, István Csabai, Alex Szalay, Tamás Budavári, Nolan Li |
SSDBM | 3 |
| 2013 | Adaptive exploration for large-scale protein analysis in the molecular dynamics databaseabstractMolecular dynamics (MD) simulations generate detailed time-series data of all-atom motions. These simulations are leading users of the world's most powerful supercomputers, and are standard-bearers for a wide range of high-performance computing (HPC) methods. However, MD data exploration and analysis is in its infancy in terms of scalability, ease-of-use, and ultimately its ability to answer 'grand challenge' science questions. This demonstration introduces the Molecular Dynamics Database (MDDB) project at Johns Hopkins, to study the co-design of database methods for deep on-the-fly exploratory MD analyses with HPC simulations. Data exploration in MD suffers from a "human bottleneck", where the laborious administration of simulations leaves little room for domain experts to focus on tackling science questions. MDDB exploits the data-rich nature of MD simulations to provide adaptive control of the exploration process with machine learning techniques, specifically reinforcement learning (RL). We present MDDB's data and queries, architecture, and its use of RL methods. Our audience will co-operate with our steering algorithm and science partners, and witness MDDB's abilities to significantly reduce exploration times and direct computation resources to where they best address science questions. Sarana Nutanong, Nick Carey, Yanif Ahmad, Alex Szalay, Thomas B. Woolf |
SSDBM | 4 |
| 2012 | A Parallel Page Cache: IOPS and Caching for Multicore Systems
Da Zheng 0004, Randal C. Burns, Alex Szalay |
HotStorage | 3 |
| 2012 | The Future of Scientific Data BasesabstractFor many decades, users in scientific fields (domain scientists) have resorted to either home-grown tools or legacy software for the management of their data. Technological advancements nowadays necessitate many of the properties such as data independence, scalability, and functionality found in the roadmap of DBMS technology, DBMS products, however, are not yet ready to address scientific application and user needs. Recent efforts toward building a science DBMS indicate that there is a long way ahead of us, paved by a research agenda that is rich in interesting and challenging problems. Michael Stonebraker, Anastasia Ailamaki, Jeremy Kepner, Alex Szalay |
ICDE | 4 |
| 2012 | Data-intensive spatial filtering in large numerical simulation datasetsabstractWe present a query processing framework for the efficient evaluation of spatial filters on large numerical simulation datasets stored in a data-intensive cluster. Previously, filtering of large numerical simulations stored in scientific databases has been impractical owing to the immense data requirements. Rather, filtering is done during simulation or by loading snapshots into the aggregate memory of an HPC cluster. Our system performs filtering within the database and supports large filter widths. We present two complementary methods of execution: I/O streaming computes a batch filter query in a single sequential pass using incremental evaluation of decomposable kernels, summed volumes generates an intermediate data set and evaluates each filtered value by accessing only eight points in this dataset. We dynamically choose between these methods depending upon workload characteristics. The system allows us to perform filters against large data sets with little overhead: query performance scales with the cluster's aggregate I/O throughput. Kalin Kanov, Randal C. Burns, Gregory L. Eyink, Charles Meneveau, Alex Szalay |
SC | 5 |
| 2012 | SkyQuery: An Implementation of a Parallel Probabilistic Join Engine for Cross-Identification of Multiple Astronomical Databases
László Dobos 0001, Tamás Budavári, Nolan Li, Alex Szalay, István Csabai |
SSDBM | 4 |
| 2012 | Just-in-Time Analytics on Large File SystemsabstractAs file systems reach the petabytes scale, users and administrators are increasingly interested in acquiring high-level analytical information for file management and analysis. Two particularly important tasks are the processing of aggregate and top-k queries which, unfortunately, cannot be quickly answered by hierarchical file systems such as ext3 and NTFS. Existing preprocessing-based solutions, e.g., file system crawling and index building, consume a significant amount of time and space (for generating and maintaining the indexes) which in many cases cannot be justified by the infrequent usage of such solutions. In this paper, we advocate that user interests can often be sufficiently satisfied by approximate-i.e., statistically accurate-answers. We develop Glance, a just-in-time sampling-based system which, after consuming a small number of disk accesses, is capable of producing extremely accurate answers for a broad class of aggregate and top-k queries over a file system without the requirement of any prior knowledge. We use a number of real-world file systems to demonstrate the efficiency, accuracy, and scalability of Glance. H. Howie Huang, Nan Zhang 0004, Wei Wang 0082, Gautam Das 0001, Alex Szalay |
IEEE Trans. Computers | 5 |
| 2012 | Turbulence Visualization at the Terascale on Desktop PCsabstractDespite the ongoing efforts in turbulence research, the universal properties of the turbulence small-scale structure and the relationships between small- and large-scale turbulent motions are not yet fully understood. The visually guided exploration of turbulence features, including the interactive selection and simultaneous visualization of multiple features, can further progress our understanding of turbulence. Accomplishing this task for flow fields in which the full turbulence spectrum is well resolved is challenging on desktop computers. This is due to the extreme resolution of such fields, requiring memory and bandwidth capacities going beyond what is currently available. To overcome these limitations, we present a GPU system for feature-based turbulence visualization that works on a compressed flow field representation. We use a wavelet-based compression scheme including run-length and entropy encoding, which can be decoded on the GPU and embedded into brick-based volume ray-casting. This enables a drastic reduction of the data to be streamed from disk to GPU memory. Our system derives turbulence properties directly from the velocity gradient tensor, and it either renders these properties in turn or generates and renders scalar feature volumes. The quality and efficiency of the system is demonstrated in the visualization of two unsteady turbulence simulations, each comprising a spatio-temporal resolution of 10244. On a desktop computer, the system can visualize each time step in 5 seconds, and it achieves about three times this rate for the visualization of a scalar feature volume. Marc Treib, Kai Bürger, Florian Reichl, Charles Meneveau, Alex Szalay, Rüdiger Westermann |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2011 | Just-in-Time Analytics on Large File Systems
H. Howie Huang, Nan Zhang 0004, Wei Wang 0082, Gautam Das 0001, Alex Szalay |
FAST | 5 |
| 2011 | Performance modeling and analysis of flash-based storage devicesabstractFlash-based solid-state drives (SSDs) will become key components in future storage systems. An accurate performance model will not only help understand the state-of-the-art of SSDs, but also provide the research tools for exploring the design space of such storage systems. Although over the years many performance models were developed for hard drives, the architectural differences between two device families prevent these models from being effective for SSDs. The hard drive performance models cannot account for several unique characteristics of SSDs, e.g., low latency, slow update, and expensive block-level erase. In this paper, we utilize the black-box modeling approach to analyze and evaluate SSD performance, including latency, bandwidth, and throughput, as it requires minimal a priori information about the storage devices. We construct the black-box models, using both synthetic workloads and real-world traces, on three SSDs, as well as an SSD RAID. We find that, while the black-box approach may produce less desirable performance predictions for hard disks, a black-box SSD model with a comprehensive set of workload characteristics can produce accurate predictions for latency, bandwidth, and throughput with small errors. H. Howie Huang, Alex Szalay, Andreas Terzis |
MSST | 3 |
| 2011 | MPI-DB, A Parallel Database Services Software Library for Scientific Computing
Edward Givelberg, Alex Szalay, Kalin Kanov, Randal C. Burns |
EuroMPI | 2 |
| 2011 | I/O streaming evaluation of batch queries for data-intensive computational turbulenceabstractWe describe a method for evaluating computational turbulence queries, including Lagrange Polynomial interpolation, based on partial sums that allows the underlying data to be accessed in any order and in parts. We exploit these properties to stream data from disk in a single pass and concurrently evaluate batch queries. The combination of sequential I/O and data sharing improves performance by an order of magnitude when compared with direct evaluation of each query. The technique also supports distributed evaluation of queries in a database cluster, assembling the partial sums from each node at the query mediator. Interpolation is fundamental to computational turbulence, over 95% of queries use these routines, and the partial sums method allows the JHU Turbulence Database Cluster to realize scale and throughput for our scientists' data-intensive workloads. Kalin Kanov, Eric A. Perlman, Randal C. Burns, Yanif Ahmad, Alex Szalay |
SC | 5 |
| 2011 | Implementing a General Spatial Indexing Library for Relational Databases of Large Numerical Simulations
Gerard Lemson, Tamás Budavári, Alex Szalay |
SSDBM | 3 |
| 2010 | Phoenix: An Epidemic Approach to Time Reconstruction
Jayant Gupchup, Douglas Carlson, Razvan Musaloiu-Elefteri, Alex Szalay, Andreas Terzis |
EWSN | 4 |
| 2010 | An overview of the Open Science Data CloudabstractThe Open Science Data Cloud is a distributed cloud based infrastructure for managing, analyzing, archiving and sharing scientific datasets. We introduce the Open Science Data Cloud, give an overview of its architecture, provide an update on its current status, and briefly describe some research areas of relevance. Robert L. Grossman, Yunhong Gu, Joe Mambretti, Michal Sabala, Alex Szalay, Kevin P. White |
HPDC | 5 |
| 2010 | Migrating a (large) science database to the cloudabstractWe report on attempts to put an existing scientific (astronomical) database -- the Sloan Digital Sky Survey (SDSS) science archive [1] - in the cloud. Based on our experience, it is either very frustrating or impossible at this time to migrate an existing, complex SQL Server database into current cloud service offerings such as Amazon (EC2) and Microsoft (SQL Azure). Certainly it is impossible to migrate a large database in excess of a TB, but even with (much) smaller databases, the limitations of cloud services make it very difficult to migrate the data to the cloud without making changes to the schema and settings (for example, inability to migrate a spatial indexing library, and several other user-defined functions and stored procedures) that would invalidate performance comparisons between cloud and on-premise versions. So it is not surprising that our preliminary performance comparisons show a very large (an order of magnitude) performance discrepancy with the Amazon cloud version of the SDSS database. We have also not yet investigated the performance tweaks that could be possible within the cloud. Ani Thakar, Alex Szalay |
HPDC | 2 |
| 2010 | JAWS: Job-Aware Workload Scheduling for the Exploration of Turbulence SimulationsabstractWe present JAWS, a job-aware, data-driven batch scheduler that improves query throughput for data-intensive scientific database clusters. As datasets reach petabyte-scale, workloads that scan through vast amounts of data to extract features are gaining importance in the sciences. However, acute performance bottlenecks result when multiple queries execute simultaneously and compete for I/O resources. Our solution, JAWS, divides queries into I/O-friendly sub-queries for scheduling. It then identifies overlapping data requirements within the workload and executes sub-queries in batches to maximize data sharing and reduce redundant I/O. JAWS extends our previous work by supporting workflows in which queries exhibit data dependencies, exploiting workload knowledge to coordinate caching decisions, and combating starvation through adaptive and incremental trade-offs between query throughput and response time. Instrumenting JAWS in the Turbulence Database Cluster yields nearly three-fold improvement in query throughput when contention in the workload is high. Eric A. Perlman, Randal C. Burns, Tanu Malik, Tamás Budavári, Charles Meneveau, Alex Szalay |
SC | 7 |
| 2010 | VisWeek Capstone Address
Alex Szalay |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2009 | Building Reliable Data Pipelines for Managing Community Data Using Scientific WorkflowsabstractThe growing amount of scientific data from sensors and field observations is posing a challenge to ¿data valets¿ responsible for managing them in data repositories. These repositories built on commodity clusters need to reliably ingest data continuously and ensure its availability to a wide user community. Workflows provide several benefits to modeling data-intensive science applications and many of these benefits can help manage the data ingest pipelines too. But using workflows is not panacea in itself and data valets need to consider several issues when designing workflows that behave reliably on fault prone hardware while retaining the consistency of the scientific data. In this paper, we propose workflow designs for reliable data ingest in a distributed environment and identify workflow framework features to support resilience. We illustrate these using the data pipeline for the Pan-STARRS repository, one of the largest digital surveys that accumulates 100TB of data annually to support 300 astronomers. Yogesh L. Simmhan, Catharine van Ingen, Alex Szalay, Roger S. Barga, Jim Heasley |
eScience | 3 |
| 2009 | Introduction
Alex Szalay, Djoerd Hiemstra, Alfons Kemper, Manuel Prieto 0001 |
Euro-Par | 1 |
| 2009 | Sundial: Using Sunlight to Reconstruct Global Timestamps
Jayant Gupchup, Razvan Musaloiu-Elefteri, Alex Szalay, Andreas Terzis |
EWSN | 3 |
| 2008 | On Building Scientific Workflow Systems for Data Management in the CloudabstractScientific workflows have become an archetype to model in silico experiments in the Cloud by scientists. There is a class of workflows that are used to by "data valets" to prepare raw data from scientific instruments into a science-ready form for use by scientists. These share data-intensive traits with traditional scientific workflows, yet differ significantly, for example, in the required degree of reliability and the type of provenance collected. We compare and contrast science application and data valet workflows through exemplar eScience projects to drive shared and unique requirements for scientific workflows across diverse users in a Science Cloud. Yogesh L. Simmhan, Roger S. Barga, Catharine van Ingen, Edward D. Lazowska, Alex Szalay |
eScience | 5 |
| 2008 | New Challenges in Petascale Scientific Databases
Alex Szalay |
SSDBM | 1 |
| 2007 | Spatial Indexing of Large Multidimensional Databases
István Csabai, Márton Trencséni, László Dobos 0001, Peter Jozsa, Geza Herczegh, Norbert Purger, Tamás Budavári, Alex Szalay |
CIDR | 8 |
| 2006 | Distributing the Sloan Digital Sky Survey Using UDT and SectorabstractIn this paper, we describe a peer-to-peer storage system called Sector that is designed to access and transport large data sets over wide area high performance networks. We also describe our recent experience using Sector to distribute the Sloan Digital Sky Survey BESTDR4 catalog data. Yunhong Gu, Robert L. Grossman, Alex Szalay, Ani Thakar |
e-Science | 3 |
| 2006 | Bandwidth challenge - Transporting sloan digital sky survey data using SECTORabstractNational Center for Data Mining at UICIn our SC06 BWC entry, we will transfer SDSS (Sloan Digital Sky Survey) Data Release 5 (DR5) between the SC06 show floor in Tampa and one of the NCDM labs on the UIC campus. We will use SECTOR, our newly developed distributed data space management system, to transfer DR5 in parallel between two Linux clusters in Tampa and Chicago, respectively. SECTOR transparently manages the file locating and data moving, while it employs UDT for actual data transfer. The data transfer will be from disk to disk over a 10Gb/s shared, router link between SC06 and UIC, via StarLight. We expect to reach 5Gb/s disk-to-disk data transfer rate between the two sites. Robert L. Grossman, Yunhong Gu, Michal Sabala, Shirley Connelly, David Hanley, Joe Mambretti, Alex Szalay, Ani Thakar, Jan vandenBerg, Alainna Wonders |
SC | 7 |
| 2006 | Data management and query - Estimating query result sizes for proxy caching in scientific database federationsabstractIn a proxy cache for federations of scientific databases it is important to estimate the size of a query before making a caching decision. With accurate estimates, near-optimal cache performance can be obtained. On the other extreme, inaccurate estimates can render the cache totally ineffective. We present classification and regression over templates (CAROT), a general method for estimating query result sizes, which is suited to the resource-limited environment of proxy caches and the distributed nature of database federations. CAROT estimates query result sizes by learning the distribution of query results, not by examining or sampling data, but from observing workload. We have integrated CAROT into the proxy cache of the National Virtual Observatory (NVO) federation of astronomy databases. Experiments conducted in the NVO show that CAROT dramatically outperforms conventional estimation techniques and provides near-optimal cache performance. Tanu Malik, Randal C. Burns, Nitesh V. Chawla, Alex Szalay |
SC | 4 |
| 2006 | Poster reception - Harnessing grid resources to enable the dynamic analysis of large astronomy datasetsabstractAstronomy datasets are generally terabytes in size and contain hundreds of millions of objects separated into millions of files-factors which makes many analyses impractical to perform on small computers. The key question we answer in this paper is: How can we leverage Grid resources to make the analysis of large astronomy datasets a reality for the astronomy community? To address this question, we have developed a Web Services-based system, AstroPortal, that uses grid computing to federate large computing and storage resources for dynamic analysis of large datasets. Building on the GT4, we have built a prototype and implemented a first analysis, stacking, that sums multiple regions of the sky, a function that can help both identify variable sources and detect faint objects. AstroPortal gives the astronomy community a new tool to advance their research and to open new doors to opportunities never before possible on such a large scale. Ioan Raicu, Ian T. Foster, Alex Szalay |
SC | 3 |
| 2006 | Data analysis tools for sensor-based scienceabstractScience is increasingly driven by data collected automatically from arrays of inexpensive sensors. The collected data volumes require a different approach from the scientists' current Excel spreadsheet storage and analysis model. Spreadsheets work well for small data sets; but scientists want high level summaries of their data for various statistical analyses without sacrificing the ability to drill down to every bit of the raw data. This demonstration describes our prototype data analysis system that is suitable for browsing and visualization - like a spreadsheet - but scalable to much larger data sets. Stuart Ozer, Jim Gray 0001, Alex Szalay, Andreas Terzis, Razvan Musaloiu-Elefteri, Katalin Szlavecz, Randal C. Burns, Joshua Cogan |
SenSys | 3 |
| 2006 | Data mining middleware for wide-area high-performance networks
Robert L. Grossman, Yunhong Gu, David Hanley, Michal Sabala, Joe Mambretti, Alex Szalay, Ani Thakar, Kazumi Kumazoe, Yuji Oie, Yoonjoo Kwon, Woojin Seok |
Future Gener. Comput. Syst. | 6 |
| 2005 | When Database Systems Meet the Grid
María A. Nieto-Santisteban, Jim Gray 0001, Alex Szalay, James Annis, Ani Thakar, William O'Mullane |
CIDR | 3 |
| 2005 | Batch is Back: CasJobs, Serving Multi-TB Data on the WebabstractThe Sloan Digital Sky Survey (SDSS) science database describes over 230 million objects and is over 1.6 TB in size. The SDSS Catalog Archive Server (CAS) provides several levels of query interface to the SDSS data via the SkyServer website. Most queries execute in seconds or minutes. However, some queries can take hours or days, either because they require non-index scans of the largest tables, or because they request very large result sets, or because they represent very complex aggregations of the data. These "monster queries" not only take a long time, they also affect response times for everyone else - one or more of them can clog the entire system. To ameliorate this problem, we developed a multiserver multiqueue batch job submission, execution, and tracking system for the CAS called CasJobs. The transfer of very large result sets from queries over the network is another serious problem. Statistics suggested that much of this data transfer is unnecessary; users would prefer to store results locally in order to allow further joins and filtering. To allow local analysis, a system was developed that gives users their own personal databases (MyDB) at the server side. Users may transfer data to their MyDB, and then perform further analysis before extracting it to their own machine. MyDB tables also provide a convenient way to share results of queries with collaborators without downloading them. CasJobs is built using SOAP XML Web services and has been in operation since May 2004. William O'Mullane, Nolan Li, María A. Nieto-Santisteban, Alex Szalay, Ani Thakar |
ICWS | 4 |
| 2003 | SkyQuery: A Web Service Approach to Federate Databases
Tanu Malik, Alex Szalay, Tamás Budavári, Ani Thakar |
CIDR | 2 |
| 2002 | The SDSS skyserver: public access to the sloan digital sky server dataabstractThe SkyServer provides Internet access to the public Sloan Digital Sky Survey (SDSS) data for both astronomers and for science education. This paper describes the SkyServer goals and architecture. It also describes our experience operating the SkyServer on the Internet. The SDSS data is public and well-documented so it makes a good test platform for research on database algorithms and performance. Alex Szalay, Jim Gray 0001, Ani Thakar, Peter Z. Kunszt, Tanu Malik, M. Jordan Raddick, Christopher Stoughton, Jan vandenBerg |
SIGMOD Conference | 1 |
| 2000 | Designing and Mining Multi-Terabyte Astronomy Archives: The Sloan Digital Sky SurveyabstractThe next-generation astronomy digital archives will cover most of the sky at fine resolution in many wavelengths, from X-rays, through ultraviolet, optical, and infrared. The archives will be stored at diverse geographical locations. One of the first of these projects, the Sloan Digital Sky Survey (SDSS) is creating a 5-wavelength catalog over 10,000 square degrees of the sky (see http://www.sdss.org/). The 200 million objects in the multi-terabyte database will have mostly numerical attributes in a 100+ dimensional space. Points in this space have highly correlated distributions. Alex Szalay, Peter Z. Kunszt, Ani Thakar, Jim Gray 0001, Donald R. Slutz, Robert J. Brunner |
SIGMOD Conference | 1 |
| 1999 | Astronomical archives of the future: a Virtual ObservatoryabstractAstronomy is entering a new era as multiple, large area, digital sky surveys are in production. The resulting datasets are truly remarkable in their own right; however, a revolutionary step arises in the aggregation of complimentary multi-wavelength surveys (i.e. the cross-identification of a billion sources). Federating these different datasets, however, is an extremely challenging task. With this task in mind, we have identified several areas where community standardization can provide enormous benefits in order to develop the techniques and technologies necessary to solve the problems inherent in federating these large databases, as well as the mining of the resultant aggregate data. Several of these areas are domain specific, however, the majority of them are not. We feel that the inclusion of non-astronomical partnerships can provide tremendous insights. Alex Szalay, Robert J. Brunner |
Future Gener. Comput. Syst. | 1 |