VLDB 2026 Research / reviewers in the wild / expert
Stephen J. Smith
dblp:83/5500
· DBLP profile ↗
10ranked-venue papers
1as first author
1since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3Applied, interdisciplinary, general and emerging computing · 3Theory of computation · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Minimum Wiener index of triangulations and quadrangulations
Éva Czabarka, Trevor Olsen, Stephen J. Smith, László A. Székely |
Discret. Appl. Math. | 3 |
| 2017 | Probabilistic fluorescence-based synapse detectionabstractDeeper exploration of the brain's vast synaptic networks will require new tools for high-throughput structural and molecular profiling of the diverse populations of synapses that compose those networks. Fluorescence microscopy (FM) and electron microscopy (EM) offer complementary advantages and disadvantages for single-synapse analysis. FM combines exquisite molecular discrimination capacities with high speed and low cost, but rigorous discrimination between synaptic and non-synaptic fluorescence signals is challenging. In contrast, EM remains the gold standard for reliable identification of a synapse, but offers only limited molecular discrimination and is slow and costly. To develop and test single-synapse image analysis methods, we have used datasets from conjugate array tomography (cAT), which provides voxel-conjugate FM and EM (annotated) images of the same individual synapses. We report a novel unsupervised probabilistic method for detection of synapses from multiplex FM (muxFM) image data, and evaluate this method both by comparison to EM gold standard annotated data and by examining its capacity to reproduce known important features of cortical synapse distributions. The proposed probabilistic model-based synapse detector accepts molecular-morphological synapse models as user queries, and delivers a volumetric map of the probability that each voxel represents part of a synapse. Taking human annotation of cAT EM data as ground truth, we show that our algorithm detects synapses from muxFM data alone as successfully as human annotators seeing only the muxFM data, and accurately reproduces known architectural features of cortical synapse distributions. This approach opens the door to data-driven discovery of new synapse types and their density. We suggest that our probabilistic synapse detector will also be useful for analysis of standard confocal and super-resolution FM images, where EM cross-validation is not practical. Anish K. Simhal, Cecilia Aguerrebere, Forrest Collman, Joshua T. Vogelstein, Kristina D. Micheva, Richard J. Weinberg, Stephen J. Smith, Guillermo Sapiro |
PLoS Comput. Biol. | 7 |
| 2013 | The open connectome project data cluster: scalable analysis and vision for high-throughput neuroscienceabstract- neural connectivity maps of the brain-using the parallel execution of computer vision algorithms on high-performance compute clusters. These services and open-science data sets are publicly available at openconnecto.me. The system design inherits much from NoSQL scale-out and data-intensive computing architectures. We distribute data to cluster nodes by partitioning a spatial index. We direct I/O to different systems-reads to parallel disk arrays and writes to solid-state storage-to avoid I/O interference and maximize throughput. All programming interfaces are RESTful Web services, which are simple and stateless, improving scalability and usability. We include a performance evaluation of the production system, highlighting the effec-tiveness of spatial data organization. Randal C. Burns, Kunal Lillaney, Daniel R. Berger, Logan Grosenick, Karl Deisseroth, R. Clay Reid, William R. Gray Roncal, Priya Manavalan, Davi Bock, Narayanan Kasthuri, Michael M. Kazhdan, Stephen J. Smith, Dean Kleissas, Eric A. Perlman, Kwanghun Chung, Nicholas C. Weiler, Jeff Lichtman, Alex Szalay, Joshua T. Vogelstein, R. Jacob Vogelstein |
SSDBM | 12 |
| 2013 | Automated Analysis of a Diverse Synapse PopulationabstractSynapses of the mammalian central nervous system are highly diverse in function and molecular composition. Synapse diversity per se may be critical to brain function, since memory and homeostatic mechanisms are thought to be rooted primarily in activity-dependent plastic changes in specific subsets of individual synapses. Unfortunately, the measurement of synapse diversity has been restricted by the limitations of methods capable of measuring synapse properties at the level of individual synapses. Array tomography is a new high-resolution, high-throughput proteomic imaging method that has the potential to advance the measurement of unit-level synapse diversity across large and diverse synapse populations. Here we present an automated feature extraction and classification algorithm designed to quantify synapses from high-dimensional array tomographic data too voluminous for manual analysis. We demonstrate the use of this method to quantify laminar distributions of synapses in mouse somatosensory cortex and validate the classification process by detecting the presence of known but uncommon proteomic profiles. Such classification and quantification will be highly useful in identifying specific subpopulations of synapses exhibiting plasticity in response to perturbations from the environment or the sensory periphery. Brad Busse, Stephen J. Smith |
PLoS Comput. Biol. | 2 |
| 2012 | Sub-diffraction Limit Localization of Proteins in Volumetric Space Using Bayesian Restoration of Fluorescence Images from Ultrathin SpecimensabstractPhoton diffraction limits the resolution of conventional light microscopy at the lateral focal plane to 0.61λ/NA (λ = wavelength of light, NA = numerical aperture of the objective) and at the axial plane to 1.4nλ/NA(2) (n = refractive index of the imaging medium, 1.51 for oil immersion), which with visible wavelengths and a 1.4NA oil immersion objective is -220 nm and -600 nm in the lateral plane and axial plane respectively. This volumetric resolution is too large for the proper localization of protein clustering in subcellular structures. Here we combine the newly developed proteomic imaging technique, Array Tomography (AT), with its native 50-100 nm axial resolution achieved by physical sectioning of resin embedded tissue, and a 2D maximum likelihood deconvolution method, based on Bayes' rule, which significantly improves the resolution of protein puncta in the lateral plane to allow accurate and fast computational segmentation and analysis of labeled proteins. The physical sectioning of AT allows tissue specimens to be imaged at the physical optimum of modern high NA plan-apochormatic objectives. This translates to images that have little out of focus light, minimal aberrations and wave-front distortions. Thus, AT is able to provide images with truly invariant point spread functions (PSF), a property critical for accurate deconvolution. We show that AT with deconvolution increases the volumetric analytical fidelity of protein localization by significantly improving the modulation of high spatial frequencies up to and potentially beyond the spatial frequency cut-off of the objective. Moreover, we are able to achieve this improvement with no noticeable introduction of noise or artifacts and arrive at object segmentation and localization accuracies on par with image volumes captured using commercial implementations of super-resolution microscopes. Gordon Wang, Stephen J. Smith |
PLoS Comput. Biol. | 2 |
| 1998 | Sorting Algorithms
Bruce M. Maggs, C. Greg Plaxton, Stephen J. Smith, Marco Zagha |
Theory Comput. Syst. | 3 |
| 1994 | Handwritten Character Classification Using Nearest Neighbor in Large DatabasesabstractShows that systems built on a simple statistical technique and a large training database can be automatically optimized to produce classification accuracies of 99% in the domain of handwritten digits. It is also shown that the performance of these systems scale consistently with the size of the training database, where the error rate is cut by more than half for every tenfold increase in the size of the training set from 10 to 100,000 examples. Three distance metrics for the standard nearest neighbor classification system are investigated: a simple Hamming distance metric, a pixel distance metric, and a metric based on the extraction of penstroke features. Systems employing these metrics were trained and tested on a standard, publicly available, database of nearly 225,000 digits provided by the National Institute of Standards and Technology. Additionally, a confidence metric is both introduced by the authors and also discovered and optimized by the system. The new confidence measure proves to be superior to the commonly used nearest neighbor distance.> Stephen J. Smith, Mario O. Bourgoin, Karl Sims, Harry Voorhees |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1993 | A practical external sort for shared disk MPPsabstractNo abstract available. Xiqing Li, Gordon Linoff, Stephen J. Smith, Craig Stanfill, Kurt H. Thearling |
SC | 3 |
| 1992 | An Improved Supercomputer Sorting BenchmarkabstractThe authors propose that the process of sorting be more formally adopted as a performance benchmark for commercial supercomputer applications. To this end they have investigated the use of entropy as a measure of data distribution and propose that it, along with larger datasets, be added to existing sorting benchmarks (such as NAS). Some of the key points in adopting such a benchmark are presented, and the results of applying such a benchmark to the CM-5 supercomputer are discussed. As a result of carefully examining this problem, the authors were able to sort 1 billion 32-b keys in less than 17 s on a 1024 processor CM-5.> Kurt H. Thearling, Stephen J. Smith |
SC | 2 |
| 1991 | A Comparison of Sorting Algorithms for the Connection Machine CM-2abstractWe have implemented three parallel sorting algorithms on the Connection Machine Supercomputer model CM-2: B atcher's bitonic sort, a parallel radix sor~and a sample sort similar to Reif and Valiant's flashsort.We have also evaluated the implementation of many other sorting algorithms proposed in the literature.Our computational experiments show that the sample sort algorithm, which is a theoretically efficient "randomized" algorithm, is the fastest of the three algorithms on large data sets.On a 64Kprocessor CM-2, our sample sort implementation can sort 32 x 106 64-bit keys in 5.1 seconds, which is over 10 times faster than the CM-2 library sort.Our implementation of radix sort, although not as fast on large data sets, is deterministic, much simpler to code, stable, faster with small keys, and faster on small data sets (few elements per processor), Our implementation of bitonic sor~which is pipelined to use all the hypercube wires simultaneously, is the least efficient of the three on large data sets, but is the most efficient on small data sets, and is considerably more space efficient.This paper analyzes the three algorithms in detail and discusses many practical issues that led us to the particular implementations. Guy E. Blelloch, Charles E. Leiserson, Bruce M. Maggs, C. Greg Plaxton, Stephen J. Smith, Marco Zagha |
SPAA | 5 |