VLDB 2026 Research / reviewers in the wild / expert
Jonathan P. Arnold
dblp:23/2575
· DBLP profile ↗
17ranked-venue papers
0as first author
1since 2021 · last 2025
0000-0001-5845-1048ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 1 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Bioinformatics and computational biology · 99% Computational science and engineering · 1% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
metabolomics |
0.4 | 1 | 2020 | RTExtract: time-series NMR spectra quantification based on 3D surface ridge tracking · Bioinform. 2020 |
Bioinformatics and computational biology › genomics
physical mapping |
0.1 | 7 | 2001 | Physical mapping with automatic capture of hybridization data · Bioinform. 2001 Reconstructing distances in physical maps of chromosomes with nonoverlapping probes · RECOMB 2000 ODS_BOOTSTRAP: assessing the statistical reliability of physical maps by bootstrap resampling · Comput. Appl. Biosci. 1994 |
Bioinformatics and computational biology
genomics |
0.1 | 2 | 2001 | Physical mapping with automatic capture of hybridization data · Bioinform. 2001 Reconstructing distances in physical maps of chromosomes with nonoverlapping probes · RECOMB 2000 |
Bioinformatics and computational biology › sequence analysis › sequence assembly
contig assembly |
0.0 | 3 | 2001 | Physical mapping with automatic capture of hybridization data · Bioinform. 2001 CMAP: contig mapping and analysis package, a relational database for chromosome reconstruction · Comput. Appl. Biosci. 1992 PCAP: probe choice and analysis package - a set of programs to aid in choosing synthetic oligomers for contig mapping · Comput. Appl. Biosci. 1993 |
Bioinformatics and computational biology › genome annotation
gene prediction |
0.0 | 1 | 2001 | An analysis of gene-finding programs for Neurospora crassa · Bioinform. 2001 |
Bioinformatics and computational biology › genome annotation › gene prediction
gene prediction evaluation |
0.0 | 1 | 2001 | An analysis of gene-finding programs for Neurospora crassa · Bioinform. 2001 |
Bioinformatics and computational biology › genome annotation
gene structure prediction |
0.0 | 1 | 2001 | An analysis of gene-finding programs for Neurospora crassa · Bioinform. 2001 |
Bioinformatics and computational biology › genomics › physical mapping
clone ordering |
0.0 | 2 | 1994 | ODS_BOOTSTRAP: assessing the statistical reliability of physical maps by bootstrap resampling · Comput. Appl. Biosci. 1994 ODS: ordering DNA sequences - a physical mapping algorithm based on simulated annealing · Comput. Appl. Biosci. 1993 |
Parallel and multicore computing
parallel algorithms |
0.0 | 1 | 1996 | PARODS - a study of parallel algorithms for ordering DNA sequences · Comput. Appl. Biosci. 1996 |
Parallel and multicore computing › parallel algorithms › parallel combinatorial optimization
parallel simulated annealing |
0.0 | 1 | 1996 | PARODS - a study of parallel algorithms for ordering DNA sequences · Comput. Appl. Biosci. 1996 |
Computational science and engineering › statistical computing
bootstrap resampling |
0.0 | 1 | 1994 | ODS_BOOTSTRAP: assessing the statistical reliability of physical maps by bootstrap resampling · Comput. Appl. Biosci. 1994 |
Bioinformatics and computational biology › genomics › genomic data management
genome database |
0.0 | 1 | 1993 | Design of an Object-Oriented Database for Reverse Genetics · ISMB 1993 |
Bioinformatics and computational biology › sequence analysis › primer and probe design
oligonucleotide probe selection |
0.0 | 1 | 1993 | PCAP: probe choice and analysis package - a set of programs to aid in choosing synthetic oligomers for contig mapping · Comput. Appl. Biosci. 1993 |
Bioinformatics and computational biology
biological database |
0.0 | 1 | 1992 | CMAP: contig mapping and analysis package, a relational database for chromosome reconstruction · Comput. Appl. Biosci. 1992 |
Mathematical optimization
continuous optimization |
0.0 | 1 | 2000 | Reconstructing distances in physical maps of chromosomes with nonoverlapping probes · RECOMB 2000 |
Mathematical optimization › statistical estimation
maximum likelihood estimation |
0.0 | 1 | 2000 | Reconstructing distances in physical maps of chromosomes with nonoverlapping probes · RECOMB 2000 |
Data models and query languages
object-oriented database |
0.0 | 1 | 1993 | Design of an Object-Oriented Database for Reverse Genetics · ISMB 1993 |
Methods — techniques the papers use, named apart from their topics
ridge tracking · 0.9greedy algorithm · 0.93d surface analysis · 0.9maximum likelihood · 0.1continuous optimization · 0.1simulated annealing · 0.0simulation · 0.0hybridization intensity analysis · 0.0hidden markov model · 0.0comparative evaluation · 0.0markov chain decomposition · 0.0SIMD · 0.0MIMD · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MINE: a new way to design genetics experiments for discoveryabstractThe Maximally Informative Next Experiment or MINE is a new experimental design approach for experiments, such as those in omics, in which the number of effects or parameters p greatly exceeds the number of samples n (p > n). Classical experimental design presumes n > p for inference about parameters and its application to p > n can lead to over-fitting. To overcome p > n, MINE is an ensemble method, which makes predictions about future experiments from an existing ensemble of models consistent with available data in order to select the most informative next experiment. Its advantages are in exploration of the data for new relationships with n < p and being able to integrate smaller and more tractable experiments to replace adaptively one large classic experiment as discoveries are made. Thus, using MINE is model-guided and adaptive over time in a large omics study. Here, MINE is illustrated in two distinct multiyear experiments, one involving genetic networks in Neurospora crassa and a second one involving a genome-wide association study in Sorghum bicolor as a comparison to classic experimental design in an agricultural setting. Isaac Torres, Amanda Bouffier, Michael Skaro, Yue Wu 0029, Lauren Stupp, Jonathan P. Arnold, Y. Anny Chung, Heinz-Bernd Schüttler |
Briefings Bioinform. | 7 |
| 2020 | RTExtract: time-series NMR spectra quantification based on 3D surface ridge trackingabstractMOTIVATION: Time-series nuclear magnetic resonance (NMR) has advanced our knowledge about metabolic dynamics. Before analyzing compounds through modeling or statistical methods, chemical features need to be tracked and quantified. However, because of peak overlap and peak shifting, the available protocols are time consuming at best or even impossible for some regions in NMR spectra. RESULTS: We introduce Ridge Tracking-based Extract (RTExtract), a computer vision-based algorithm, to quantify time-series NMR spectra. The NMR spectra of multiple time points were formulated as a 3D surface. Candidate points were first filtered using local curvature and optima, then connected into ridges by a greedy algorithm. Interactive steps were implemented to refine results. Among 173 simulated ridges, 115 can be tracked (RMSD < 0.001). For reproducing previous results, RTExtract took less than 2 h instead of ∼48 h, and two instead of seven parameters need tuning. Multiple regions with overlapping and changing chemical shifts are accurately tracked. AVAILABILITY AND IMPLEMENTATION: Source code is freely available within Metabolomics toolbox GitHub repository (https://github.com/artedison/Edison_Lab_Shared_Metabolomics_UGA/tree/master/metabolomics_toolbox/code/ridge_tracking) and is implemented in MATLAB and R. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yue Wu 0029, Michael T. Judge, Jonathan P. Arnold, Suchendra M. Bhandarkar, Arthur Edison |
Bioinform. | 3 |
| 2015 | Secondary structural entropy in RNA switch (Riboswitch) identificationabstractBACKGROUND: RNA regulatory elements play a significant role in gene regulation. Riboswitches, a widespread group of regulatory RNAs, are vital components of many bacterial genomes. These regulatory elements generally function by forming a ligand-induced alternative fold that controls access to ribosome binding sites or other regulatory sites in RNA. Riboswitch-mediated mechanisms are ubiquitous across bacterial genomes. A typical class of riboswitch has its own unique structural and biological complexity, making de novo riboswitch identification a formidable task. Traditionally, riboswitches have been identified through comparative genomics based on sequence and structural homology. The limitations of structural-homology-based approaches, coupled with the assumption that there is a great diversity of undiscovered riboswitches, suggests the need for alternative methods for riboswitch identification, possibly based on features intrinsic to their structure. As of yet, no such reliable method has been proposed. RESULTS: We used structural entropy of riboswitch sequences as a measure of their secondary structural dynamics. Entropy values of a diverse set of riboswitches were compared to that of their mutants, their dinucleotide shuffles, and their reverse complement sequences under different stochastic context-free grammar folding models. Significance of our results was evaluated by comparison to other approaches, such as the base-pairing entropy and energy landscapes dynamics. Classifiers based on structural entropy optimized via sequence and structural features were devised as riboswitch identifiers and tested on Bacillus subtilis, Escherichia coli, and Synechococcus elongatus as an exploration of structural entropy based approaches. The unusually long untranslated region of the cotH in Bacillus subtilis, as well as upstream regions of certain genes, such as the sucC genes were associated with significant structural entropy values in genome-wide examinations. CONCLUSIONS: Various tests show that there is in fact a relationship between higher structural entropy and the potential for the RNA sequence to have alternative structures, within the limitations of our methodology. This relationship, though modest, is consistent across various tests. Understanding the behavior of structural entropy as a fairly new feature for RNA conformational dynamics, however, may require extensive exploratory investigation both across RNA sequences and folding models. Amirhossein Manzourolajdad, Jonathan P. Arnold |
BMC Bioinform. | 2 |
| 2004 | Quality of service for workflows and web service processes
Jorge Cardoso 0001, Amit P. Sheth, John A. Miller 0001, Jonathan P. Arnold, Krys J. Kochut |
J. Web Semant. | 4 |
| 2003 | IntelliGEN: A Distributed Workflow System for Discovering Protein-Protein Interactions
Krys J. Kochut, Jonathan P. Arnold, Amit P. Sheth, John A. Miller 0001, Eileen T. Kraemer, Ismailcem Budak Arpinar, Jorge Cardoso 0001 |
Distributed Parallel Databases | 2 |
| 2002 | J3DV: A Java-based 3D database visualization toolabstractAbstract Database visualization helps users process, interpret and act upon large stored data sets. In this paper, we present a Java‐based 3D database visualization tool called J3DV. The J3DV tool successfully solved the problem of data management faced by many other visualization systems by integrating multiple data sources with the visualization tool. The tool utilizes two‐level mapping to transform the data into intermediate data that can be used to render graphs. Intermediate data offers better performance with a two‐tier cache. This visualization tool presents a sound framework, which has good extensibility for plugging in new data sources, supporting new data models and visual presentation types and allowing new graph layout algorithms. Copyright © 2002 John Wiley & Sons, Ltd. John A. Miller 0001, Jonathan P. Arnold |
Softw. Pract. Exp. | 3 |
| 2001 | Physical mapping with automatic capture of hybridization dataabstractMOTIVATION: Contig maps are a type of physical map that show the native order of a set of overlapping genomic clones. Overlaps between clones can be detected by finding common sequences using a number of experimental protocols including hybridization of probes. All current mapping algorithms of which we are aware require that hybridizations be scored using a fixed number of discrete values (typically 0/1 or high/medium/low). When hybridization data is captured automatically using digital equipment, this provides the opportunity for hybridization intensities to be used in map construction. More fine-grained distinctions in the levels of hybridization may be exploited by algorithms to generate more accurate physical maps. RESULTS: We describe an approach to creating contig maps that uses measured hybridization intensities instead of data scored with a fixed number of discrete values. We describe and compare four algorithms for creating physical maps with hybridization intensities. Simulations using measured intensities sampled from actual data on Aspergillus nidulans indicate that using hybridization intensities rather than data that is automatically scored with respect to threshold values may yield more accurate physical maps. Suchendra M. Bhandarkar, Jonathan P. Arnold, Tongzhang Jiang |
Bioinform. | 3 |
| 2001 | An analysis of gene-finding programs for Neurospora crassaabstractMOTIVATION: Computational gene identification plays an important role in genome projects. The approaches used in gene identification programs are often tuned to one particular organism, and accuracy for one organism or class of organism does not necessarily translate to accurate predictions for other organisms. In this paper we evaluate five computer programs on their ability to locate coding regions and to predict gene structure in Neurospora crassa. One of these programs (FFG) was designed specifically for gene-finding in N.crassa, but the model parameters have not yet been fully 'tuned', and the program should thus be viewed as an initial prototype. The other four programs were neither designed nor tuned for N.crassa. RESULTS: We describe the data sets on which the experiments were performed, the approaches employed by the five algorithms: GenScan, HMMGene, GeneMark, Pombe and FFG, the methodology of our evaluation, and the results of the experiments. Our results show that, while none of the programs consistently performs well, overall the GenScan program has the best performance on sensitivity and Missing Exons (ME) while the HMMGene and FFG programs have good performance in locating the exons roughly. Additional work motivated by this study includes the creation of a tool for the automated evaluation of gene-finding programs, the collection of larger and more reliable data sets for N.crassa, parameterization of the model used in FFG to produce a more accurate gene-finding program for this species, and a more in-depth evaluation of the reasons that existing programs generally fail for N.crassa. AVAILABILITY: Data sets, the FFG program source code, and links to the other programs analyzed are available at http://jerry.cs.uga.edu/~wang/genefind.html. CONTACT: [email protected]. Eileen T. Kraemer, Jinhua Guo 0001, Samuel Hopkins, Jonathan P. Arnold |
Bioinform. | 5 |
| 2000 | Parallel Computation for Chromosome Reconstruction on a Cluster of WorkstationsabstractReconstructing a physical map of a chromosome from a genomic library presents a central computational problem in genetics. Physical map reconstruction in the presence of errors is a problem of high computational complexity which provides the motivation for parallel computing. Parallelization strategies for a maximum likelihood estimation-based approach to physical map reconstruction are presented. The estimation procedure entails gradient descent search for determining the optimal spacings between probes for a given probe ordering. The optimal probe ordering is determined using a stochastic optimization algorithm. A two-tier parallelization strategy is proposed wherein the gradient descent search is parallelized at the lower level and the stochastic optimization algorithm is simultaneously parallelized at the higher level. Implementation and experimental results on a distributed memory multiprocessor cluster running the Parallel Virtual Machine (PVM) environment are presented. Suchendra M. Bhandarkar, Salem Machaka, Sanjay Shete, Jonathan P. Arnold |
IPDPS | 4 |
| 2000 | Reconstructing distances in physical maps of chromosomes with nonoverlapping probesabstractWe present a new method for reconstructing the distances between probes in physical maps of chromosomes constructed by hybridizing pairs of clones under the so-called sampling-without-replacement protocol. In this protocol, which is simple, inexpensive, and has been used to successfully map several organisms, equal-length clones are hybridized against a clone-subset called the probes. The probes are chosen by a sequential process that is designed to generate a pairwise-nonoverlapping subset of the clones. We derive a likelihood function on probe spacings and orders for this protocol under a natural model of hybridization error, and describe how to reconstruct the most likely spacing for a given order under this objective using continuous optimization. The approach is tested on simulated data and real data from chromosome VI of Aspergillus nidulans. On simulated data we recover the true order and close to the true spacing; on the real data, for which the true order and spacing is unknown, we recover a probe order differing significantly from the published one. To our knowledge this is the first practical approach for computing a globally-optimal maximum-likelihood reconstruction of interprobe distances from clone-probe hybridization data. John D. Kececioglu, Sanjay Shete, Jonathan P. Arnold |
RECOMB | 3 |
| 1998 | Parallel Computing for Chromosome Reconstruction via Ordering of DNA Sequences
Suchendra M. Bhandarkar, Salem Machaka, Sridhar Chirravuri, Jonathan P. Arnold |
Parallel Comput. | 4 |
| 1996 | PARODS - a study of parallel algorithms for ordering DNA sequencesabstractA suite of parallel algorithms for ordering DNA sequences (termed PARODS) is presented. The algorithms in PARODS are based on an earlier serial algorithm, ODS, which is a physical mapping algorithm based on simulated annealing. Parallel algorithms for simulated annealing based on Markov chain decomposition are proposed and applied to the problem of physical mapping. Perturbation methods and problem-specific annealing heuristics are proposed and described. Implementations of parallel Single Instruction Multiple Data (SIMD) algorithms on a 2048 processor MasPar MP-2 system and implementations of parallel Multiple Instruction Multiple Data (MIMD) algorithms on an 8 processor Intel iPSC/860 system are presented. The convergence, speedup and scalability characteristics of the aforementioned algorithms are analyzed and discussed. The best SIMD algorithm is shown to have a speedup of approximately 1000 on the 2048 processor MasPar MP-2 system, whereas the best MIMD algorithm is shown to have a speedup of approximately 5 on the 8 processor Intel iPSC/860 system. Suchendra M. Bhandarkar, Sridhar Chirravuri, Jonathan P. Arnold |
Comput. Appl. Biosci. | 3 |
| 1994 | ODS_BOOTSTRAP: assessing the statistical reliability of physical maps by bootstrap resamplingabstractIn the program ODS_BOOTSTRAP we provide a methodology for quickly ordering clones in a genomic library into a physical map and for applying a statistical tool known as the bootstrap to assess the statistical reliability of a clonal ordering. Each clone is assigned a binary fingerprint by one of a variety of experimental approaches to physical mapping. For example, the binary fingerprints might be generated by hybridizing a panel of m probes to a library of n clones. The resulting n x m binary data matrix, X, is input to ODS_BOOTSTRAP, which utilizes the similarity in binary fingerprints of clones to construct a physical map. Under this particular implementation of bootstrap resampling, the m probes (or columns of the data matrix) are sampled randomly with replacement in the computer to generate a new n x m data matrix, X*, from which a second physical map is constructed. The resampling process is repeated 100 or more times to generate 100 or more X* matrices. The resulting 100 or more physical maps are compared with the original physical map based on the original data matrix X by counting how often links in the original physical map reappear. Three confidence statistics are introduced for each link in a physical map. The statistic C1 is defined as the percentage of time two neighboring clones on the original map reappear as neighbors under resampling. The statistic C2 is defined as the percentage of time that two neighboring clones i and j on the original map reappear as neighbors or that a clone with an identical binary fingerprint to clone i reappears as a neighbor to clone j. The statistic C3 is defined as the percentage of time that two neighboring clones on the original map reappear in the same contig under resampling. Rolf A. Prade, James Griffith, William E. Timberlake, Jonathan P. Arnold |
Comput. Appl. Biosci. | 5 |
| 1993 | Design of an Object-Oriented Database for Reverse Genetics
Krys J. Kochut, Jonathan P. Arnold, John A. Miller 0001, Walter D. Potter |
ISMB | 2 |
| 1993 | PCAP: probe choice and analysis package - a set of programs to aid in choosing synthetic oligomers for contig mappingabstractIn the program, PCAP, we provide a methodology for choosing synthetic oligonucleotide probes to be used in contig mapping experiments. The package serves the purpose of presenting a series of short oligonucleotides (8-12mers) that are chosen based on constraints with respect to frequency of occurrence within a particular genome and the G+C content of the oligonucleotides. The four programs contained within the package: (i) convert GenBank files to a format useable by the package; (ii) calculate trinucleotide and tetranucleotide frequencies in available sequence data on a particular species; (iii) present the user with upper and lower bounds on the frequencies of hybridization sites for oligonucleotide probes of length 8-12, (iv) allow the user to place constraints on site frequency and G+C content and provides a list of short probe sequences that fit these criteria. These sequences can then be synthetically produced and used in hybridization experiments to carry out contig mapping. A. Jamie Cuticchia, Jonathan P. Arnold, William E. Timberlake |
Comput. Appl. Biosci. | 2 |
| 1993 | ODS: ordering DNA sequences - a physical mapping algorithm based on simulated annealingabstractIn the program ODS we provide a methodology for quickly ordering random clones into a physical map. The process of ordering individual clones with respect to their position along a chromosome is based on the similarity of binary signatures assigned to each clone. This binary signature is obtained by hybridizing each clone to a panel of oligonucleotide probes. By using the fact that the amount of overlap between any two clones is reflected in the similarity of their binary signatures, it is possible to reconstruct a chromosome by minimizing the sum of linking distances between an ordered sequence of clones. Unlike other programs for physical mapping, ODS is very general in the types of data that can be utilized for chromosome reconstruction. Any trait that can be scored in a presence--absence manner, such as hybridized synthetic oligonucleotides, restriction endonuclease recognition sites or single copy landmarks, can be used for analysis. Furthermore, the computational requirements for the construction of large physical maps can be measured in a matter of hours on work-stations such as the VAX2000. A. Jamie Cuticchia, Jonathan P. Arnold, William E. Timberlake |
Comput. Appl. Biosci. | 2 |
| 1992 | CMAP: contig mapping and analysis package, a relational database for chromosome reconstructionabstractIn the contig mapping and analysis package, CMAP, we provide a foundation for reverse genetics by organizing information about DNA fragments obtained from an organism's genome into a physical map. The user can store information about a particular segment of DNA. This information can be both descriptive, such as any genes contained in a particular DNA fragment, or experimental, such as hybridization profiles or restriction digest patterns for comparison with other fragments. The package can then be instructed to update the physical map or provide information on a DNA fragment within the map, such as its location. The user interface is designed to minimize the learning curve associated with database usage, while eliminating the possibility of entering data outside the ranges of fields through error-checking protocols. Queries are currently accomplished by the use of dynamic SQL (structured query language), which gives the user the ability to build queries based on any combination of the attributes contained within the database without requiring that all possible queries be permanently programmed within the query software. In order to eliminate the need for knowledge of SQL, an interface was designed to allow users to build queries by menu choices. Thus, CMAP is a software package supporting a database for both the production and storage of a physical map as well as being the first step toward the production of a physical mapping workstation. A. Jamie Cuticchia, Jonathan P. Arnold, H. Brody, William E. Timberlake |
Comput. Appl. Biosci. | 2 |